Skip to content
OpenAIOpenAI

Software Engineer, Core Network Engineering

Builds and operates high-performance networking infrastructure for OpenAI's large-scale AI training and inference, focusing on host networking, datacenter fabrics, and WAN systems. Optimizes latency, reliability, and scalability using technologies like RDMA, InfiniBand, and RoCE; requires strong systems programming in C++, Python, or Go.

About the job

Responsibilities

  • Design, build, and operate networking systems that support large-scale AI training and inference infrastructure
  • Improve performance, reliability, and scalability across host networking, datacenter fabrics, and WAN systems
  • Develop automation for provisioning, configuration management, validation, upgrades, and lifecycle management of networking infrastructure
  • Build tooling and observability systems for network health, performance analysis, debugging, and automated remediation
  • Optimize network performance across technologies such as RDMA, RoCE, InfiniBand, Ethernet, and high-performance GPU interconnects
  • Define and operationalize networking protocols, readiness criteria, and continuous validation systems
  • Partner closely with compute, storage, hardware, and infrastructure teams to ensure networking scales predictably with fleet growth
  • Contribute to architecture decisions around topology design, capacity planning, failure domains, and network reliability
  • Diagnose complex distributed systems and networking issues across large heterogeneous compute environments

Requirements

  • Experience building or operating large-scale networking or distributed systems infrastructure
  • Comfortable working close to the hardware/software boundary
  • Experience with Linux networking, kernel systems, NICs, RDMA, or performance-sensitive infrastructure software
  • Worked with high-performance networking technologies such as InfiniBand, RoCE, DPDK, or large-scale Ethernet fabrics
  • Experience with datacenter networking, WAN systems, or host networking stacks
  • Enjoy debugging complex systems and performance bottlenecks across multiple layers of the stack
  • Comfortable writing production software in languages such as C++, Python, or Go
  • Strong systems fundamentals across networking, operating systems, distributed systems, or infrastructure engineering

Skills

Linux Networking, Rdma, InfiniBand, Roce, Dpdk, C++, Python, Go, Kubernetes, Ethernet

Perplexity

Perplexity

San Francisco, CA
Member of Technical Staff
$220k+/yrRemote4+ YOEDevOps / SRE

Build and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.

Firecrawl

Firecrawl

San Francisco, CA

Cloud DevOps Engineer
$240k+/yrHybrid5+ YOEDevOps / SRE

Build and operate the Kubernetes-based cloud and on-premises infrastructure powering large-scale crawling, search, and ML workloads. The role requires 5+ years in DevOps, platform engineering, or cloud infrastructure, with strong Kubernetes, cloud, Docker, Terraform, and distributed-systems experience.

Fireworks AI

Fireworks AI

San Mateo, CA
Member of Technical Staff - Reliability Engineering
$240k+/yrHybrid5+ YOEDevOps / SRE

Owns reliability standards, incident management, observability, failure testing, and automation for a high-throughput AI infrastructure platform. The role requires deep Linux, networking, software, cloud-native, and distributed-systems experience, along with the ability to influence teams across the organization.

Tessera Labs

Tessera Labs

San Francisco, CA

AI Platform Engineer
$200k+/yrRemote5+ YOEDevOps / SRE

Build and own production-grade AI agent infrastructure across multiple clouds, with responsibility for Kubernetes, Terraform, observability, security, reliability, and automation. Requires 5+ years of cloud infrastructure experience and strong CI/CD, networking, and production operations expertise.

Crusoe

Crusoe

United States

Electrical Field Engineer - Data Center
$196k+/yrRemote5+ YOEDevOps / SRE

Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.