Infrastructure Engineer building low-latency, high-reliability APIs, streaming gateways, and observability for Arena's real-world AI model evaluation platform. Requires 4+ years backend/distributed systems experience with Go/Rust, LLM APIs, and cloud infra (K8s/Terraform).
Salary not listed
Hybrid4+ YOEDevOps / SRE
About the role
What You'll Do
Build API-based products from the ground up. Design and implement low-latency, high-reliability APIs for leaderboards, models, and arenas.
Solve hard streaming problems. Handle SSE/streaming responses across heterogeneous providers, including partial failure recovery, mid-stream fallback, and consistent response normalization.
Ship enterprise-grade infrastructure. Build the systems enterprise customers expect: rate limiting, authentication, usage metering, cost attribution, audit logging, and SOC 2 compliance.
Build deep observability. Instrument infrastructure with distributed tracing, latency breakdowns, token-level usage tracking, and real-time dashboards so customers (and we) can see exactly what's happening.
Build AI-centered products. Integrate with our core evaluation platform, Arena data, and customer-specific benchmarks. Collaborate with the research team to turn novel ideas into full-featured products.
Flex across the stack. Contribute to the backend of our Leaderboards and Evals platforms when needed, helping unify our public and private data architectures.
Requirements
4+ years of backend engineering experience, with meaningful time spent on distributed systems, infrastructure, or developer-facing platforms.
Strong proficiency in Go and/or Rust, with hands-on experience building high-throughput APIs or proxy/gateway systems.
Experience with LLM provider APIs (OpenAI, Anthropic, Google, etc.) and a working understanding of the challenges: streaming, token management, rate limits, model-specific quirks.
Solid cloud infrastructure skills — comfortable with AWS or GCP, Kubernetes, Terraform, and database systems like Postgres and Redis.
A product-oriented mindset. You think about the developer experience of your APIs, not just the implementation. You ask "why" before "how."
Comfort with ambiguity. We're a startup. Scope is fluid, context shifts, and you'll wear many hats.
Nice to Have
Experience building API gateways, proxies, or developer tools (Bifrost, Kong, Envoy, Tyk, or custom).
Background in ML infrastructure, model serving, or evaluation frameworks.
Experience building enterprise-ready features: SSO, RBAC, audit logs, multi-tenancy.
Experience building billing infrastructure around systems like Stripe, Metronome and Orb.
Familiarity with the modern AI infra stack (vLLM, LiteLLM, LangChain, etc.).
Systems Integration Engineer, Build Systems | Consumer Devices
OpenAISan Francisco, CA
Build and evolve Bazel, Yocto, and Buildkite-based CI systems for OpenAI consumer device software. Focus on hermetic builds, remote caching, test optimization, observability, and AI-powered failure analysis to accelerate reliable shipping. Requires 5+ years building developer infrastructure at scale.
293k – 325k/yr
Hybrid5+ YOEDevOps / SRE
Software Engineer, CI Platform Infrastructure
AirbnbUnited States
Build and optimize a next-generation CI platform infrastructure for workflow orchestration, scheduling, caching, and autoscaling to accelerate software development for engineers and AI coding agents at scale. Requires interest in distributed systems and knowledge of Kubernetes, EC2, Golang, and Docker.
162k – 190k/yr
RemoteDevOps / SRE
Release Management Engineer, Mobile
PhantomUnited States
Own and continuously improve the end-to-end release process for Phantom's iOS and Android mobile apps, including scheduling, CI/CD automation, app store submissions, rollout monitoring, and cross-team coordination.
Salary not listed
Remote3+ YOEDevOps / SRE
Software Engineer
xAIPalo Alto, CA
Build and optimize large-scale distributed systems powering xAI's massive supercomputing clusters for AI training. Requires strong systems programming in Rust/C++ and deep Kubernetes/Linux expertise.
180k – 440k/yr
On-site5+ YOEDevOps / SRE
AI Infrastructure Engineer, Sandbox Platform
Scale AISan Francisco, CA +2
Build and evolve a secure, high-performance agent sandboxing platform for code execution. Combine deep systems expertise in isolation/virtualization with strong focus on developer experience, APIs, and internal partnerships. Requires 4+ years in high-performance systems software.