Build and operate the scalable, low-latency infrastructure powering Arena's real-world AI model evaluation platform, including API gateways, observability, and enterprise features for frontier model routing and evaluation.
Salary not listed
Hybrid6+ YOEDevOps / SRE
About the role
What You'll Do
Build and orchestrate API-based products from the ground up. Design and implement low-latency, high-reliability APIs for leaderboards, models, and arenas.
Build deep observability with distributed tracing, latency breakdowns, token-level usage tracking, and real-time dashboards.
Build AI-centered products by integrating with the core evaluation platform, Arena data, and customer-specific benchmarks. Collaborate with the research team to turn novel ideas into full-featured products.
What We're Looking For
6+ years of backend engineering experience, with meaningful time spent on distributed systems, infrastructure, or developer-facing platforms.
Strong proficiency in Go and/or Rust, with hands-on experience building high-throughput APIs or proxy/gateway systems.
Experience with LLM provider APIs (OpenAI, Anthropic, Google, etc.) and understanding of challenges like streaming, token management, rate limits, model-specific quirks.
Solid cloud infrastructure skills with AWS or GCP, Kubernetes, Terraform, and database systems like Postgres and Redis.
A product-oriented mindset focused on developer experience.
Comfort with ambiguity in a startup environment.
Nice to Have
Experience building API gateways, proxies, or developer tools (Bifrost, Kong, Envoy, Tyk, or custom).
Background in ML infrastructure, model serving, or evaluation frameworks.
Experience building enterprise-ready features: SSO, RBAC, audit logs, multi-tenancy.
Experience building billing infrastructure around Stripe, Metronome and Orb.
Familiarity with the modern AI infra stack (vLLM, LiteLLM, LangChain, etc.).
Skills
GoRustKubernetesTerraformPostgresRedisAWSGCPapi gatewaysDistributed Systems
Senior engineer building and operating foundational Dev Tools infrastructure at Airbnb, including cloud development environments, Kubernetes platforms for AI agents, source control, code review, and polyglot monorepo tooling. Requires 5+ years building high-scale distributed systems with a focus on developer productivity, reliability, and cost efficiency.
196k – 230k/yr
Remote5+ YOEDevOps / SRE
Network Production Engineering Lead
FluidstackSan Francisco, CA
Lead the network production engineering team responsible for availability, performance, and automation of fabrics supporting 100k+ accelerator clusters at massive scale. Own SLOs, build remediation automation, and set operating models between design and site teams.
242k – 284k/yr
On-site7+ YOEDevOps / SRE
AI Enablement Engineer
Sprinter HealthSan Francisco, CA
Build and enable company-wide AI adoption by creating agents, workflows, prompt libraries, evaluation frameworks, and training programs. Partner with engineering, clinical, operations and other teams to identify opportunities, deliver production AI tools, and ensure safe, measurable impact in a regulated healthcare environment.
180k – 260k/yr
Hybrid7+ YOEDevOps / SRE
Senior Performance Engineer
CrusoeSan Francisco, CA
Senior Performance Engineer responsible for Linux kernel optimization, system benchmarking, and low-level performance tuning to enhance Crusoe's AI cloud infrastructure. Requires deep Linux kernel expertise, proficiency in Go/C/C++, and hands-on experience with performance optimization in complex environments.
170k – 205k/yr
On-site5+ YOEDevOps / SRE
Senior Software Engineer
Grafana LabsUnited States
Senior SRE embedded with Mimir and Loki squads to own production reliability and SLOs for Grafana Cloud's high-SLA database products (Mimir, Loki, Tempo, Pyroscope) running on AWS/GCP/Azure. Requires 6+ years engineering experience including 3+ in SRE/production engineering, strong Kubernetes and multi-tenant systems experience, and on-call participation.