# Site Reliability Engineer

**Company:** [Arena](https://hotfix.jobs/companies/arena)
**Location:** California
**Role:** DevOps / SRE
**Experience:** 6+ years
**Skills:** Go, Rust, Kubernetes, Terraform, Postgres, Redis, AWS, GCP, api gateways, Distributed Systems
**Posted:** 2026-07-16

> Build and operate the scalable, low-latency infrastructure powering Arena's real-world AI model evaluation platform, including API gateways, observability, and enterprise features for frontier model routing and evaluation.

## Job Description

## What You'll Do
- Build and orchestrate API-based products from the ground up. Design and implement low-latency, high-reliability APIs for leaderboards, models, and arenas.
- Ship enterprise-grade infrastructure: rate limiting, authentication, usage metering, cost attribution, audit logging, and SOC 2 compliance.
- Build deep observability with distributed tracing, latency breakdowns, token-level usage tracking, and real-time dashboards.
- Build AI-centered products by integrating with the core evaluation platform, Arena data, and customer-specific benchmarks. Collaborate with the research team to turn novel ideas into full-featured products.

## What We're Looking For
- 6+ years of backend engineering experience, with meaningful time spent on distributed systems, infrastructure, or developer-facing platforms.
- Strong proficiency in Go and/or Rust, with hands-on experience building high-throughput APIs or proxy/gateway systems.
- Experience with LLM provider APIs (OpenAI, Anthropic, Google, etc.) and understanding of challenges like streaming, token management, rate limits, model-specific quirks.
- Solid cloud infrastructure skills with AWS or GCP, Kubernetes, Terraform, and database systems like Postgres and Redis.
- A product-oriented mindset focused on developer experience.
- Comfort with ambiguity in a startup environment.

## Nice to Have
- Experience building API gateways, proxies, or developer tools (Bifrost, Kong, Envoy, Tyk, or custom).
- Background in ML infrastructure, model serving, or evaluation frameworks.
- Experience building enterprise-ready features: SSO, RBAC, audit logs, multi-tenancy.
- Experience building billing infrastructure around Stripe, Metronome and Orb.
- Familiarity with the modern AI infra stack (vLLM, LiteLLM, LangChain, etc.).

## Similar roles

- [Senior Software Engineer, Dev Tools](https://hotfix.jobs/jobs/dc931946-25e4-4987-a47b-f88c09268fd8) - Airbnb - Remote - $196k – $230k/yr
- [Network Production Engineering Lead](https://hotfix.jobs/jobs/51065552-0f77-47bc-a8f2-edf2c84caecc) - Fluidstack - San Francisco, CA - $242k – $284k/yr
- [AI Enablement Engineer](https://hotfix.jobs/jobs/307d7d5b-7858-4c37-bce0-624c01b780ac) - Sprinter Health - San Francisco, CA - $180k – $260k/yr
- [Senior Performance Engineer](https://hotfix.jobs/jobs/6629d9ee-877d-4387-8d8f-e75835fd70c0) - Crusoe - San Francisco, CA - $170k – $205k/yr
- [Senior Software Engineer](https://hotfix.jobs/jobs/fe40327f-1648-4f0e-ba97-f600e8bae0ef) - Grafana Labs - Remote - $154k – $185k/yr

**Apply:** https://hotfix.jobs/jobs/b130465a-b419-4c26-943c-d52f0a07b8f5
**Canonical:** https://hotfix.jobs/jobs/b130465a-b419-4c26-943c-d52f0a07b8f5