# Infrastructure Engineer, TL

**Company:** [Arena](https://hotfix.jobs/companies/arena)
**Location:** California
**Role:** DevOps / SRE
**Experience:** 4+ years
**Skills:** Go, Rust, Kubernetes, Terraform, Postgres, Redis, AWS, GCP, api gateways, LLM APIs
**Posted:** 2026-07-19

> Infrastructure Engineer building low-latency, high-reliability APIs, streaming gateways, and observability for Arena's real-world AI model evaluation platform. Requires 4+ years backend/distributed systems experience with Go/Rust, LLM APIs, and cloud infra (K8s/Terraform).

## Job Description

## What You'll Do
- Build API-based products from the ground up. Design and implement low-latency, high-reliability APIs for leaderboards, models, and arenas.
- Solve hard streaming problems. Handle SSE/streaming responses across heterogeneous providers, including partial failure recovery, mid-stream fallback, and consistent response normalization.
- Ship enterprise-grade infrastructure. Build the systems enterprise customers expect: rate limiting, authentication, usage metering, cost attribution, audit logging, and SOC 2 compliance.
- Build deep observability. Instrument infrastructure with distributed tracing, latency breakdowns, token-level usage tracking, and real-time dashboards so customers (and we) can see exactly what's happening.
- Build AI-centered products. Integrate with our core evaluation platform, Arena data, and customer-specific benchmarks. Collaborate with the research team to turn novel ideas into full-featured products.
- Flex across the stack. Contribute to the backend of our Leaderboards and Evals platforms when needed, helping unify our public and private data architectures.

## Requirements
- 4+ years of backend engineering experience, with meaningful time spent on distributed systems, infrastructure, or developer-facing platforms.
- Strong proficiency in Go and/or Rust, with hands-on experience building high-throughput APIs or proxy/gateway systems.
- Experience with LLM provider APIs (OpenAI, Anthropic, Google, etc.) and a working understanding of the challenges: streaming, token management, rate limits, model-specific quirks.
- Solid cloud infrastructure skills — comfortable with AWS or GCP, Kubernetes, Terraform, and database systems like Postgres and Redis.
- A product-oriented mindset. You think about the developer experience of your APIs, not just the implementation. You ask "why" before "how."
- Comfort with ambiguity. We're a startup. Scope is fluid, context shifts, and you'll wear many hats.

## Nice to Have
- Experience building API gateways, proxies, or developer tools (Bifrost, Kong, Envoy, Tyk, or custom).
- Background in ML infrastructure, model serving, or evaluation frameworks.
- Experience building enterprise-ready features: SSO, RBAC, audit logs, multi-tenancy.
- Experience building billing infrastructure around systems like Stripe, Metronome and Orb.
- Familiarity with the modern AI infra stack (vLLM, LiteLLM, LangChain, etc.).

## Similar roles

- [Systems Integration Engineer, Build Systems | Consumer Devices](https://hotfix.jobs/jobs/4504a8ef-1c33-4742-b405-0358b1508d76) - OpenAI - San Francisco, CA - $293k – $325k/yr
- [Software Engineer, CI Platform Infrastructure](https://hotfix.jobs/jobs/fdfecbed-6550-4136-926b-f191623c3751) - Airbnb - Remote - $162k – $190k/yr
- [Release Management Engineer, Mobile](https://hotfix.jobs/jobs/7a1fb6c9-5c10-477c-9932-97c167947d8b) - Phantom - Remote
- [Software Engineer](https://hotfix.jobs/jobs/24949960-ef02-4961-8eb2-14082cd46e85) - xAI - Palo Alto, CA - $180k – $440k/yr
- [AI Infrastructure Engineer, Sandbox Platform](https://hotfix.jobs/jobs/4235c3a2-e042-48e4-8e5e-d6872fd48ad9) - Scale AI - San Francisco, CA - $180k – $225k/yr

**Apply:** https://hotfix.jobs/jobs/4c25d7f7-fad9-42fb-ae21-64920ca59d11
**Canonical:** https://hotfix.jobs/jobs/4c25d7f7-fad9-42fb-ae21-64920ca59d11