What You'll Do
- Own core infrastructure — Databases, AWS services, networking, CI/CD; the systems every engineer depends on
- Build for scale — Thousands of concurrent AI operations, sub-second response times, cost-effective under load
- Design observability that works — System-level errors AND business-level silent-failure detection; we find out something is broken from an alert, not from a person hours later
- Orchestrate agentic workloads — N parallel agent instances doing non-deterministic work, running reliably and cost-effectively as N grows
- Multiply developer velocity — The tooling and abstractions that let product engineers ship in days instead of weeks
- Own reliability end-to-end — SLOs, on-call, incident response, and the post-mortems that make it rarer next time
- Build the eval infrastructure — Systems that measure whether AI outputs are getting better, not just running
Requirements
- Software engineering experience with a platform, infrastructure, or SRE focus (level determined during interviews)
- Proficiency in Python, TypeScript, Go, or similar
- Production experience with cloud infrastructure (AWS), databases (Postgres), and observability tooling
- Track record of building systems and tooling that other engineers depend on
- Based in San Francisco or willing to relocate
Nice to Have
- Experience orchestrating AI/ML workloads at scale (agent frameworks, LLM inference infrastructure)
- Voice AI or real-time systems experience
- Deep AWS experience (ECS, Lambda, RDS, VPC design)
- Prior startup experience—especially at companies that scaled through hypergrowth
Compensation & Logistics
Salary: $140,000–$280,000 depending on experience + performance bonuses & equity
Location: San Francisco, in-office. Based in SF or willing to relocate.
Schedule: Monday–Friday, very early morning start, in-office five days a week.
Benefits: Uber commuter benefits; breakfast, lunch, and dinner provided; snacks, drinks, and coffee daily; free gym membership; health, dental, and vision insurance.