Site Reliability Engineer
Site Reliability Engineer responsible for the reliability, observability, performance, and security of a core AI agent platform including code sandboxes. Requires 5+ years software engineering experience (3+ in SRE/DevOps), strong CS fundamentals, expertise in containers, cloud IaC, monitoring, and distributed systems.
About the job
Responsibilities
- Design and maintain our production infrastructure on cloud platforms like AWS, GCP, or Azure
- Monitor and respond to system alerts and incidents using Grafana and Prometheus, ensuring high availability and a secure environment for our users' code
- Collaborate with developers to ensure new features and services are designed with scalability and reliability in mind
- Troubleshoot and resolve complex issues related to our infrastructure, networking, and the sandbox environment
- Participate in an on-call rotation to support our production systems
- Define and track SLIs/SLOs, manage error budgets, and proactively monitor distributed systems with logging and tracing
- Automate deployments, scaling, provisioning, and recovery tasks to reduce toil and build self-healing systems
- Lead incident response, conduct root-cause analysis, and facilitate blameless post-mortems to drive continual improvement
- Collaborate cross-functionally with product, engineering, and developer relations to ensure reliable releases and an outstanding developer experience
- Plan for capacity growth, forecast system usage, and contribute to safe release and change management processes
Requirements
- Strong computer science fundamentals, backed by a degree from a top-tier CS/EE program, or equivalent experience
- 5+ years of experience in software engineering, with at least 3 years focused explicitly on site reliability, DevOps, or infrastructure operations
- Strong programming skills in languages like Python or Go
- Deep expertise in containerization technologies such as Docker and Kubernetes
- Experience with cloud infrastructure and tools like Terraform and/or Pulumi
- Familiarity with monitoring and alerting tools like Prometheus, Grafana, or Datadog
- A solid understanding of networking, security, and Linux systems administration
- Experience designing, scaling, and maintaining distributed systems (backend platforms, APIs, or front-end infrastructure)
- Proficiency in implementing observability frameworks (metrics, logging, tracing) and aligning reliability goals with developer velocity
- Hands-on experience managing incidents, running on-call operations, and producing actionable post-mortems
- Ability to mentor engineers and influence reliability practices across teams, especially for front-end infrastructure and performance
Nice-to-Haves
- Experience with chaos engineering techniques, front-end observability tools (e.g., Sentry, RUM, synthetic monitoring), or building CI/CD pipelines for front-end delivery
Benefits
- Competitive salary and equity
- Comprehensive health, dental, and vision insurance for employee and dependents
- Opportunity to work on cutting-edge technology and make a real impact on the future of software engineering
- Daily catered lunch for all employees and a fridge full of your favorite snacks and drinks
- Onsite 4 days a week in San Francisco; Optional 1 day a week remote
Skills
Python, Go, Docker, Kubernetes, Terraform, Pulumi, Prometheus, Grafana, Datadog, AWS, GCP, Azure, Linux
Similar jobs
DevOps / SRE jobsBuild and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.
Build IT workflow automation and security tooling using Go, Temporal, Kubernetes, Terraform, and shell scripting. The role also supports internal IT systems and integrations across endpoint management, identity, access, and security administration.
Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.
Build and operate AWS cloud, ML, LLM, RAG, and IoT infrastructure, including deployment platforms, data pipelines, vector search, observability, security, and cost controls. The role requires deep AWS experience and production experience with LLM-powered applications.