Site Reliability Engineer
Owns production reliability (SLOs, monitoring, incident response) and platform engineering (CI/CD, infrastructure as code) for AI developer tools used by hundreds of thousands. Requires deep production systems experience, strong coding skills, and cloud proficiency.
About the job
What You'll Accomplish
Production Reliability
- Define and own SLOs, SLIs, and error budgets for Devin and Windsurf.
- Build the monitoring, alerting, and observability systems that give the team a clear, honest picture of service health at all times.
Incident Response and On-Call
- Lead incident response with speed and clarity.
- Run blameless postmortems that turn outages into durable improvements.
- Build the runbooks and tooling that make on-call sustainable and effective.
Platform Engineering and CI/CD
- Own the deployment pipelines, release infrastructure, and internal developer tooling that let the team ship fast without breaking things.
- Reduce toil systematically so engineers spend time on work that matters.
Infrastructure as Code
- Manage cloud infrastructure through code.
- Build reproducible, auditable, version-controlled environments that scale with the product and eliminate configuration drift.
Capacity Planning and Performance
- Model growth, forecast resource needs, and ensure the infrastructure stays ahead of demand.
- Profile and improve system performance before users feel it.
Security and Reliability as One
- Treat security not as a separate concern but as a reliability requirement.
- Ensure that misconfigurations, vulnerabilities, and access failures are caught and remediated with the same urgency as outages.
Reliability Culture
- Partner closely with product and engineering teams to build reliability in from the start.
- Be the person who catches the single point of failure in the architecture review before it becomes a page at 2am.
Exceptional Candidates Have Demonstrated
- Deep experience running production systems at scale: SLOs, error budgets, on-call rotations, and incident command
- Strong software engineering fundamentals; SRE at Cognition means writing real code, not just configuring tools
- Proficiency with cloud infrastructure (AWS, GCP, or Azure), container orchestration (Kubernetes), and infrastructure as code (Terraform or equivalent)
- Experience building and owning CI/CD pipelines and deployment infrastructure for fast-moving product teams
- Strong observability instincts: knows how to instrument systems, build useful dashboards, and design alerts that surface signal without generating noise
- A track record of reducing toil systematically through automation, not just working around it
- Comfort owning incidents end to end: detection, triage, mitigation, resolution, and postmortem
- Enough product empathy to understand what reliability means from a user's perspective, not just an infrastructure one
- Experience with developer-facing products or platforms is a strong plus
Skills
Kubernetes, Terraform, AWS, GCP, Azure, CI/CD, SLOs, Slis, Observability, Incident Response
Similar jobs
DevOps / SRE jobsBuild and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.
Build IT workflow automation and security tooling using Go, Temporal, Kubernetes, Terraform, and shell scripting. The role also supports internal IT systems and integrations across endpoint management, identity, access, and security administration.
Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.
Build and operate AWS cloud, ML, LLM, RAG, and IoT infrastructure, including deployment platforms, data pipelines, vector search, observability, security, and cost controls. The role requires deep AWS experience and production experience with LLM-powered applications.