Staff Site Reliability Engineer
Staff SRE embedding early with product/platform teams to design reliability, observability, and scalability into systems across multi-cloud environments. Define SLIs/SLOs, build golden paths with Terraform, lead incidents/postmortems, influence org-wide standards, and share real on-call rotation.
About the job
How You'll Contribute
- Embed with product and platform teams from design and architecture reviews through launch, bringing SRE perspective early to ensure reliability is designed in.
- Define and evolve production-readiness standards including design reviews, launch checklists, and operational acceptance criteria.
- Collaborate to define meaningful SLIs, SLOs, and error budgets to drive prioritization decisions.
- Build frameworks, tooling, and golden paths across AWS, GCP, and Azure (using Terraform) to make the reliable path the easy path.
- Provide cross-team leadership: influence roadmaps, resolve technical disagreements, identify and address technical/process debt, mentor senior and mid-level engineers.
- Mature incident management practices including blameless postmortems, turning incidents into systematic improvements.
- Build relationships with cloud providers to influence roadmaps; represent the company in customer trust conversations and the reliability community.
- Participate in on-call rotation (currently one week per month) and respond to incidents.
Qualifications
- Multi-cloud fluency across AWS, GCP, and Azure (Terraform as IaC backbone preferred over deep specialization in one).
- Comfort supporting and contributing to TypeScript (frontend/backend) and Ruby on Rails backend services.
- Significant experience as SRE, production/platform engineer, or software engineer with deep reliability focus at scale.
- Strong software engineering fundamentals; able to write production-quality code and partner deeply with teams.
- Track record of technical leadership and influence across team boundaries without formal authority.
- Ability to drive ambiguous, high-scope problems to completion with minimal oversight.
- Systems thinking to identify and resolve organizational debt.
- Experience building measurement frameworks, analyzing operational data, and driving improvements.
- Strong verbal and written English communication skills.
Bonus Points
- Experience standing up or maturing an SRE practice at a growth-stage company.
- Background as an embedded SRE or partnering closely with product teams.
- Experience with chaos/resilience testing or progressive delivery practices.
Notes: No college degree required. Fully remote (global).
Skills
AWS, GCP, Azure, Terraform, TypeScript, Ruby on Rails, SRE, Sli, Slo, On-Call, Incident Management, Chaos Engineering
Similar jobs
DevOps / SRE jobsStaff-level site reliability engineer responsible for safely deploying and operating safeguards infrastructure across model releases and cloud platforms. The role emphasizes production change management, high-stakes incident response, and automating manual launch and validation processes.
Staff Infrastructure Engineer responsible for designing and operating scalable infrastructure for growth systems, including onboarding, referrals, and user acquisition. The role requires 7+ years of production infrastructure experience, strong reliability instincts, and independent judgment in a high-autonomy environment.
Leads the architecture, development, and operation of cloud, Kubernetes, on-premises, and hybrid infrastructure, while building developer platforms and CI/CD automation. Requires at least six years of infrastructure or related engineering experience, deep Kubernetes expertise, strong programming skills, and technical leadership.
Own the network architecture and standards for a multi-cloud enterprise AI platform deployed across Kubernetes environments and customer-controlled networks. The role requires deep cloud and Kubernetes networking expertise, strong security fundamentals, and the judgment to establish scalable, supportable connectivity patterns.
Staff Platform Engineer will build and improve automated delivery pipelines, developer environments, infrastructure, and release systems across the engineering organization. The role requires 6+ years of engineering experience, a bachelor’s degree, and expertise with CI/CD, cloud infrastructure, containers, and infrastructure as code.