Senior Site Reliability Engineer
Senior SRE improves platform reliability using AI-driven automation, leads incident response and oncall for critical services, and implements safe deployment practices. Requires 5+ years experience in backend/platform engineering, observability, and high-availability systems.
About the job
Responsibilities
- Build and extend platforms to improve system reliability
- Work on team goals that encompass reliability for the entire company
- Standardize reliability tools across multiple platforms and organizations
- Triage, coordinate, and lead stabilization of sev 0–1 incidents
- Serve as primary oncall, maintaining structured escalation paths and exercising leadership escalation
- Drive platform-wide reliability improvements, shared operational tooling, and deploy-safety patterns
- Use AI-driven systems to improve signal detection, reduce noise, and accelerate root cause analysis
- Design and implement safe deployment patterns (progressive delivery, automated rollback, guardrails)
Requirements
- Drive to root cause systems with many moving parts and take the necessary steps to fix them
- Demonstrated technical initiative and leadership on previous projects, especially those with a backend/platform focus
- Familiarity with AI-driven tooling for observability, incident analysis, or automation
- A mindset that naturally reaches for AI to accelerate problem-solving and reduce toil
- Experience running production oncall for high-availability systems
- Strong incident management skills — structured triage, mitigation under pressure, blameless postmortems
- Fluency with CI/CD pipelines, progressive rollout strategies, and rollback automation
- Monitoring & observability expertise — building/tuning alerts for uptime, error rates, latency regression, and resource exhaustion
- Ability to create and maintain evidence-based maturity assessments using trailing 90-day data windows
- Comfort with vendor/dependency management — maintaining validated escalation contacts reachable within ≤ 5 minutes
- Boundless curiosity, autonomy, and a strong sense of accountability
- A strong desire to perform and grow as an engineer
- 5+ years of software development experience
Technologies
- Kotlin, Modern Java (11+)
- HTTP, JSON, gRPC, and Protocol Buffers
- MySQL / Vitess / DynamoDB
- Event driven architectures
- DataDog
- LaunchDarkly
- Terraform, Kubernetes, Istio/Envoy
- Amazon Web Services
Compensation & Benefits
- Healthcare coverage (Medical, Vision and Dental insurance)
- Health Savings Account and Flexible Spending Account
- Retirement Plans including company match
- Employee Stock Purchase Program
- Wellness programs, including access to mental health, 1:1 financial planners, and a monthly wellness allowance
- Paid parental and caregiving leave
- Paid time off (including 12 paid holidays)
- Paid sick leave
- Learning and Development resources
- Paid Life insurance, AD&D, and disability benefits
- Equity plan eligibility
- Possible sign-on bonus
Pay varies by location zone:
- Zone A: USD $189,000 - $283,600
- Zone B: USD $179,600 - $269,400
- Zone C: USD $170,100 - $255,100
- Zone D: USD $160,700 - $241,100
Skills
Kubernetes, Terraform, AWS, Datadog, Istio, Launchdarkly, Kotlin, Java, gRPC, MySQL, Vitess, DynamoDB, CI/CD, Observability, AI Tools
Similar jobs
DevOps / SRE jobsSenior software engineer responsible for operating and evolving Voltus’s infrastructure platform across AWS, Kubernetes, Nomad, observability, stateful systems, and developer tooling. The role requires 6+ years of engineering experience, deep production Kubernetes and AWS expertise, and strong Go or Python skills.
Own and evolve VSCO’s AWS/EKS platform, including infrastructure as code, GitOps, CI/CD, observability, networking, and production reliability. The role requires 5+ years of hands-on infrastructure or SRE experience and strong Kubernetes, Terraform, and AWS expertise.
Build and operate highly available, distributed platform services and cloud infrastructure for petabyte-scale observability products. The role requires 6+ years of experience, strong Java and AWS expertise, Kubernetes and Terraform production experience, and a bachelor’s degree or equivalent.
Own foundational cloud infrastructure and the internal developer platform supporting Commure’s engineering teams. The role requires 6+ years of infrastructure, platform, or SRE experience and hands-on expertise across Kubernetes, infrastructure as code, GitOps, observability, and cloud environments.
Leads cloud infrastructure, platform strategy, deployment pipelines, and infrastructure automation for a growing consumer platform. Requires 5+ years in infrastructure, DevOps, platform engineering, or SRE, plus deep AWS, coding, containerization, and infrastructure-as-code experience.