Senior Site Reliability Engineer
Senior Site Reliability Engineer responsible for operating and improving large-scale distributed systems, infrastructure, reliability, and incident response. Requires 5+ years of SRE or DevOps experience plus programming and container orchestration expertise.
About the job
Responsibilities
- Collaborate with internal teams to identify instability in distributed systems and drive operational excellence.
- Support core infrastructure by understanding, diagnosing, and debugging production systems.
- Provide system design consulting, develop software platforms and frameworks, and conduct launch reviews and root cause analyses.
- Maintain and document sustainable postmortem and incident response practices.
- Implement changes that improve reliability, scalability, and engineering velocity.
- Reduce toil through iterative development of tooling and automation.
- Collaborate with engineering teams to release new features and build expertise in supported services.
Requirements
- 5+ years of experience in site reliability engineering or DevOps for a product with millions of users.
- Experience identifying and resolving issues in large-scale distributed systems.
- Experience with Java, Kotlin, Python, or Go.
- Understanding of containerization and container orchestration technologies, such as Docker, Mesos, Kubernetes, or Nomad.
Nice-to-haves
- Experience improving automation and tooling to reduce service-maintenance toil.
- Experience driving improvements to incident response processes.
- Experience assessing reliability and troubleshooting Dynamo, MySQL, and/or PostgreSQL databases.
Compensation
- Base salary range: $182,800–$247,300 USD.
- Equity compensation supplements the base salary.
Skills
Site Reliability Engineering, DevOps, Distributed Systems, Java, Kotlin, Python, Go, Docker, Kubernetes, Mesos, Nomad, Dynamo, MySQL, Postgres
Similar jobs
DevOps / SRE jobsOwn the design, scaling, reliability, and automation of a multi-region storage platform supporting AI workloads. The role requires 8+ years of production infrastructure or storage engineering experience, distributed storage expertise, strong Linux and networking knowledge, and production programming skills.
Senior Site Reliability Engineer responsible for building fault-tolerant infrastructure, scaling a Nomad-based service fabric, and strengthening observability for critical brokerage systems. The role requires production experience with distributed systems, Linux, networking, instrumentation, on-call operations, and reliability practices.
Own and modernize the build, CI, test automation, and ephemeral environment platform for a large TypeScript, React, and Go monorepo. The role requires 6+ years of large-scale build-system experience, strong Bazel or comparable tooling expertise, and deep knowledge of hermetic, reproducible development workflows.
Senior platform engineer responsible for reliable, secure, and scalable infrastructure, developer tooling, observability, and AI enablement. The role requires 6+ years in platform engineering, SRE, or DevOps, with strong AWS and incident leadership experience.
Own reliability, scalability, security, observability, and incident response for mission-critical applications across Kubernetes, AWS, and on-premise DoD environments. Requires an active Top Secret clearance and at least five years of infrastructure-focused SRE, DevOps, or platform engineering experience.