Senior Software Engineer, Site Reliability
Senior SRE measures software performance, defines SLOs/SLAs, optimizes infrastructure with Temporal/Kubernetes/AWS, handles on-call, and improves developer experience/scalability for growing B2B workflows. Requires 5+ years SRE/DevOps experience.
About the job
Responsibilities
- Monitor core business-logic software via on-call and non-urgent situations, define SLOs/SLAs.
- Extend and monitor infrastructure stack for business-critical B2B usage.
- Maintain mental and reified models of systems for risk estimation, project planning, and debugging.
- Collaborate with engineering team and leadership on site reliability expertise.
- Work on core orchestration logic using Temporal to run thousands of workflows.
- Advise on major backend projects from planning to release.
- Optimize services for scalability, stability, and observability.
- Improve developer experience, engineering processes, especially with LLM coding agents.
Requirements
- 5+ years of SRE, DevOps, or Platform engineering experience.
- Proven record building efficient, performant, extensible systems.
- Experience with quantitative metrics, SLOs/SLAs, navigating tradeoffs.
- Familiarity with containerization/orchestration (Docker, Kubernetes), Linux, backend services.
- Experience implementing/managing AWS infrastructure.
- Strong communicator, team player.
Nice to Haves
- Experience with Temporal.
- Building AI platforms/tooling.
- Kubernetes/Helm (Amazon EKS).
- IaaS tools for production cloud.
- Managing CI/CD pipelines, developer experience.
- Cloud storage/data modeling products.
- Early stage startups.
Compensation
Salary Range: $180,000 - $200,000 (SF/NY base, plus equity, health benefits).
Skills
Temporal, Kubernetes, Docker, AWS, Amazon Eks, Linux, CI/CD, Helm, SLOs, Slas
Similar jobs
DevOps / SRE jobsOwn the design, scaling, reliability, and automation of a multi-region storage platform supporting AI workloads. The role requires 8+ years of production infrastructure or storage engineering experience, distributed storage expertise, strong Linux and networking knowledge, and production programming skills.
Senior Site Reliability Engineer responsible for building fault-tolerant infrastructure, scaling a Nomad-based service fabric, and strengthening observability for critical brokerage systems. The role requires production experience with distributed systems, Linux, networking, instrumentation, on-call operations, and reliability practices.
Own and modernize the build, CI, test automation, and ephemeral environment platform for a large TypeScript, React, and Go monorepo. The role requires 6+ years of large-scale build-system experience, strong Bazel or comparable tooling expertise, and deep knowledge of hermetic, reproducible development workflows.
Senior platform engineer responsible for reliable, secure, and scalable infrastructure, developer tooling, observability, and AI enablement. The role requires 6+ years in platform engineering, SRE, or DevOps, with strong AWS and incident leadership experience.
Own reliability, scalability, security, observability, and incident response for mission-critical applications across Kubernetes, AWS, and on-premise DoD environments. Requires an active Top Secret clearance and at least five years of infrastructure-focused SRE, DevOps, or platform engineering experience.