Site Reliability Engineer (Senior or Staff), Atlas
Senior or Staff Site Reliability Engineer maintains and scales the Atlas platform in a multi-cloud environment, focusing on automation, on-call reliability, and collaboration with engineering teams. Requires 5+ years experience with Linux, cloud providers, and programming languages like Go, Python, or Ruby.
About the job
Responsibilities
- Participate in the development of a reliable and resilient multi-cloud platform that hosts business critical applications for a wide & varied range of customer applications
- Collaborate with service-owning teams to provide internal support, solve technical challenges and adapt or build tooling to solve novel use cases in a generic fashion
- Participate in a 24/7 on-call rotation to swiftly resolve issues related to any disruption of our customer facing Atlas fleet, ensuring minimal disruption and high availability
Requirements
- 5+ years of experience running critical systems at scale
- Value efficiency in processes and operations, and display a preference for automation over manual processes
- Familiar with a major cloud provider (AWS, Azure, or GCP) and possess the ability to build and operate systems in a multi-cloud environment
- Strong understanding of how to run a large scale Linux environment, including low level fundamentals
- Firm grasp of at least one modern programming language, beyond basic scripting (Go, Ruby, Python)
- Solid understanding of web and network protocols and standards (HTTP, TLS, DNS, etc)
Skills
Linux, AWS, Azure, GCP, Go, Ruby, Python, Http, Tls, DNS, Kubernetes
Similar jobs
DevOps / SRE jobsOwn and scale Nango’s cloud platform, customer-controlled deployments, infrastructure automation, reliability, and data layer. The role requires 10+ years in platform, infrastructure, DevOps, or SRE work, with deep Kubernetes, AWS, Terraform, database, and compliance experience.
Own and scale the company’s cloud platform, BYOC deployments, infrastructure automation, reliability, data layer, and infrastructure security. Requires 10+ years in platform, infrastructure, DevOps, or SRE roles, with deep Kubernetes, AWS, Terraform, and database expertise.
Leads the architecture, automation, observability, and reliability of multi-region AWS infrastructure supporting mission-critical payment systems. Requires 8+ years of distributed-systems experience and deep expertise in infrastructure as code, Kubernetes, automation, and cloud networking.
Leads the design and deployment of AI-enabled manufacturing systems, MES, connected-factory infrastructure, and automation for aircraft production. Requires a bachelor’s degree and 8+ years of experience in digital manufacturing, industrial automation, or software-enabled operations.
Designs, automates, and operates AWS infrastructure, shared development environments, and container platforms. The role requires strong experience with Kubernetes, infrastructure as code, environment lifecycle automation, cloud security, compliance, and cost optimization.