Site Reliability Engineer II
Site Reliability Engineer II responsible for designing, deploying, and maintaining multi-cloud infrastructure (Azure primary, AWS/GCP) for Illumio's SaaS products. Focus on IaC, CI/CD pipelines, monitoring, incident response, automation, and improving reliability/scalability in collaboration with engineering and security teams. Requires 2+ years SRE/DevOps experience with Azure.
About the job
Your Impact
As an SRE Engineer II, you will be responsible for managing our multi-cloud infrastructure on Azure, AWS and/or GCP. As and when required, you will be responsible for designing new services and applications in the cloud(s) and take them from development to production while working closely with Engineering, SRE/OPS, and Security teams.
On a day-to-day basis, you will work on enhancing system reliability and scalability of Illumio SaaS products, and drive continuous improvement initiatives.
- Design, deploy, and maintain cloud infrastructure solutions on Azure, AWS, and/or GCP to support our applications and services
- Implement infrastructure as code (IaC) principles using tools such as Terraform, ARM templates, or CloudFormation to automate provisioning and configuration management
- Develop and maintain CI/CD pipelines for automated software delivery and deployment, leveraging tools such as Azure DevOps, AWS CodePipeline, or Jenkins
- Monitor system performance, application health, and infrastructure metrics using cloud monitoring and logging services, and implement proactive measures to optimize performance and availability
- Support incident response and resolution efforts, conduct root cause analysis, implement corrective actions, and document post-incident reviews
- Collaborate with Engineering teams to design and implement scalable and reliable architectures, providing guidance on best practices for cloud-native application development
- Implement security best practices and controls in cloud environments to protect data, applications, and infrastructure, and ensure compliance with regulatory requirements
- Drive automation initiatives to streamline operational tasks, reduce manual effort, and improve overall efficiency in cloud operations
- Stay current with cloud platform updates, trends, and best practices, and evaluate emerging technologies for potential adoption to drive innovation and efficiency
- Provide support and guidance to junior team members, fostering a culture of learning, collaboration, and continuous improvement within the SRE/DevOps team
Your Toolkit
- Bachelor's degree in Computer Science, Engineering, or related field; or equivalent work experience
- 2+ years of experience working as an SRE, DevOps Engineer, or similar role, with hands-on experience in Azure cloud platform in a production environment setting
- Exposure to AWS and/or GCP cloud platforms is preferred
- Proficiency in scripting and programming languages such as PowerShell, Python, or Go for automation and infrastructure management tasks
- Experience with CI/CD tools and methodologies, containerization technologies, and microservices architecture in cloud environments
- Strong analytical, problem-solving, and communication skills, with the ability to collaborate effectively with cross-functional teams
- Azure certifications such as Azure Administrator, Azure Developer, or AWS/GCP certifications are a plus
Skills
Azure, AWS, GCP, Terraform, CloudFormation, Arm Templates, CI/CD, Jenkins, Python, PowerShell, Go, Kubernetes, Docker, Microservices
Similar jobs
DevOps / SRE jobsBackend engineers build and scale Airtable's infrastructure across teams like Base, Compute, Data, Storage, and Traffic. Requires 2-8 years experience in distributed systems, databases; CS degree; hybrid work in SF, NYC, Seattle, or LA areas.
Production Engineer builds and operates large-scale systems, focusing on automation, monitoring, infrastructure management, and resilient operations. Requires 2+ years in SRE/DevOps, expertise in Linux, AWS, Kubernetes, and programming in Python or Golang.
Build Mercury’s secure, observable infrastructure platform across AWS, networking, containers, and developer tooling. The role requires strong Linux fundamentals, cloud-native experience, technical writing ability, and software development skills, with opportunities to support AI-agent infrastructure.
Infrastructure and site reliability intern building and operating on-premises backend infrastructure for a semiconductor fabrication environment. The role emphasizes systems programming, Linux, networking, reliability, observability, automation, and performance engineering.
Winter infrastructure and site reliability internship focused on building and operating minimal, on-premises backend infrastructure for a semiconductor fabrication facility. The role requires systems programming, Linux, networking, distributed systems, and hands-on infrastructure or automation experience.