Software Engineer-II SRE
Build and operate reliable, scalable production systems across AWS, Kubernetes, infrastructure automation, CI/CD, and observability. The role requires 2–4 years of SRE, DevOps, platform, or cloud infrastructure experience and strong automation skills.
About the job
Responsibilities
- Build, operate, and improve production infrastructure on AWS.
- Work with Kubernetes/EKS-based applications and platforms.
- Manage and automate infrastructure using Terraform.
- Build and improve CI/CD and GitOps workflows.
- Automate repetitive operational work using Python, Go, or scripting.
- Monitor and troubleshoot production systems using metrics, logs, traces, and alerts.
- Participate in on-call, incident response, root-cause analysis, and long-term remediation.
- Contribute to SLIs/SLOs and reliability improvements.
- Identify opportunities to improve scalability, reliability, and operational efficiency.
- Collaborate with application, platform, security, and other engineering teams.
Requirements
- 2–4 years of hands-on experience in SRE, DevOps, platform engineering, cloud infrastructure, or a related engineering role.
- Strong hands-on experience with AWS.
- Strong understanding of Linux and networking fundamentals.
- Hands-on experience with Kubernetes; EKS experience preferred.
- Experience with Terraform or infrastructure as code.
- Exposure to CI/CD and GitOps tools.
- Experience with observability tools.
- Ability to automate using Python, Go, Bash, or similar languages.
Skills
AWS, Kubernetes, Amazon Eks, Terraform, Infrastructure As Code, CI/CD, GitOps, Python, Go, Bash, Linux, Networking, Datadog, Prometheus, Grafana
Similar jobs
DevOps / SRE jobsSupports the reliability and day-to-day operation of Okta’s Customer Identity Cloud by monitoring platform health, handling service requests, executing runbooks, and troubleshooting production issues. Requires cloud operations experience, infrastructure knowledge, and familiarity with Kubernetes and monitoring tools.
Owns reliability, scalability, and performance for managed gateway services by automating cloud operations, monitoring production systems, and resolving incidents. Requires at least two years of production SRE experience plus proficiency in Golang or Python, Kubernetes, and major cloud platforms.
Supports reliable, secure, and scalable cloud platforms across AWS, GCP, and Azure, with a focus on Kubernetes workloads. The role monitors services, troubleshoots incidents, supports deployments, and automates operations while requiring 1–2 years of SRE, DevOps, cloud operations, or infrastructure experience.
Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Own reliability, scalability, and operational excellence for DataHub Cloud and enterprise deployment offerings. The role requires 5+ years in DevOps, platform engineering, or SRE, with expertise in cloud platforms, Kubernetes, infrastructure as code, observability, and deployment automation.