Site Reliability Engineer 2, Managed Gateways
Owns reliability, scalability, and performance for managed gateway services by automating cloud operations, monitoring production systems, and resolving incidents. Requires at least two years of production SRE experience plus proficiency in Golang or Python, Kubernetes, and major cloud platforms.
About the job
Responsibilities
- Implement and maintain automation for deploying and operating managed gateways across cloud environments.
- Monitor system health, performance, and uptime, targeting 99.99% availability for core infrastructure.
- Resolve complex production incidents and participate in on-call rotations.
- Build resilient tools and systems that improve platform reliability and operational efficiency.
- Prevent technical debt and support sustainable, scalable operations.
- Collaborate with engineering teams to design, review, and implement resilient, highly scalable services.
Requirements
- 2+ years of Site Reliability Engineering experience in a production environment.
- Proficiency in at least one of Golang or Python for automation, tooling, and infrastructure as code.
- Hands-on experience with Kubernetes and major cloud platforms such as AWS, Google Cloud, or Azure.
- Familiarity with monitoring, logging, and alerting tools such as Prometheus, Grafana, or Datadog.
- Understanding of networking concepts, distributed systems, and API gateways.
Nice-to-haves
- Experience with Kong Gateway or other API management platforms.
- Relevant cloud certifications, such as AWS Certified DevOps Engineer or Kubernetes Administrator.
- Contributions to open-source projects or developer communities.
Skills
Go, Python, Kubernetes, AWS, GCP, Microsoft Azure, Prometheus, Grafana, Datadog, Infrastructure As Code, Distributed Systems, Api Gateways
Similar jobs
DevOps / SRE jobsBuild and operate reliable, scalable production systems across AWS, Kubernetes, infrastructure automation, CI/CD, and observability. The role requires 2–4 years of SRE, DevOps, platform, or cloud infrastructure experience and strong automation skills.
Supports the reliability and day-to-day operation of Okta’s Customer Identity Cloud by monitoring platform health, handling service requests, executing runbooks, and troubleshooting production issues. Requires cloud operations experience, infrastructure knowledge, and familiarity with Kubernetes and monitoring tools.
Supports reliable, secure, and scalable cloud platforms across AWS, GCP, and Azure, with a focus on Kubernetes workloads. The role monitors services, troubleshoots incidents, supports deployments, and automates operations while requiring 1–2 years of SRE, DevOps, cloud operations, or infrastructure experience.
Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Own reliability, scalability, and operational excellence for DataHub Cloud and enterprise deployment offerings. The role requires 5+ years in DevOps, platform engineering, or SRE, with expertise in cloud platforms, Kubernetes, infrastructure as code, observability, and deployment automation.