Site Reliability Engineer
Owns reliability, observability, incident response, and automation for cloud-based production systems. The role requires 5+ years in SRE, DevOps, platform engineering, or related infrastructure work, with strong Kubernetes, cloud, distributed-systems, and infrastructure-automation experience.
About the job
Responsibilities
- Own the reliability, availability, and performance of production systems running in cloud environments.
- Define and monitor SLIs/SLOs and help manage error budgets across the platform.
- Lead incident response, including detection, triage, mitigation, and postmortems.
- Improve observability through logging, monitoring, alerting, and dashboards.
- Automate operational workflows and reduce manual toil.
- Partner with engineering teams to improve system resiliency and scalability.
- Assist with capacity planning, infrastructure optimization, and performance tuning.
- Build internal tooling, runbooks, and operational best practices.
- Support Kubernetes-based infrastructure and distributed systems at scale.
- Act as an escalation point for complex production and platform issues.
Requirements
- 5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or related infrastructure roles.
- Strong experience with cloud platforms such as AWS, Google Cloud, or Azure.
- Hands-on experience with Kubernetes and containerized environments.
- Strong understanding of distributed systems and microservices architecture.
- Experience with observability tools such as Prometheus, Grafana, Datadog, ELK, or OpenTelemetry.
- Proficiency with infrastructure automation and scripting, including Terraform, Python, or Bash.
- Experience managing CI/CD pipelines and deployment automation.
- Strong troubleshooting and incident management skills.
- Ability to work cross-functionally and communicate effectively during high-pressure situations.
Nice to Have
- Experience supporting large-scale SaaS or cloud-native platforms.
- Familiarity with workflow orchestration technologies such as Conductor, Temporal, or Camunda.
- Experience with Kafka, messaging systems, or event-driven architectures.
- Knowledge of security best practices and cloud infrastructure hardening.
- Open-source contributions or a strong systems engineering background.
Compensation and Benefits
- Base salary: $125,000–$250,000 USD.
- Compensation varies based on skills, experience, job scope, location and cost of living, and competitive market data for the country and role.
- Comprehensive health coverage, including medical, dental, and vision.
- Flexible PTO.
- Personal development support.
- Expected travel: 15–20%.
Skills
AWS, GCP, Azure, Kubernetes, Docker, Prometheus, Grafana, Datadog, OpenTelemetry, Terraform, Python, Bash, CI/CD, Kafka, Distributed Systems
Similar jobs
DevOps / SRE jobsOwn large-scale ClickHouse cluster upgrades and production operations while building tooling that improves release safety and automation. The role requires 5+ years operating stateful distributed systems, cloud and Kubernetes experience, strong debugging skills, and Go development experience.
Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Platform engineer responsible for forecasting and automating compute capacity across regions, including reservations, fleet reconciliation, observability, and cost optimization. Requires 5+ years in infrastructure, SRE, platform, or capacity engineering plus production software and AWS EC2 experience.
Provides hands-on L2 technical escalation support for enterprise customers in the APAC region, troubleshooting distributed systems and APIs while leading root-cause analysis, support process improvements, and technical documentation. Requires 4+ years of support or escalation engineering experience.
Build and operate distributed cloud infrastructure and platform services that support product teams globally. The role requires 5+ years of software development experience, strong distributed-systems expertise, and experience with cloud infrastructure, reliability, and observability.