Staff SRE Engineer
The Staff SRE Engineer will drive reliability, scalability, observability, and operational efficiency across highly available cloud and distributed systems. The role requires 5+ years of SRE, DevOps, or platform experience, advanced Kubernetes expertise, strong automation skills, and leadership in incident management.
About the job
Responsibilities
- Administer and maintain container orchestration platforms and containerized workloads.
- Monitor and troubleshoot production systems, participating in on-call rotations to ensure reliability.
- Drive observability improvements by enhancing monitoring, logging, and alerting capabilities across systems and data platforms.
- Administer and optimize cloud-based environments across multiple providers.
- Manage and support distributed data platforms and real-time processing systems.
- Develop and maintain continuous integration and delivery pipelines for efficient and reliable deployments.
- Own and implement Infrastructure as Code (IaC) practices to ensure consistency and scalability.
- Automate and orchestrate infrastructure using programming and scripting languages.
- Perform system administration and networking tasks to support internal and external environments.
- Collaborate effectively with engineers and stakeholders across different time zones.
Requirements
- 5+ years of experience in Site Reliability Engineering, DevOps, or Platform Engineering roles.
- Proven success leading large-scale production systems in cloud environments, including AWS, GCP, Azure, or OCI.
- Demonstrated leadership in driving incident response, on-call best practices, and a reliability-focused culture.
- Strong experience with production on-call operations and incident management.
- Advanced proficiency in Kubernetes administration and troubleshooting.
- Hands-on experience with Prometheus, Grafana, Loki, and Alertmanager.
- Knowledge of chat-based operations interfaces and/or auto-remediation controllers using AI agentic frameworks.
- Understanding of AI agents for auto-triaging alerts, correlating signals, and suggesting root-cause hypotheses.
- Expertise in operating data platforms including Elasticsearch, MongoDB, Spark, Kafka, and Redis.
- Proficiency with public cloud services such as AWS, Azure, GCP, or OCI.
- Strong programming and automation skills in Python and Bash.
- Deep understanding of Infrastructure as Code, including Terraform and Helm.
- Experience with CI/CD pipelines and tools such as GitHub Actions, Bitbucket, and ArgoCD.
- Strong technical background in distributed systems, databases, networking, and Linux administration.
- Excellent problem-solving, communication, and leadership abilities.
- Bachelor's degree in Computer Science, Engineering, or a related technical field.
Nice-to-haves
- Certifications in AWS, GCP, observability, Linux, or Kubernetes.
Skills
Kubernetes, AWS, GCP, Azure, Oci, Prometheus, Grafana, Loki, Alertmanager, Elasticsearch, MongoDB, Spark, Kafka, Redis, Python
Similar jobs
DevOps / SRE jobsLeads the technical direction of multi-cloud Kubernetes capacity management and workload placement across Datadog’s large-scale infrastructure. The role requires strong systems programming experience, ideally in Go, cloud infrastructure expertise, and the ability to influence architecture across teams.
Build and operate high-performance customer compute environments spanning bare metal, Kubernetes, Slurm, GPUs, networking, storage, and observability. The role requires 5+ years of production Linux infrastructure experience and strong expertise in bare-metal Kubernetes, NVIDIA GPUs, virtualization, and networking.
Owns and evolves CI/CD, mobile release, testing, and deployment infrastructure for a production fintech application. The role requires 8+ years in DevOps or related platform disciplines, strong AWS and Kubernetes expertise, and experience with secure mobile release systems.
Leads production reliability for Grafana Cloud’s multi-tenant database products, partnering with product engineering teams to improve SLOs, scalability, observability, automation, and incident response. Requires 8+ years of engineering experience, including substantial SRE or production engineering work, plus strong Kubernetes and cloud expertise.
The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.