Skip to content

Staff SRE Engineer

The Staff SRE Engineer will drive reliability, scalability, observability, and operational efficiency across highly available cloud and distributed systems. The role requires 5+ years of SRE, DevOps, or platform experience, advanced Kubernetes expertise, strong automation skills, and leadership in incident management.

About the job

Responsibilities

  • Administer and maintain container orchestration platforms and containerized workloads.
  • Monitor and troubleshoot production systems, participating in on-call rotations to ensure reliability.
  • Drive observability improvements by enhancing monitoring, logging, and alerting capabilities across systems and data platforms.
  • Administer and optimize cloud-based environments across multiple providers.
  • Manage and support distributed data platforms and real-time processing systems.
  • Develop and maintain continuous integration and delivery pipelines for efficient and reliable deployments.
  • Own and implement Infrastructure as Code (IaC) practices to ensure consistency and scalability.
  • Automate and orchestrate infrastructure using programming and scripting languages.
  • Perform system administration and networking tasks to support internal and external environments.
  • Collaborate effectively with engineers and stakeholders across different time zones.

Requirements

  • 5+ years of experience in Site Reliability Engineering, DevOps, or Platform Engineering roles.
  • Proven success leading large-scale production systems in cloud environments, including AWS, GCP, Azure, or OCI.
  • Demonstrated leadership in driving incident response, on-call best practices, and a reliability-focused culture.
  • Strong experience with production on-call operations and incident management.
  • Advanced proficiency in Kubernetes administration and troubleshooting.
  • Hands-on experience with Prometheus, Grafana, Loki, and Alertmanager.
  • Knowledge of chat-based operations interfaces and/or auto-remediation controllers using AI agentic frameworks.
  • Understanding of AI agents for auto-triaging alerts, correlating signals, and suggesting root-cause hypotheses.
  • Expertise in operating data platforms including Elasticsearch, MongoDB, Spark, Kafka, and Redis.
  • Proficiency with public cloud services such as AWS, Azure, GCP, or OCI.
  • Strong programming and automation skills in Python and Bash.
  • Deep understanding of Infrastructure as Code, including Terraform and Helm.
  • Experience with CI/CD pipelines and tools such as GitHub Actions, Bitbucket, and ArgoCD.
  • Strong technical background in distributed systems, databases, networking, and Linux administration.
  • Excellent problem-solving, communication, and leadership abilities.
  • Bachelor's degree in Computer Science, Engineering, or a related technical field.

Nice-to-haves

  • Certifications in AWS, GCP, observability, Linux, or Kubernetes.

Skills

Kubernetes, AWS, GCP, Azure, Oci, Prometheus, Grafana, Loki, Alertmanager, Elasticsearch, MongoDB, Spark, Kafka, Redis, Python

Datadog

Datadog

Dublin, Ireland
Staff Engineer, Compute
No salary listedHybrid7+ YOEDevOps / SRE

Leads the technical direction of multi-cloud Kubernetes capacity management and workload placement across Datadog’s large-scale infrastructure. The role requires strong systems programming experience, ideally in Go, cloud infrastructure expertise, and the ability to influence architecture across teams.

Fal

Fal

Remote

Senior/Staff Kubernetes Infrastructure Engineer
$180k+/yrRemote5+ YOEDevOps / SRE

Build and operate high-performance customer compute environments spanning bare metal, Kubernetes, Slurm, GPUs, networking, storage, and observability. The role requires 5+ years of production Linux infrastructure experience and strong expertise in bare-metal Kubernetes, NVIDIA GPUs, virtualization, and networking.

Phantom

Phantom

Remote

Staff DevOps Engineer
No salary listedRemote8+ YOEDevOps / SRE

Owns and evolves CI/CD, mobile release, testing, and deployment infrastructure for a production fintech application. The role requires 8+ years in DevOps or related platform disciplines, strong AWS and Kubernetes expertise, and experience with secure mobile release systems.

Grafana Labs

Grafana Labs

United Kingdom
Staff Software Engineer - Databases SRE
£104k+/yrRemote8+ YOEDevOps / SRE

Leads production reliability for Grafana Cloud’s multi-tenant database products, partnering with product engineering teams to improve SLOs, scalability, observability, automation, and incident response. Requires 8+ years of engineering experience, including substantial SRE or production engineering work, plus strong Kubernetes and cloud expertise.

Lightning AI

Lightning AI

Remote

Senior Network Engineer
$150k+/yrRemote5+ YOEDevOps / SRE

The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.