Skip to content
FlexAIFlexAI

Staff DevOps Engineer/SRE

The Staff DevOps/SRE Engineer will define infrastructure strategy and SRE practices while building reliable, scalable systems for distributed, multi-cloud AI workloads. The role requires 8+ years of experience, deep Kubernetes and infrastructure-as-code expertise, and strong capabilities in automation, observability, and incident response.

About the job

Responsibilities

Reliability and Architecture

  • Design and evolve the infrastructure backbone for an AI and PaaS platform.
  • Build highly available, fault-tolerant, and scalable systems.
  • Define and drive SRE practices, including SLIs, SLOs, and error budgets.

Infrastructure at Scale

  • Lead infrastructure as code using Pulumi.
  • Own and scale Kubernetes clusters and containerized workloads.
  • Standardize and automate infrastructure for global deployments.

CI/CD and Automation

  • Design and scale CI/CD pipelines for fast, reliable releases.
  • Build self-healing systems and automated remediation workflows.
  • Drive GitOps and platform engineering practices.

Observability and Performance

  • Implement end-to-end observability using VictoriaMetrics and Grafana for metrics, logs, and traces.
  • Identify and resolve performance bottlenecks involving latency, throughput, and cost.
  • Lead incident response, root-cause analysis, and postmortems.

Leadership and Collaboration

  • Partner with backend, AI, runtime, and security teams.
  • Guide infrastructure decisions and scaling strategy.
  • Mentor engineers and raise reliability and engineering standards.

Security and Resilience

  • Embed security into infrastructure and deployment workflows.
  • Design for resilience through disaster recovery, chaos testing, and capacity planning.

Requirements

  • 8+ years of experience in DevOps, SRE, or infrastructure engineering.
  • Experience operating large-scale, distributed systems in production.
  • Deep expertise in Kubernetes and container orchestration.
  • Deep expertise in Pulumi or similar infrastructure-as-code tools.
  • Experience with cloud or hybrid environments, including AWS, GCP, Azure, or on-premises infrastructure.
  • Experience with observability stacks such as Prometheus, Grafana, and OpenTelemetry.
  • Strong experience with CI/CD, automation, and release engineering.
  • Proficiency in Python, Go, or Bash.
  • Strong systems thinking and debugging skills in high-scale environments.
  • Experience defining and operating with SLOs and SLAs.
  • Experience in startup environments.
  • Comfortable leveraging AI coding tools and agents.

Nice to Have

  • Experience with AI/ML infrastructure or GPU workloads.
  • Familiarity with distributed or high-performance compute systems.
  • Exposure to platform engineering or internal developer platforms.
  • Experience scaling systems from beta to production.

Benefits and Compensation

  • Work on cutting-edge AI infrastructure.
  • Build systems that power developers and enterprises.
  • High ownership, fast execution, and real impact.
  • Collaborative, high-caliber team.

Skills

Kubernetes, Pulumi, AWS, GCP, Microsoft Azure, Prometheus, Grafana, OpenTelemetry, Victoriametrics, CI/CD, GitOps, Python, Go, Bash, Docker

Together AI

Together AI

London, United Kingdom
Staff Software Engineer, Inference / Compute Infrastructure Engineering
No salary listedRemote7+ YOEDevOps / SRE

Build and operate a Kubernetes-native control plane for provisioning, scheduling, self-healing, and optimizing GPU inference infrastructure. The role requires strong software engineering, durable workflow orchestration, reconciliation systems, event-driven architecture, and platform API experience.

Okta

Okta

Bengaluru, India

Staff DevSecOps Engineer, Enterprise Technology
No salary listedOn-site8+ YOEDevOps / SRE

Owns enterprise DevSecOps architecture across Salesforce, NetSuite, Workday, AEM, and modern web platforms. The role requires 8+ years of DevSecOps, SRE, or security engineering experience, strong CI/CD and edge-security expertise, and leadership in secure automation, observability, identity, and compliance.

Together AI

Together AI

London, United Kingdom
Staff Software Engineer, Inference / Compute Infrastructure Engineering
No salary listedOn-site7+ YOEDevOps / SRE

Build and operate declarative control planes, durable workflows, and self-healing systems that provision and manage GPU inference infrastructure. The role requires strong software engineering, reconciliation or orchestration experience, and event-driven systems expertise.

Okta

Okta

Bengaluru, India

Staff Software Engineer
No salary listedHybrid7+ YOEDevOps / SRE

Builds and mentors development of scalable cloud tooling, Continuous Delivery platforms, Infrastructure as Code automation, and supporting microservices across AWS environments. The role requires substantial backend software development experience with Java, Go, or Python, plus Terraform, CI/CD, containers, and distributed systems expertise.

Fal

Fal

Remote

Senior/Staff Kubernetes Infrastructure Engineer
$180k+/yrRemote5+ YOEDevOps / SRE

Build and operate high-performance customer compute environments spanning bare metal, Kubernetes, Slurm, GPUs, networking, storage, and observability. The role requires 5+ years of production Linux infrastructure experience and strong expertise in bare-metal Kubernetes, NVIDIA GPUs, virtualization, and networking.