Skip to content
OktaOkta

Staff Site Reliability Engineer

Leads the design, operation, and modernization of multi-cloud infrastructure across AWS and Google Cloud. The role requires deep Kubernetes expertise, SRE practices, infrastructure as code, automation, observability, and at least eight years of relevant experience.

About the job

Responsibilities

  • Design, build, and operate highly scalable, reliable, and secure infrastructure powering production systems across AWS and Google Cloud.
  • Lead reliability and modernization initiatives, including container platform migrations such as ECS to EKS/GKE and microservice enablement across multi-cloud environments.
  • Serve as a technical authority in Kubernetes, cloud infrastructure, and modern CI/CD practices including GitOps and automation pipelines.
  • Partner with development teams to architect and enable microservice-based applications, ensuring production readiness, scalability, and observability.
  • Implement and manage infrastructure as code with Terraform and Ansible to automate provisioning, scaling, and configuration management across cloud providers.
  • Improve observability, performance, and cost efficiency through monitoring, logging, and alerting across AWS and Google Cloud.
  • Define SLOs and SLIs, conduct blameless postmortems, and continuously improve incident response.
  • Lead complex technical projects from conception to completion, managing timelines and technical dependencies across teams.
  • Mentor engineers and foster a culture of reliability, automation, and continuous learning.
  • Collaborate with security and compliance partners on infrastructure standards, including IAM Federation and Workload Identity.
  • Participate in the on-call rotation and use incidents to improve systems and processes.

Requirements

  • 8+ years in SRE, DevOps, or infrastructure engineering roles.
  • 3–5 years of production experience with Kubernetes, including EKS and GKE, and ecosystem tools such as Helm and Karpenter.
  • 3–5 years of experience with AWS and Google Cloud.
  • 3–5 years using Terraform to manage multi-cloud infrastructure.
  • 5+ years of coding experience in Python, Go, or similar languages.
  • Hands-on experience architecting and operating cloud-native distributed systems.
  • Experience leading ECS-to-EKS/GKE migration projects and enabling microservice architectures.
  • Proficiency with Terraform, Ansible, or CloudFormation.
  • Advanced understanding of CI/CD pipelines, Linux systems, networking fundamentals, and Redis.
  • Experience managing databases and caching systems such as RDS, Cloud SQL, Redis/Memorystore, PostgreSQL, and MySQL.
  • Experience with observability tools including Prometheus, Grafana, ELK, Loki, OpenTelemetry, and Google Cloud Operations.
  • Working knowledge of container security, secrets management, and production compliance.
  • Strong communication and problem-solving skills, including cross-team project leadership and mentoring.
  • Strong Linux and security fundamentals.
  • Bachelor’s degree in Computer Science or equivalent hands-on experience.

Nice to Have

  • Experience in SaaS or high-scale, cloud-native environments.

Benefits

  • In-person onboarding experience.
  • Well-being support.
  • Social impact opportunities.
  • Talent development and community-building programs.

Skills

AWS, GCP, Kubernetes, Amazon Eks, Google Gke, Terraform, Ansible, Python, Go, CI/CD, Argo Cd, Linux, Redis, Prometheus, Grafana

Together AI

Together AI

London, United Kingdom
Staff Software Engineer, Inference / Compute Infrastructure Engineering
No salary listedRemote7+ YOEDevOps / SRE

Build and operate a Kubernetes-native control plane for provisioning, scheduling, self-healing, and optimizing GPU inference infrastructure. The role requires strong software engineering, durable workflow orchestration, reconciliation systems, event-driven architecture, and platform API experience.

Okta

Okta

Bengaluru, India

Staff DevSecOps Engineer, Enterprise Technology
No salary listedOn-site8+ YOEDevOps / SRE

Owns enterprise DevSecOps architecture across Salesforce, NetSuite, Workday, AEM, and modern web platforms. The role requires 8+ years of DevSecOps, SRE, or security engineering experience, strong CI/CD and edge-security expertise, and leadership in secure automation, observability, identity, and compliance.

Together AI

Together AI

London, United Kingdom
Staff Software Engineer, Inference / Compute Infrastructure Engineering
No salary listedOn-site7+ YOEDevOps / SRE

Build and operate declarative control planes, durable workflows, and self-healing systems that provision and manage GPU inference infrastructure. The role requires strong software engineering, reconciliation or orchestration experience, and event-driven systems expertise.

Okta

Okta

Bengaluru, India

Staff Software Engineer
No salary listedHybrid7+ YOEDevOps / SRE

Builds and mentors development of scalable cloud tooling, Continuous Delivery platforms, Infrastructure as Code automation, and supporting microservices across AWS environments. The role requires substantial backend software development experience with Java, Go, or Python, plus Terraform, CI/CD, containers, and distributed systems expertise.

Fal

Fal

Remote

Senior/Staff Kubernetes Infrastructure Engineer
$180k+/yrRemote5+ YOEDevOps / SRE

Build and operate high-performance customer compute environments spanning bare metal, Kubernetes, Slurm, GPUs, networking, storage, and observability. The role requires 5+ years of production Linux infrastructure experience and strong expertise in bare-metal Kubernetes, NVIDIA GPUs, virtualization, and networking.