Skip to content
OktaOkta

Staff Site Reliability Engineer - Ecosystem

Leads the design, operation, and modernization of multi-cloud infrastructure across AWS and Google Cloud, with deep Kubernetes, automation, observability, and SRE expertise. The role requires 8+ years in SRE, DevOps, or infrastructure engineering and experience leading large-scale migrations and cross-team reliability initiatives.

About the job

Responsibilities

  • Design, build, and operate highly scalable, reliable, and secure infrastructure powering production systems across AWS and Google Cloud.
  • Lead reliability and modernization initiatives, including container platform migrations such as ECS to EKS/GKE and microservice enablement across multi-cloud environments.
  • Serve as a technical authority in Kubernetes, cloud infrastructure, and modern CI/CD practices including GitOps and automation pipelines.
  • Partner with development teams to architect and enable microservice-based applications, ensuring production readiness, scalability, and observability.
  • Implement and manage infrastructure as code with Terraform and Ansible to automate provisioning, scaling, and configuration management across cloud providers.
  • Improve observability, performance, and cost efficiency through monitoring, logging, and alerting systems spanning AWS and Google Cloud.
  • Champion SRE practices by defining SLOs and SLIs, conducting blameless postmortems, and improving incident response.
  • Lead complex technical projects from conception to completion, managing timelines and dependencies across teams.
  • Mentor engineers and foster a culture of reliability, automation, and continuous learning.
  • Collaborate with security and compliance partners to ensure infrastructure follows applicable standards, including IAM Federation and Workload Identity.
  • Participate in the on-call rotation and use incidents to improve systems and processes.

Requirements

  • 8+ years of experience in SRE, DevOps, or infrastructure engineering roles.
  • 3–5 years of production experience with Kubernetes, including EKS, GKE, and related tools such as Helm and Karpenter.
  • 3–5 years of experience with AWS and Google Cloud.
  • 3–5 years of experience using Terraform to manage multi-cloud infrastructure.
  • 3+ years of coding experience in Python, Go, or similar languages.
  • Experience leading high-impact migration projects, particularly ECS-to-EKS/GKE migrations, and enabling microservice architectures.
  • Experience implementing SLOs/SLIs, performing root-cause analyses, and improving operational resilience.
  • Strong Linux and security fundamentals.
  • Bachelor’s degree in Computer Science or equivalent hands-on experience.
  • Strong communication and problem-solving skills, including experience leading cross-team projects and mentoring peers.

Technical Expertise

  • Cloud-native distributed systems across AWS and Google Cloud.
  • Infrastructure as code using Terraform, Ansible, or CloudFormation.
  • CI/CD with Argo CD, GitLab CI, or Spinnaker.
  • Linux systems and networking fundamentals, including Direct Connect/Interconnect, DNS, routing, and load balancing.
  • Databases and caching systems such as RDS/Cloud SQL, Redis/Memorystore, PostgreSQL, and MySQL.
  • Observability tools including Prometheus, Grafana, ELK, Loki, OpenTelemetry, and Google Cloud Operations.
  • Container security and secrets management using HashiCorp Vault, AWS Secrets Manager, or Google Secret Manager.

Nice to Have

  • Experience in SaaS or high-scale, cloud-native environments.

Skills

AWS, GCP, Kubernetes, Amazon Eks, Google Gke, Terraform, Ansible, Python, Go, Argo Cd, Gitlab Ci, Prometheus, Grafana, Linux, Hashicorp Vault

Together AI

Together AI

London, United Kingdom
Staff Software Engineer, Inference / Compute Infrastructure Engineering
No salary listedRemote7+ YOEDevOps / SRE

Build and operate a Kubernetes-native control plane for provisioning, scheduling, self-healing, and optimizing GPU inference infrastructure. The role requires strong software engineering, durable workflow orchestration, reconciliation systems, event-driven architecture, and platform API experience.

Okta

Okta

Bengaluru, India

Staff DevSecOps Engineer, Enterprise Technology
No salary listedOn-site8+ YOEDevOps / SRE

Owns enterprise DevSecOps architecture across Salesforce, NetSuite, Workday, AEM, and modern web platforms. The role requires 8+ years of DevSecOps, SRE, or security engineering experience, strong CI/CD and edge-security expertise, and leadership in secure automation, observability, identity, and compliance.

Together AI

Together AI

London, United Kingdom
Staff Software Engineer, Inference / Compute Infrastructure Engineering
No salary listedOn-site7+ YOEDevOps / SRE

Build and operate declarative control planes, durable workflows, and self-healing systems that provision and manage GPU inference infrastructure. The role requires strong software engineering, reconciliation or orchestration experience, and event-driven systems expertise.

Okta

Okta

Bengaluru, India

Staff Software Engineer
No salary listedHybrid7+ YOEDevOps / SRE

Builds and mentors development of scalable cloud tooling, Continuous Delivery platforms, Infrastructure as Code automation, and supporting microservices across AWS environments. The role requires substantial backend software development experience with Java, Go, or Python, plus Terraform, CI/CD, containers, and distributed systems expertise.

Fal

Fal

Remote

Senior/Staff Kubernetes Infrastructure Engineer
$180k+/yrRemote5+ YOEDevOps / SRE

Build and operate high-performance customer compute environments spanning bare metal, Kubernetes, Slurm, GPUs, networking, storage, and observability. The role requires 5+ years of production Linux infrastructure experience and strong expertise in bare-metal Kubernetes, NVIDIA GPUs, virtualization, and networking.