Staff Site Reliability Engineer
Leads the design, operation, and modernization of multi-cloud infrastructure across AWS and Google Cloud. The role requires deep Kubernetes expertise, SRE practices, infrastructure as code, automation, observability, and at least eight years of relevant experience.
About the job
Responsibilities
- Design, build, and operate highly scalable, reliable, and secure infrastructure powering production systems across AWS and Google Cloud.
- Lead reliability and modernization initiatives, including container platform migrations such as ECS to EKS/GKE and microservice enablement across multi-cloud environments.
- Serve as a technical authority in Kubernetes, cloud infrastructure, and modern CI/CD practices including GitOps and automation pipelines.
- Partner with development teams to architect and enable microservice-based applications, ensuring production readiness, scalability, and observability.
- Implement and manage infrastructure as code with Terraform and Ansible to automate provisioning, scaling, and configuration management across cloud providers.
- Improve observability, performance, and cost efficiency through monitoring, logging, and alerting across AWS and Google Cloud.
- Define SLOs and SLIs, conduct blameless postmortems, and continuously improve incident response.
- Lead complex technical projects from conception to completion, managing timelines and technical dependencies across teams.
- Mentor engineers and foster a culture of reliability, automation, and continuous learning.
- Collaborate with security and compliance partners on infrastructure standards, including IAM Federation and Workload Identity.
- Participate in the on-call rotation and use incidents to improve systems and processes.
Requirements
- 8+ years in SRE, DevOps, or infrastructure engineering roles.
- 3–5 years of production experience with Kubernetes, including EKS and GKE, and ecosystem tools such as Helm and Karpenter.
- 3–5 years of experience with AWS and Google Cloud.
- 3–5 years using Terraform to manage multi-cloud infrastructure.
- 5+ years of coding experience in Python, Go, or similar languages.
- Hands-on experience architecting and operating cloud-native distributed systems.
- Experience leading ECS-to-EKS/GKE migration projects and enabling microservice architectures.
- Proficiency with Terraform, Ansible, or CloudFormation.
- Advanced understanding of CI/CD pipelines, Linux systems, networking fundamentals, and Redis.
- Experience managing databases and caching systems such as RDS, Cloud SQL, Redis/Memorystore, PostgreSQL, and MySQL.
- Experience with observability tools including Prometheus, Grafana, ELK, Loki, OpenTelemetry, and Google Cloud Operations.
- Working knowledge of container security, secrets management, and production compliance.
- Strong communication and problem-solving skills, including cross-team project leadership and mentoring.
- Strong Linux and security fundamentals.
- Bachelor’s degree in Computer Science or equivalent hands-on experience.
Nice to Have
- Experience in SaaS or high-scale, cloud-native environments.
Benefits
- In-person onboarding experience.
- Well-being support.
- Social impact opportunities.
- Talent development and community-building programs.
Skills
AWS, GCP, Kubernetes, Amazon Eks, Google Gke, Terraform, Ansible, Python, Go, CI/CD, Argo Cd, Linux, Redis, Prometheus, Grafana
Similar jobs
DevOps / SRE jobsBuild and operate a Kubernetes-native control plane for provisioning, scheduling, self-healing, and optimizing GPU inference infrastructure. The role requires strong software engineering, durable workflow orchestration, reconciliation systems, event-driven architecture, and platform API experience.
Owns enterprise DevSecOps architecture across Salesforce, NetSuite, Workday, AEM, and modern web platforms. The role requires 8+ years of DevSecOps, SRE, or security engineering experience, strong CI/CD and edge-security expertise, and leadership in secure automation, observability, identity, and compliance.
Build and operate declarative control planes, durable workflows, and self-healing systems that provision and manage GPU inference infrastructure. The role requires strong software engineering, reconciliation or orchestration experience, and event-driven systems expertise.
Builds and mentors development of scalable cloud tooling, Continuous Delivery platforms, Infrastructure as Code automation, and supporting microservices across AWS environments. The role requires substantial backend software development experience with Java, Go, or Python, plus Terraform, CI/CD, containers, and distributed systems expertise.
Build and operate high-performance customer compute environments spanning bare metal, Kubernetes, Slurm, GPUs, networking, storage, and observability. The role requires 5+ years of production Linux infrastructure experience and strong expertise in bare-metal Kubernetes, NVIDIA GPUs, virtualization, and networking.