Staff Site Reliability Engineer - Kubernetes
Build and operate secure, highly available Kubernetes platforms on AWS, including cluster creation, scaling, service mesh, automation, and incident response. The Staff-level role requires deep experience with Kubernetes, Terraform, AWS, Helm, Karpenter, and Istio.
About the job
Responsibilities
- Design, implement, and maintain highly available, scalable, and fault-tolerant Kubernetes platforms for production workloads.
- Build, manage, and optimize AWS infrastructure, including EKS, ECS, S3, VPCs, RDS, IAM, and related services.
- Create, maintain, and manage Helm charts for production Kubernetes deployments.
- Implement and manage Karpenter for dynamic Kubernetes cluster scaling and resource optimization.
- Configure and manage Istio for service-to-service communication, security, observability, traffic management, service discovery, and policy enforcement.
- Automate infrastructure and application deployment, scaling, and management through CI/CD pipelines.
- Respond to incidents and troubleshoot performance, availability, and security issues.
- Implement secure cloud infrastructure with appropriate access controls, network security, and compliance practices.
- Document Kubernetes platform setup, operational procedures, and best practices, and share knowledge across teams.
Requirements
- 5+ years of experience with AWS.
- 4+ years of experience with Kubernetes and Helm.
- 4+ years of experience with Terraform.
- Experience with multi-region cloud environments and cloud-native architectures.
- Strong expertise creating and managing Kubernetes platforms, including highly available clusters, networking, and storage.
- Hands-on experience deploying and managing applications with Helm.
- Experience with Karpenter for dynamic Kubernetes scaling.
- Experience managing and securing Istio, including traffic management, security, and observability.
- Proficiency with CI/CD pipelines and automation tools such as Jenkins, GitLab, CircleCI, Ansible, and Spinnaker.
- Strong scripting and automation skills in Python, Bash, or Go.
- Experience with monitoring, logging, and alerting tools such as Prometheus, Grafana, CloudWatch, and ELK Stack.
- Ability to meet U.S. Person status requirements for access to federal environments or protected federal data.
- Ability to complete in-person onboarding and travel to the San Francisco, CA headquarters or Chicago office during the first week.
Nice-to-haves
- Knowledge of cloud and Kubernetes security practices, including RBAC, encryption, and compliance frameworks.
- Familiarity with Docker and containerization.
- Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent professional experience.
- CKA, CKAD, or AWS Certified DevOps Engineer certification.
Skills
Kubernetes, Helm, Terraform, AWS, Amazon Eks, Amazon Ecs, Amazon S3, Amazon Vpc, Amazon Rds, IAM, Karpenter, Istio, CI/CD, Python, Bash
Similar jobs
DevOps / SRE jobsLeads reliability and networking for highly available, secure cloud services in Okta’s Federal SRE organization. The role requires active TS/SCI clearance with full-scope polygraph, Federal/DoD compliance experience, and deep expertise in AWS networking, Terraform, observability, and automation.
Leads reliability engineering for highly available, FedRAMP-compliant cloud services, including infrastructure architecture, automation, observability, incident response, and operational standards. Requires extensive Kubernetes, cloud, software engineering, and cross-team technical leadership experience, plus US-person eligibility and residence on US soil.
Build and operate high-performance customer compute environments spanning bare metal, Kubernetes, Slurm, GPUs, networking, storage, and observability. The role requires 5+ years of production Linux infrastructure experience and strong expertise in bare-metal Kubernetes, NVIDIA GPUs, virtualization, and networking.
Leads strategic production engineering initiatives that improve the reliability, scalability, observability, and security of large-scale platforms. The role requires 7+ years of relevant experience, strong coding skills, and expertise in reliability practices such as SLIs, SLOs, and incident management.
Leads the operational reliability, security, observability, deployment standards, and governance of Databricks for enterprise data workloads. Requires 12+ years in platform, SRE, or cloud data infrastructure engineering plus production Databricks experience and expertise in CI/CD, secure execution, and regulated environments.