Site Reliability Engineer II
Operate and scale a cloud-native CTV advertising platform on AWS and Kubernetes. Focus on reliability, GitOps workflows, infrastructure automation, observability, and incident response.
About the job
What you’ll do
- Ensuring the reliability, availability, and performance of production infrastructure and platform services
- Operating and scaling Kubernetes platforms, including governance and support for multi-tenant workloads
- Managing GitOps-based deployment workflows using ArgoCD and Helm
- Supporting infrastructure provisioning and change management through Terraform/Terragrunt
- Building and supporting CI/CD automation and deployment workflows using GitHub Actions
- Participating in incident response, root cause analysis, and post-incident improvement initiatives
- Reducing operational toil through scripting, tooling, and process automation
- Advancing observability practices across logs, metrics, traces, dashboards, and alerting
- Supporting secure secrets integration, IAM-aware operations, and platform guardrails
- Partnering closely with application, security, and platform teams to improve reliability and delivery outcomes
What we're looking for
- 4+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Cloud Infrastructure
- Strong hands-on experience operating AWS in production environments
- Good expertise in Kubernetes, including cluster operations, troubleshooting, workload reliability, and platform administration
- Experience with Kubernetes multi-tenancy, including namespaces, RBAC, quotas, policies, and tenant isolation patterns
- Experience implementing and operating ArgoCD within a GitOps delivery model
- Strong hands-on experience with Helm
- Experience with Terraform/Terragrunt for infrastructure provisioning and environment management
- Solid scripting and automation skills using Bash and/or Python
- Experience building, maintaining, or supporting CI/CD pipelines, ideally using GitHub Actions
- Strong troubleshooting skills across Linux, containers, IAM, networking, and distributed systems
- Experience with monitoring, alerting, and observability in production environments
- Demonstrated ownership mindset with experience handling incidents and resolving production issues
- Strong collaboration and communication skills, with the ability to work effectively across engineering, security, and platform teams
- Bachelor’s degree in computer science, engineering, a related field or equivalent experience
- Demonstrated ability to use AI to improve speed and quality in your day-to-day workflow for relevant outputs
- Strong track record of critical evaluation and verification of AI-assisted work (e.g., testing, source-checking, data validation, peer review)
- High integrity and ownership: you protect sensitive data, avoid over-reliance on AI, and remain accountable for final decisions and deliverables
Skills
AWS, Kubernetes, EKS, Argo CD, Helm, Terraform, Terragrunt, GitHub Actions, Bash, Python, Linux, IAM, CI/CD, Observability
Similar jobs
DevOps / SRE jobsOperates and evolves foundational networking, compute, Kubernetes, and ingress infrastructure for PagerDuty’s real-time platform. Requires 3+ years in SRE, DevOps, or platform engineering, with Linux production operations, cloud infrastructure, programming, and Infrastructure as Code experience.
Owns secure, scalable Azure infrastructure for healthcare applications, including cloud migrations, Terraform-based automation, CI/CD pipelines, monitoring, and compliance. Requires 3–5+ years of Azure experience and strong DevOps and cloud-security expertise.
Operate and scale Kong’s multi-region SaaS platform across major cloud providers, Kubernetes, and distributed data systems. The role requires strong infrastructure automation, observability, CI/CD, and production reliability experience, with participation in a global on-call rotation.
Builds and scales highly available infrastructure using AWS, Terraform, and Docker to support rapid growth and AI workloads. Collaborates with product and research teams on architectures, CI/CD, monitoring, and performance optimization.
Build and operate Mercor’s enterprise agent platform across security, routing, isolated execution, orchestration, deployment, and production scalability. The role requires 5+ years building high-scale platforms, architectural ownership, and experience with core infrastructure primitives across multiple clouds.