Skip to content
PinterestPinterest

Site Reliability Engineer II

Operate and scale a cloud-native CTV advertising platform on AWS and Kubernetes. Focus on reliability, GitOps workflows, infrastructure automation, observability, and incident response.

About the job

What you’ll do

  • Ensuring the reliability, availability, and performance of production infrastructure and platform services
  • Operating and scaling Kubernetes platforms, including governance and support for multi-tenant workloads
  • Managing GitOps-based deployment workflows using ArgoCD and Helm
  • Supporting infrastructure provisioning and change management through Terraform/Terragrunt
  • Building and supporting CI/CD automation and deployment workflows using GitHub Actions
  • Participating in incident response, root cause analysis, and post-incident improvement initiatives
  • Reducing operational toil through scripting, tooling, and process automation
  • Advancing observability practices across logs, metrics, traces, dashboards, and alerting
  • Supporting secure secrets integration, IAM-aware operations, and platform guardrails
  • Partnering closely with application, security, and platform teams to improve reliability and delivery outcomes

What we're looking for

  • 4+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Cloud Infrastructure
  • Strong hands-on experience operating AWS in production environments
  • Good expertise in Kubernetes, including cluster operations, troubleshooting, workload reliability, and platform administration
  • Experience with Kubernetes multi-tenancy, including namespaces, RBAC, quotas, policies, and tenant isolation patterns
  • Experience implementing and operating ArgoCD within a GitOps delivery model
  • Strong hands-on experience with Helm
  • Experience with Terraform/Terragrunt for infrastructure provisioning and environment management
  • Solid scripting and automation skills using Bash and/or Python
  • Experience building, maintaining, or supporting CI/CD pipelines, ideally using GitHub Actions
  • Strong troubleshooting skills across Linux, containers, IAM, networking, and distributed systems
  • Experience with monitoring, alerting, and observability in production environments
  • Demonstrated ownership mindset with experience handling incidents and resolving production issues
  • Strong collaboration and communication skills, with the ability to work effectively across engineering, security, and platform teams
  • Bachelor’s degree in computer science, engineering, a related field or equivalent experience
  • Demonstrated ability to use AI to improve speed and quality in your day-to-day workflow for relevant outputs
  • Strong track record of critical evaluation and verification of AI-assisted work (e.g., testing, source-checking, data validation, peer review)
  • High integrity and ownership: you protect sensitive data, avoid over-reliance on AI, and remain accountable for final decisions and deliverables

Skills

AWS, Kubernetes, EKS, Argo CD, Helm, Terraform, Terragrunt, GitHub Actions, Bash, Python, Linux, IAM, CI/CD, Observability

PagerDuty

PagerDuty

Atlanta, GA

Site Reliability Engineer II
$113k+/yrHybrid3+ YOEDevOps / SRE

Operates and evolves foundational networking, compute, Kubernetes, and ingress infrastructure for PagerDuty’s real-time platform. Requires 3+ years in SRE, DevOps, or platform engineering, with Linux production operations, cloud infrastructure, programming, and Infrastructure as Code experience.

Fusion Health

Fusion Health

Woodbridge, NJ

DevOps Engineer
$120k+/yrHybrid5+ YOEDevOps / SRE

Owns secure, scalable Azure infrastructure for healthcare applications, including cloud migrations, Terraform-based automation, CI/CD pipelines, monitoring, and compliance. Requires 3–5+ years of Azure experience and strong DevOps and cloud-security expertise.

Kong

Kong

United States

Site Reliability Engineer 2
$123k+/yrRemoteDevOps / SRE

Operate and scale Kong’s multi-region SaaS platform across major cloud providers, Kubernetes, and distributed data systems. The role requires strong infrastructure automation, observability, CI/CD, and production reliability experience, with participation in a global on-call rotation.

Mercor

Mercor

San Francisco, CA
Infrastructure Engineer
$130k+/yrOn-siteDevOps / SRE

Builds and scales highly available infrastructure using AWS, Terraform, and Docker to support rapid growth and AI workloads. Collaborates with product and research teams on architectures, CI/CD, monitoring, and performance optimization.

Mercor

Mercor

San Francisco, CA

Member of Technical Staff, Mercor Enterprise Platform
$130k+/yrOn-site5+ YOEDevOps / SRE

Build and operate Mercor’s enterprise agent platform across security, routing, isolated execution, orchestration, deployment, and production scalability. The role requires 5+ years building high-scale platforms, architectural ownership, and experience with core infrastructure primitives across multiple clouds.