Site Reliability Engineer II
Operates and evolves foundational networking, compute, Kubernetes, and ingress infrastructure for PagerDuty’s real-time platform. Requires 3+ years in SRE, DevOps, or platform engineering, with Linux production operations, cloud infrastructure, programming, and Infrastructure as Code experience.
About the job
Responsibilities
- Support and improve foundational infrastructure, including networking, compute platforms, Kubernetes clusters, and ingress/traffic management systems.
- Improve the reliability and scalability of the core platform by hardening existing systems and rolling out new infrastructure capabilities.
- Participate in agile ceremonies and communicate progress and risks proactively.
- Monitor system health using metrics, logs, and alerts.
- Participate in 24/7 on-call rotations to detect, respond to, and resolve incidents.
- Stay current on technical trends and suggest innovative tools and approaches.
Requirements
- 3+ years of experience in Site Reliability Engineering, DevOps, or Platform Engineering.
- Hands-on experience operating Linux-based systems in production.
- Working knowledge of networking fundamentals, including load balancing, DNS, TLS, and ingress traffic flow.
- Experience with container orchestration such as EKS or Kubernetes.
- Experience with cloud-native infrastructure, including networking and compute concepts, on AWS, GCP, or Azure.
- Proficiency in at least one programming language, such as Python, Ruby, or Go.
- Experience with Infrastructure as Code, such as Terraform or CloudFormation.
Nice-to-haves
- Experience with AWS cloud networking, including VPCs, subnets, routing, security groups, and load balancers.
- Experience operating or contributing to production Kubernetes platforms, including cluster upgrades, networking, or ingress configuration.
- Experience with monitoring, observability, and logging platforms such as Datadog, New Relic, Sumo Logic, Splunk, Prometheus, or Grafana.
- Familiarity with service meshes, ingress controllers, or API gateways such as Envoy, Istio, or NGINX.
Compensation and Benefits
- Salary range: $113,000–$171,600 annually.
- Comprehensive benefits package, flexible work arrangements, company equity, ESPP, retirement or pension plan, paid vacation, paid holidays and sick leave, wellness days, paid parental leave, paid volunteer time, company-wide hack weeks, and mental wellness programs.
Skills
Linux, Networking, Kubernetes, Amazon Eks, AWS, GCP, Microsoft Azure, Python, Ruby, Go, Terraform, CloudFormation, Prometheus, Grafana, Istio
Similar jobs
DevOps / SRE jobsOwns secure, scalable Azure infrastructure for healthcare applications, including cloud migrations, Terraform-based automation, CI/CD pipelines, monitoring, and compliance. Requires 3–5+ years of Azure experience and strong DevOps and cloud-security expertise.
Operate and scale Kong’s multi-region SaaS platform across major cloud providers, Kubernetes, and distributed data systems. The role requires strong infrastructure automation, observability, CI/CD, and production reliability experience, with participation in a global on-call rotation.
Builds and scales highly available infrastructure using AWS, Terraform, and Docker to support rapid growth and AI workloads. Collaborates with product and research teams on architectures, CI/CD, monitoring, and performance optimization.
Build and operate Mercor’s enterprise agent platform across security, routing, isolated execution, orchestration, deployment, and production scalability. The role requires 5+ years building high-scale platforms, architectural ownership, and experience with core infrastructure primitives across multiple clouds.
Build and operate AWS cloud, ML, LLM, RAG, and IoT infrastructure, including deployment platforms, data pipelines, vector search, observability, security, and cost controls. The role requires deep AWS experience and production experience with LLM-powered applications.