# Site Reliability Engineer II

**Company:** [PagerDuty](https://hotfix.jobs/companies/pagerduty)
**Location:** Atlanta, GA
**Role:** DevOps / SRE
**Salary:** $113k – $172k/yr
**Experience:** 3+ years
**Skills:** Linux, Networking, Kubernetes, Amazon Eks, AWS, GCP, Microsoft Azure, Python, Ruby, Go, Terraform, CloudFormation, Prometheus, Grafana, Istio
**Posted:** 2026-09-01

> Operates and evolves foundational networking, compute, Kubernetes, and ingress infrastructure for PagerDuty’s real-time platform. Requires 3+ years in SRE, DevOps, or platform engineering, with Linux production operations, cloud infrastructure, programming, and Infrastructure as Code experience.

## Job Description

## Responsibilities
- Support and improve foundational infrastructure, including networking, compute platforms, Kubernetes clusters, and ingress/traffic management systems.
- Improve the reliability and scalability of the core platform by hardening existing systems and rolling out new infrastructure capabilities.
- Participate in agile ceremonies and communicate progress and risks proactively.
- Monitor system health using metrics, logs, and alerts.
- Participate in 24/7 on-call rotations to detect, respond to, and resolve incidents.
- Stay current on technical trends and suggest innovative tools and approaches.

## Requirements
- 3+ years of experience in Site Reliability Engineering, DevOps, or Platform Engineering.
- Hands-on experience operating Linux-based systems in production.
- Working knowledge of networking fundamentals, including load balancing, DNS, TLS, and ingress traffic flow.
- Experience with container orchestration such as EKS or Kubernetes.
- Experience with cloud-native infrastructure, including networking and compute concepts, on AWS, GCP, or Azure.
- Proficiency in at least one programming language, such as Python, Ruby, or Go.
- Experience with Infrastructure as Code, such as Terraform or CloudFormation.

## Nice-to-haves
- Experience with AWS cloud networking, including VPCs, subnets, routing, security groups, and load balancers.
- Experience operating or contributing to production Kubernetes platforms, including cluster upgrades, networking, or ingress configuration.
- Experience with monitoring, observability, and logging platforms such as Datadog, New Relic, Sumo Logic, Splunk, Prometheus, or Grafana.
- Familiarity with service meshes, ingress controllers, or API gateways such as Envoy, Istio, or NGINX.

## Compensation and Benefits
- Salary range: $113,000–$171,600 annually.
- Comprehensive benefits package, flexible work arrangements, company equity, ESPP, retirement or pension plan, paid vacation, paid holidays and sick leave, wellness days, paid parental leave, paid volunteer time, company-wide hack weeks, and mental wellness programs.

## Similar jobs

- [DevOps Engineer](https://hotfix.jobs/jobs/0ba7c9d4-0bad-447f-a491-48b3469ea0e4) - Fusion Health - Woodbridge, NJ - $120k – $140k/yr
- [Site Reliability Engineer 2](https://hotfix.jobs/jobs/d5711c27-d8d5-4321-bf3b-77ff440c4709) - Kong - Remote - $123k – $150k/yr
- [Infrastructure Engineer](https://hotfix.jobs/jobs/d60b169f-55da-4d33-8874-fc12682768e8) - Mercor - San Francisco, CA - $130k – $500k/yr
- [Member of Technical Staff, Mercor Enterprise Platform](https://hotfix.jobs/jobs/49f85fd5-0cb5-49e0-85c5-1bd1d60d091f) - Mercor - San Francisco, CA - $130k – $500k/yr
- [Software Engineer, Cloud Infrastructure](https://hotfix.jobs/jobs/949677d6-6d57-49e8-acf8-017a14790019) - Beacon AI - San Carlos, CA - $135k – $260k/yr

**Apply:** https://hotfix.jobs/jobs/36cfd88a-a2bd-4f65-a5c2-1c9faf7ace63
**Canonical:** https://hotfix.jobs/jobs/36cfd88a-a2bd-4f65-a5c2-1c9faf7ace63