Site Reliability Engineer I
Supports and evolves the networking, compute, Kubernetes, and ingress infrastructure powering PagerDuty’s real-time platform. Requires 0–1+ years of relevant experience, Linux production operations, cloud infrastructure knowledge, programming proficiency, and Infrastructure as Code experience.
About the job
Responsibilities
- Support and improve foundational infrastructure, including networking, compute platforms, Kubernetes clusters, and ingress/traffic management systems.
- Contribute to the reliability and scalability of the core platform by hardening existing systems and rolling out new infrastructure capabilities.
- Participate in agile standups, planning, and retrospectives; communicate progress and risks early.
- Stay current on technical trends and suggest innovative tools and approaches.
- Monitor system health using metrics, logs, and alerts.
- Participate in 24/7 on-call rotations to detect, respond to, and resolve incidents.
Requirements
- 0–1+ years of experience in Site Reliability Engineering, DevOps, or Platform Engineering.
- Hands-on experience operating Linux-based systems in production environments.
- Working knowledge of networking fundamentals, including load balancing, DNS, TLS, and ingress traffic flow.
- Experience with container orchestration such as EKS or Kubernetes.
- Experience with cloud-native infrastructure, including networking and compute concepts, on AWS, GCP, or Azure.
- Proficiency in at least one programming language, such as Python, Ruby, or Go.
- Experience with Infrastructure as Code, such as Terraform or CloudFormation.
Nice-to-haves
- Experience with AWS cloud networking concepts, including VPCs, subnets, routing, security groups, and load balancers.
- Experience operating or contributing to production Kubernetes platforms such as EKS, including cluster upgrades, networking, or ingress configuration.
- Experience with monitoring, observability, and logging platforms such as Datadog, New Relic, Sumo Logic, Splunk, Prometheus, or Grafana.
- Familiarity with service meshes, ingress controllers, or API gateways such as Envoy, Istio, or NGINX.
Compensation and benefits
- Salary range: $98,000–$148,500.
- Comprehensive benefits package, flexible work arrangements, company equity, ESPP, retirement or pension plan, paid vacation, paid holidays and sick leave, wellness days, paid parental leave, paid volunteer time, company-wide hack weeks, and mental wellness programs.
- Benefit eligibility may vary by role, region, and tenure.
Skills
Linux, Networking, Kubernetes, Amazon Eks, AWS, Python, Ruby, Go, Terraform, CloudFormation, Prometheus, Grafana, Istio, Nginx
Similar jobs
DevOps / SRE jobsSupports cloud infrastructure, automation, CI/CD, monitoring, and service reliability while learning alongside a global DevOps team. The entry-level role requires a bachelor’s degree, foundational systems knowledge, and exposure to cloud and DevOps tools.
Build and Release Engineer responsible for release orchestration, CI/CD pipelines, artifact lifecycle management, and an internal release portal. The role requires strong software development skills, Git expertise, and graduation by December 2026.
Winter infrastructure and site reliability internship focused on building and operating minimal, on-premises backend infrastructure for a semiconductor fabrication facility. The role requires systems programming, Linux, networking, distributed systems, and hands-on infrastructure or automation experience.
Infrastructure and site reliability intern building and operating on-premises backend infrastructure for a semiconductor fabrication environment. The role emphasizes systems programming, Linux, networking, reliability, observability, automation, and performance engineering.
Build Mercury’s secure, observable infrastructure platform across AWS, networking, containers, and developer tooling. The role requires strong Linux fundamentals, cloud-native experience, technical writing ability, and software development skills, with opportunities to support AI-agent infrastructure.