Site Reliability Engineer
The Site Reliability Engineer designs and operates highly available AWS infrastructure, observability systems, Kubernetes platforms, and automation while participating in incident response. Senior candidates require at least five years of relevant experience; the role also has an IC3 track requiring three years.
About the job
Responsibilities
Observability
- Design and implement observability solutions across infrastructure and applications.
- Establish metrics, logging, and tracing systems for rapid issue identification and resolution.
- Create alerting thresholds and automated responses based on service-level objectives (SLOs).
- Provide actionable insights into service-to-service communications.
Infrastructure and Reliability
- Design, build, and maintain scalable AWS cloud infrastructure.
- Implement infrastructure as code using Terraform and related tools.
- Automate operational tasks to reduce toil and improve efficiency.
- Implement security compliance and least-privilege access controls.
- Deploy and scale container-based applications using custom metrics.
- Improve system reliability, performance, scalability, and cost efficiency.
- Scale developer experiences through an opinionated platform approach.
Incident Response
- Participate in a 24/7 on-call rotation to support high system availability.
- Conduct post-incident reviews and implement preventative measures.
- Use observability data for root-cause analysis and system improvements.
Requirements
Senior-Level Track
- 5+ years of experience in Site Reliability Engineering, Platform Engineering, or equivalent experience.
- Expert knowledge of observability platforms and practices, including OpenTelemetry, Prometheus, Grafana, Jaeger, ELK, or Splunk.
- Experience with Kubernetes and container orchestration.
- Strong experience with infrastructure-as-code tools such as Terraform, Spacelift, or Pulumi.
- Proficiency in at least one programming language, including Go or Python.
- Deep understanding of cloud platforms, preferably AWS.
- Bachelor’s degree in Computer Science, Engineering, or equivalent practical experience.
IC3 Track
- 3+ years of experience in Site Reliability or Platform Engineering.
- Deep understanding of cloud platforms, particularly AWS.
- Strong experience with Kubernetes and container orchestration.
- Experience with Terraform and infrastructure-as-code tools.
- Bachelor’s degree in Computer Science, Engineering, or equivalent practical experience.
Nice-to-Haves
- Experience with distributed systems and microservice architectures.
- Experience in high-compliance environments.
- Experience instrumenting code with OpenTelemetry.
- Familiarity with service mesh technologies.
- Contributions to open-source projects.
- Experience in identity verification or financial technology.
- Application development experience.
- Holistic monitoring and alerting experience for developing platforms.
Compensation
- Site Reliability Engineer, Metro 2: $130,000–$150,000; Metro 3: $120,000–$135,000.
- Senior Site Reliability Engineer, Metro 2: $166,000–$185,000; Metro 3: $153,000–$171,000.
- Additional variable commission or company bonus may apply.
- Benefits include equity, wellness programs, 401(k) matching, flexible vacation and hours, and comprehensive medical benefits.
Skills
AWS, Terraform, Kubernetes, OpenTelemetry, Prometheus, Grafana, Jaeger, Splunk, Go, Python, Infrastructure As Code, Distributed Systems, Microservices, Service Mesh
Similar jobs
DevOps / SRE jobsOperate and scale Kong’s multi-region SaaS platform across major cloud providers, Kubernetes, and distributed data systems. The role requires strong infrastructure automation, observability, CI/CD, and production reliability experience, with participation in a global on-call rotation.
Operates and evolves foundational networking, compute, Kubernetes, and ingress infrastructure for PagerDuty’s real-time platform. Requires 3+ years in SRE, DevOps, or platform engineering, with Linux production operations, cloud infrastructure, programming, and Infrastructure as Code experience.
Builds and scales highly available infrastructure using AWS, Terraform, and Docker to support rapid growth and AI workloads. Collaborates with product and research teams on architectures, CI/CD, monitoring, and performance optimization.
Build and operate Mercor’s enterprise agent platform across security, routing, isolated execution, orchestration, deployment, and production scalability. The role requires 5+ years building high-scale platforms, architectural ownership, and experience with core infrastructure primitives across multiple clouds.
Build and operate AWS cloud, ML, LLM, RAG, and IoT infrastructure, including deployment platforms, data pipelines, vector search, observability, security, and cost controls. The role requires deep AWS experience and production experience with LLM-powered applications.