# Site Reliability Engineer

**Company:** [Prove AI](https://hotfix.jobs/companies/prove-ai)
**Location:** Remote
**Role:** DevOps / SRE
**Salary:** $120k – $185k/yr
**Experience:** 5+ years
**Skills:** AWS, Terraform, Kubernetes, OpenTelemetry, Prometheus, Grafana, Jaeger, Splunk, Go, Python, Infrastructure As Code, Distributed Systems, Microservices, Service Mesh
**Posted:** 2026-09-11

> The Site Reliability Engineer designs and operates highly available AWS infrastructure, observability systems, Kubernetes platforms, and automation while participating in incident response. Senior candidates require at least five years of relevant experience; the role also has an IC3 track requiring three years.

## Job Description

## Responsibilities

### Observability
- Design and implement observability solutions across infrastructure and applications.
- Establish metrics, logging, and tracing systems for rapid issue identification and resolution.
- Create alerting thresholds and automated responses based on service-level objectives (SLOs).
- Provide actionable insights into service-to-service communications.

### Infrastructure and Reliability
- Design, build, and maintain scalable AWS cloud infrastructure.
- Implement infrastructure as code using Terraform and related tools.
- Automate operational tasks to reduce toil and improve efficiency.
- Implement security compliance and least-privilege access controls.
- Deploy and scale container-based applications using custom metrics.
- Improve system reliability, performance, scalability, and cost efficiency.
- Scale developer experiences through an opinionated platform approach.

### Incident Response
- Participate in a 24/7 on-call rotation to support high system availability.
- Conduct post-incident reviews and implement preventative measures.
- Use observability data for root-cause analysis and system improvements.

## Requirements

### Senior-Level Track
- 5+ years of experience in Site Reliability Engineering, Platform Engineering, or equivalent experience.
- Expert knowledge of observability platforms and practices, including OpenTelemetry, Prometheus, Grafana, Jaeger, ELK, or Splunk.
- Experience with Kubernetes and container orchestration.
- Strong experience with infrastructure-as-code tools such as Terraform, Spacelift, or Pulumi.
- Proficiency in at least one programming language, including Go or Python.
- Deep understanding of cloud platforms, preferably AWS.
- Bachelor’s degree in Computer Science, Engineering, or equivalent practical experience.

### IC3 Track
- 3+ years of experience in Site Reliability or Platform Engineering.
- Deep understanding of cloud platforms, particularly AWS.
- Strong experience with Kubernetes and container orchestration.
- Experience with Terraform and infrastructure-as-code tools.
- Bachelor’s degree in Computer Science, Engineering, or equivalent practical experience.

## Nice-to-Haves
- Experience with distributed systems and microservice architectures.
- Experience in high-compliance environments.
- Experience instrumenting code with OpenTelemetry.
- Familiarity with service mesh technologies.
- Contributions to open-source projects.
- Experience in identity verification or financial technology.
- Application development experience.
- Holistic monitoring and alerting experience for developing platforms.

## Compensation
- Site Reliability Engineer, Metro 2: $130,000–$150,000; Metro 3: $120,000–$135,000.
- Senior Site Reliability Engineer, Metro 2: $166,000–$185,000; Metro 3: $153,000–$171,000.
- Additional variable commission or company bonus may apply.
- Benefits include equity, wellness programs, 401(k) matching, flexible vacation and hours, and comprehensive medical benefits.

## Similar jobs

- [Site Reliability Engineer 2](https://hotfix.jobs/jobs/d5711c27-d8d5-4321-bf3b-77ff440c4709) - Kong - Remote - $123k – $150k/yr
- [Site Reliability Engineer II](https://hotfix.jobs/jobs/36cfd88a-a2bd-4f65-a5c2-1c9faf7ace63) - PagerDuty - Atlanta, GA - $113k – $172k/yr
- [Infrastructure Engineer](https://hotfix.jobs/jobs/d60b169f-55da-4d33-8874-fc12682768e8) - Mercor - San Francisco, CA - $130k – $500k/yr
- [Member of Technical Staff, Mercor Enterprise Platform](https://hotfix.jobs/jobs/49f85fd5-0cb5-49e0-85c5-1bd1d60d091f) - Mercor - San Francisco, CA - $130k – $500k/yr
- [Software Engineer, Cloud Infrastructure](https://hotfix.jobs/jobs/949677d6-6d57-49e8-acf8-017a14790019) - Beacon AI - San Carlos, CA - $135k – $260k/yr

**Apply:** https://hotfix.jobs/jobs/a92e9ece-5bd0-4e1e-9b52-ea4f36c00585
**Canonical:** https://hotfix.jobs/jobs/a92e9ece-5bd0-4e1e-9b52-ea4f36c00585