# Site Reliability Engineer

**Company:** [Orkes](https://hotfix.jobs/companies/orkes)
**Location:** Remote
**Role:** DevOps / SRE
**Salary:** $125k – $250k/yr
**Experience:** 5+ years
**Skills:** AWS, GCP, Azure, Kubernetes, Docker, Prometheus, Grafana, Datadog, OpenTelemetry, Terraform, Python, Bash, CI/CD, Kafka, Distributed Systems
**Posted:** 2026-08-12

> Owns reliability, observability, incident response, and automation for cloud-based production systems. The role requires 5+ years in SRE, DevOps, platform engineering, or related infrastructure work, with strong Kubernetes, cloud, distributed-systems, and infrastructure-automation experience.

## Job Description

## Responsibilities
- Own the reliability, availability, and performance of production systems running in cloud environments.
- Define and monitor SLIs/SLOs and help manage error budgets across the platform.
- Lead incident response, including detection, triage, mitigation, and postmortems.
- Improve observability through logging, monitoring, alerting, and dashboards.
- Automate operational workflows and reduce manual toil.
- Partner with engineering teams to improve system resiliency and scalability.
- Assist with capacity planning, infrastructure optimization, and performance tuning.
- Build internal tooling, runbooks, and operational best practices.
- Support Kubernetes-based infrastructure and distributed systems at scale.
- Act as an escalation point for complex production and platform issues.

## Requirements
- 5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or related infrastructure roles.
- Strong experience with cloud platforms such as AWS, Google Cloud, or Azure.
- Hands-on experience with Kubernetes and containerized environments.
- Strong understanding of distributed systems and microservices architecture.
- Experience with observability tools such as Prometheus, Grafana, Datadog, ELK, or OpenTelemetry.
- Proficiency with infrastructure automation and scripting, including Terraform, Python, or Bash.
- Experience managing CI/CD pipelines and deployment automation.
- Strong troubleshooting and incident management skills.
- Ability to work cross-functionally and communicate effectively during high-pressure situations.

## Nice to Have
- Experience supporting large-scale SaaS or cloud-native platforms.
- Familiarity with workflow orchestration technologies such as Conductor, Temporal, or Camunda.
- Experience with Kafka, messaging systems, or event-driven architectures.
- Knowledge of security best practices and cloud infrastructure hardening.
- Open-source contributions or a strong systems engineering background.

## Compensation and Benefits
- Base salary: **$125,000–$250,000 USD**.
- Compensation varies based on skills, experience, job scope, location and cost of living, and competitive market data for the country and role.
- Comprehensive health coverage, including medical, dental, and vision.
- Flexible PTO.
- Personal development support.
- Expected travel: 15–20%.

## Similar jobs

- [Release Engineer - Data Plane Internal Tooling and Productivity](https://hotfix.jobs/jobs/ee184847-0b76-42f4-8c1f-157f199626d3) - Clickhouse - Remote
- [Software Engineer, Infrastructure](https://hotfix.jobs/jobs/585da47a-e02d-4c63-9816-248a2faa9b5b) - Granica - Remote
- [Platform Engineer - Compute Capacity](https://hotfix.jobs/jobs/4d270716-1d19-4d06-a9a3-4fdc533a14c4) - Supabase - Remote
- [Production Support Engineer](https://hotfix.jobs/jobs/328c6ec8-1e18-42d6-99c0-515d7801c9ac) - Alpaca - Remote
- [Software Engineer, Core Infrastructure](https://hotfix.jobs/jobs/4d0a1e35-8a83-4b83-a45b-4bfdcbbcc41a) - Stripe - Sydney, Australia

**Apply:** https://hotfix.jobs/jobs/628749a2-4c75-43cf-b5e2-4e12f985a834
**Canonical:** https://hotfix.jobs/jobs/628749a2-4c75-43cf-b5e2-4e12f985a834