Skip to content
OrkesOrkes

Site Reliability Engineer

Owns reliability, observability, incident response, and automation for cloud-based production systems. The role requires 5+ years in SRE, DevOps, platform engineering, or related infrastructure work, with strong Kubernetes, cloud, distributed-systems, and infrastructure-automation experience.

About the job

Responsibilities

  • Own the reliability, availability, and performance of production systems running in cloud environments.
  • Define and monitor SLIs/SLOs and help manage error budgets across the platform.
  • Lead incident response, including detection, triage, mitigation, and postmortems.
  • Improve observability through logging, monitoring, alerting, and dashboards.
  • Automate operational workflows and reduce manual toil.
  • Partner with engineering teams to improve system resiliency and scalability.
  • Assist with capacity planning, infrastructure optimization, and performance tuning.
  • Build internal tooling, runbooks, and operational best practices.
  • Support Kubernetes-based infrastructure and distributed systems at scale.
  • Act as an escalation point for complex production and platform issues.

Requirements

  • 5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or related infrastructure roles.
  • Strong experience with cloud platforms such as AWS, Google Cloud, or Azure.
  • Hands-on experience with Kubernetes and containerized environments.
  • Strong understanding of distributed systems and microservices architecture.
  • Experience with observability tools such as Prometheus, Grafana, Datadog, ELK, or OpenTelemetry.
  • Proficiency with infrastructure automation and scripting, including Terraform, Python, or Bash.
  • Experience managing CI/CD pipelines and deployment automation.
  • Strong troubleshooting and incident management skills.
  • Ability to work cross-functionally and communicate effectively during high-pressure situations.

Nice to Have

  • Experience supporting large-scale SaaS or cloud-native platforms.
  • Familiarity with workflow orchestration technologies such as Conductor, Temporal, or Camunda.
  • Experience with Kafka, messaging systems, or event-driven architectures.
  • Knowledge of security best practices and cloud infrastructure hardening.
  • Open-source contributions or a strong systems engineering background.

Compensation and Benefits

  • Base salary: $125,000–$250,000 USD.
  • Compensation varies based on skills, experience, job scope, location and cost of living, and competitive market data for the country and role.
  • Comprehensive health coverage, including medical, dental, and vision.
  • Flexible PTO.
  • Personal development support.
  • Expected travel: 15–20%.

Skills

AWS, GCP, Azure, Kubernetes, Docker, Prometheus, Grafana, Datadog, OpenTelemetry, Terraform, Python, Bash, CI/CD, Kafka, Distributed Systems

Clickhouse

Clickhouse

Singapore
Release Engineer - Data Plane Internal Tooling and Productivity
No salary listedRemote5+ YOEDevOps / SRE

Own large-scale ClickHouse cluster upgrades and production operations while building tooling that improves release safety and automation. The role requires 5+ years operating stateful distributed systems, cloud and Kubernetes experience, strong debugging skills, and Go development experience.

Granica

Granica

Remote

Software Engineer, Infrastructure
No salary listedRemote5+ YOEDevOps / SRE

Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.

Supabase

Supabase

Remote

Platform Engineer - Compute Capacity
No salary listedRemote5+ YOEDevOps / SRE

Platform engineer responsible for forecasting and automating compute capacity across regions, including reservations, fleet reconciliation, observability, and cost optimization. Requires 5+ years in infrastructure, SRE, platform, or capacity engineering plus production software and AWS EC2 experience.

Alpaca

Alpaca

Remote

Production Support Engineer
No salary listedRemote4+ YOEDevOps / SRE

Provides hands-on L2 technical escalation support for enterprise customers in the APAC region, troubleshooting distributed systems and APIs while leading root-cause analysis, support process improvements, and technical documentation. Requires 4+ years of support or escalation engineering experience.

Stripe

Stripe

Sydney, Australia

Software Engineer, Core Infrastructure
No salary listedOn-site5+ YOEDevOps / SRE

Build and operate distributed cloud infrastructure and platform services that support product teams globally. The role requires 5+ years of software development experience, strong distributed-systems expertise, and experience with cloud infrastructure, reliability, and observability.