# Site Reliability Engineer 2

**Company:** [Kong](https://hotfix.jobs/companies/kong)
**Location:** Remote
**Role:** DevOps / SRE
**Salary:** $123k – $150k/yr
**Skills:** Kubernetes, Terraform, Terragrunt, Helm, Argo CD, CI/CD, GitOps, Go, Python, Bash, Linux, AWS, GCP, Azure, Postgres
**Posted:** 2026-08-25

> Operate and scale Kong’s multi-region SaaS platform across major cloud providers, Kubernetes, and distributed data systems. The role requires strong infrastructure automation, observability, CI/CD, and production reliability experience, with participation in a global on-call rotation.

## Job Description

## Responsibilities
- Operate and scale Kong’s global, multi-region SaaS platform across AWS, Google Cloud, and Azure.
- Build and maintain Kubernetes-based infrastructure and deployment workflows using Terraform/Terragrunt, Helm, and ArgoCD.
- Design and optimize high-availability, low-latency data and caching layers, including PostgreSQL, Redis, ClickHouse, and Druid.
- Operate Kong Gateway and Kong Mesh environments supporting hybrid and distributed architectures.
- Develop CI/CD pipelines and GitOps workflows for consistent service delivery and infrastructure changes.
- Improve observability and incident response using Datadog, Prometheus, Grafana, and Thanos; define and track SLOs.
- Collaborate with development and security teams to meet reliability, security, and regulatory standards.
- Participate in a global 24/7 on-call rotation and improve operational playbooks and postmortem practices.
- Lead scaling initiatives that improve elasticity, reliability, and cost efficiency.

## Requirements
- Bachelor’s degree in Computer Science or equivalent practical experience.
- Experience managing enterprise-scale SaaS or PaaS systems in multi-region, multi-tenant, secure environments.
- Deep Kubernetes expertise, including cluster and networking troubleshooting and fault-tolerant, scalable design.
- Strong proficiency with Terraform or Terragrunt.
- Experience with CI/CD pipelines and GitOps workflows, including ArgoCD, Atlantis, or Helm.
- Proficiency in Go, Python, or Bash for automation and tooling.
- Understanding of Linux/Unix systems, DNS, TLS/SSL, HTTP, load balancers, and distributed systems.
- Experience with API gateway and service mesh technologies.
- Familiarity with Kafka and observability platforms such as Datadog, Prometheus, and Grafana.
- Experience in a 24/7/365 production support environment.

## Nice-to-Haves
- Experience with Kong Gateway, Kong Mesh, or similar service connectivity technologies.
- Experience operating ClickHouse, Druid, or other time-series and analytics databases.
- Experience managing PostgreSQL and Redis in multi-region configurations.
- Knowledge of AWS networking, Azure VNet, or Google Cloud Network Connectivity Center.
- Understanding of disaster recovery, resiliency testing, and compliance-driven reliability practices.

## Similar jobs

- [DevOps Engineer](https://hotfix.jobs/jobs/0ba7c9d4-0bad-447f-a491-48b3469ea0e4) - Fusion Health - Woodbridge, NJ - $120k – $140k/yr
- [Infrastructure Engineer](https://hotfix.jobs/jobs/d60b169f-55da-4d33-8874-fc12682768e8) - Mercor - San Francisco, CA - $130k – $500k/yr
- [Member of Technical Staff, Mercor Enterprise Platform](https://hotfix.jobs/jobs/49f85fd5-0cb5-49e0-85c5-1bd1d60d091f) - Mercor - San Francisco, CA - $130k – $500k/yr
- [Site Reliability Engineer II](https://hotfix.jobs/jobs/36cfd88a-a2bd-4f65-a5c2-1c9faf7ace63) - PagerDuty - Atlanta, GA - $113k – $172k/yr
- [Software Engineer, Cloud Infrastructure](https://hotfix.jobs/jobs/949677d6-6d57-49e8-acf8-017a14790019) - Beacon AI - San Carlos, CA - $135k – $260k/yr

**Apply:** https://hotfix.jobs/jobs/d5711c27-d8d5-4321-bf3b-77ff440c4709
**Canonical:** https://hotfix.jobs/jobs/d5711c27-d8d5-4321-bf3b-77ff440c4709