# Senior SRE, Managed Gateways

**Company:** [Kong](https://hotfix.jobs/companies/kong)
**Location:** Remote
**Role:** DevOps / SRE
**Salary:** $118k – $167k/yr
**Experience:** 7+ years
**Skills:** Kubernetes, Go, AWS, GCP, Azure, Terraform, Ansible, Prometheus, Grafana, elk stack, Datadog, CI/CD, service mesh, istio, Postgres
**Posted:** 2026-07-23

> Senior Site Reliability Engineer owning production reliability and enterprise customer implementations for Kong's fast-growing Managed Gateways SaaS product across AWS, GCP, and Azure. Requires deep Kubernetes, cloud-native, and Golang expertise plus customer-facing technical leadership.

## Job Description

## What You’ll Do

### Platform & Reliability Engineering
- Lead, mentor, and inspire a high-performing team of Site Reliability Engineers dedicated to Kong's Managed Gateway offerings.
- Architect and implement robust, scalable, and fault-tolerant cloud-native systems using technologies like Kubernetes, Golang, and major cloud providers.
- Own the end-to-end operational lifecycle, from proactive monitoring and alerting to incident response and blameless post-mortems, ensuring continuous service improvement.
- Drive a culture of developer delight by implementing automation, self-service tooling, and streamlined workflows for deploying and managing API gateways.
- Define, track, and report on key SLOs and SLIs to ensure optimal performance and reliability of Managed Gateways.
- Champion technical debt prevention and advocate for architectural best practices that enhance system resilience and reduce operational toil.
- Collaborate cross-functionally with Product, engineering, and Customer Success to influence roadmap decisions and ensure operational readiness for new features.

### Enterprise Implementation Engineering
- Partner directly with enterprise customers — working alongside Product leadership, Professional Services, and Customer Success — to drive end-to-end onboarding and implementation of Cloud Gateways, and productize recurring implementation patterns into repeatable playbooks and platform capabilities.
- Bring deep, cross-cloud breadth (AWS, GCP, Azure) to handle unique customer topologies and turn complex setups into successful, production-ready deployments.
- Serve as the technical owner of the customer relationship through implementation, primarily supporting our North America customer base, and be the escalation point Customer Success leans on for technically complex accounts.
- Feed real-world implementation patterns and customer constraints back to Product to further contribute the roadmap.

## What You’ll Bring

### The Toolkit
- Extensive experience as a Site Reliability Engineer, focusing on highly available and distributed systems.
- Deep expertise with Kubernetes and cloud-native architectures, preferably across multiple public cloud providers (AWS, GCP, Azure).
- Strong proficiency in Golang or similar modern programming languages for automation and tool development.
- Proven track record in building and maintaining CI/CD pipelines and infrastructure as code (Terraform, Ansible).
- In-depth knowledge of monitoring, logging, and alerting systems (e.g., Prometheus, Grafana, ELK stack, Datadog).
- Experience with managed services, API gateways, or similar network infrastructure is highly desirable.

### The Kong DNA
- You take immense ownership of your systems, treating reliability as a first-class feature.
- You operate with a sense of urgency, especially in critical situations, and drive quick, effective resolutions.
- You thrive in a collaborative environment, actively sharing knowledge and elevating the entire team.
- Kong moves fast, and our team's spread across continents and time zones — plans shift mid-flight, and things don't always line up neatly. You don't need everything settled to do good work. You bring your own calm to the noise, figure things out as you go, and help the people around you do the same.

## Bonus Points
- Experience with Service Mesh technologies (e.g., Istio, Linkerd).
- Familiarity with database administration for high-throughput systems (PostgreSQL, Cassandra).
- Contributions to open-source SRE tools or projects.
- Relevant cloud certifications (e.g., AWS Certified DevOps Engineer, CKA).

## Similar roles

- [Senior Engineer, Platform Infrastructure](https://hotfix.jobs/jobs/cf15e8c3-ee2a-4f6a-abd6-8a54bce776fd) - Shield AI - San Diego, CA - $120k – $180k/yr
- [Senior Site Reliability Engineer](https://hotfix.jobs/jobs/a6949459-b2bc-44d0-99bf-23c9890dcf26) - PrizePicks - Remote - $120k – $175k/yr
- [Senior Engineer, Software Engineering Tools (R4913)](https://hotfix.jobs/jobs/06ff6c1e-853f-4bd7-a741-7865d8b9c5c0) - Shield AI - Dallas, TX - $120k – $190k/yr
- [Senior Infrastructure Engineer](https://hotfix.jobs/jobs/c8f2b95b-8d32-4d8e-8ba2-0a29dfc5158a) - Bland AI - San Francisco, CA - $120k – $200k/yr
- [Senior Infrastructure Engineer](https://hotfix.jobs/jobs/d3847fd2-6f31-4a9c-bd48-e35523ad26a1) - LiveKit - Remote - $120k – $250k/yr

**Apply:** https://hotfix.jobs/jobs/ad32fcba-eefa-4554-820d-95dba7cbf225
**Canonical:** https://hotfix.jobs/jobs/ad32fcba-eefa-4554-820d-95dba7cbf225