Skip to content
KongKongOlympia, WA

Senior SRE, Managed Gateways

Senior Site Reliability Engineer owning production reliability and enterprise customer implementations for Kong's fast-growing Managed Gateways SaaS product across AWS, GCP, and Azure. Requires deep Kubernetes, cloud-native, and Golang expertise plus customer-facing technical leadership.

118k – 167k/yr
Remote7+ YOEDevOps / SRE

About the role

What You’ll Do

Platform & Reliability Engineering

  • Lead, mentor, and inspire a high-performing team of Site Reliability Engineers dedicated to Kong's Managed Gateway offerings.
  • Architect and implement robust, scalable, and fault-tolerant cloud-native systems using technologies like Kubernetes, Golang, and major cloud providers.
  • Own the end-to-end operational lifecycle, from proactive monitoring and alerting to incident response and blameless post-mortems, ensuring continuous service improvement.
  • Drive a culture of developer delight by implementing automation, self-service tooling, and streamlined workflows for deploying and managing API gateways.
  • Define, track, and report on key SLOs and SLIs to ensure optimal performance and reliability of Managed Gateways.
  • Champion technical debt prevention and advocate for architectural best practices that enhance system resilience and reduce operational toil.
  • Collaborate cross-functionally with Product, engineering, and Customer Success to influence roadmap decisions and ensure operational readiness for new features.

Enterprise Implementation Engineering

  • Partner directly with enterprise customers — working alongside Product leadership, Professional Services, and Customer Success — to drive end-to-end onboarding and implementation of Cloud Gateways, and productize recurring implementation patterns into repeatable playbooks and platform capabilities.
  • Bring deep, cross-cloud breadth (AWS, GCP, Azure) to handle unique customer topologies and turn complex setups into successful, production-ready deployments.
  • Serve as the technical owner of the customer relationship through implementation, primarily supporting our North America customer base, and be the escalation point Customer Success leans on for technically complex accounts.
  • Feed real-world implementation patterns and customer constraints back to Product to further contribute the roadmap.

What You’ll Bring

The Toolkit

  • Extensive experience as a Site Reliability Engineer, focusing on highly available and distributed systems.
  • Deep expertise with Kubernetes and cloud-native architectures, preferably across multiple public cloud providers (AWS, GCP, Azure).
  • Strong proficiency in Golang or similar modern programming languages for automation and tool development.
  • Proven track record in building and maintaining CI/CD pipelines and infrastructure as code (Terraform, Ansible).
  • In-depth knowledge of monitoring, logging, and alerting systems (e.g., Prometheus, Grafana, ELK stack, Datadog).
  • Experience with managed services, API gateways, or similar network infrastructure is highly desirable.

The Kong DNA

  • You take immense ownership of your systems, treating reliability as a first-class feature.
  • You operate with a sense of urgency, especially in critical situations, and drive quick, effective resolutions.
  • You thrive in a collaborative environment, actively sharing knowledge and elevating the entire team.
  • Kong moves fast, and our team's spread across continents and time zones — plans shift mid-flight, and things don't always line up neatly. You don't need everything settled to do good work. You bring your own calm to the noise, figure things out as you go, and help the people around you do the same.

Bonus Points

  • Experience with Service Mesh technologies (e.g., Istio, Linkerd).
  • Familiarity with database administration for high-throughput systems (PostgreSQL, Cassandra).
  • Contributions to open-source SRE tools or projects.
  • Relevant cloud certifications (e.g., AWS Certified DevOps Engineer, CKA).

Skills

KubernetesGoAWSGCPAzureTerraformAnsiblePrometheusGrafanaelk stackDatadogCI/CDservice meshistioPostgres

Similar roles

DevOps / SRE jobs
Shield AI

Senior Engineer, Platform Infrastructure

Shield AISan Diego, CA +2

Build and evolve the infrastructure platform that deploys and operates customer environments. The role focuses on Kubernetes, infrastructure as code, deployment automation, observability, reliability, security, and collaborative continuous delivery practices.

120k – 180k/yrOn-site5+ YOEDevOps / SRE
PrizePicks

Senior Site Reliability Engineer

PrizePicksUnited States

Senior Site Reliability Engineer responsible for designing, operating, and improving reliable, scalable production systems. The role requires 5+ years of reliability-focused engineering experience plus expertise in cloud platforms, infrastructure as code, Kubernetes, programming, observability, and critical incident response.

120k – 175k/yrRemote5+ YOEDevOps / SRE
Shield AI

Senior Engineer, Software Engineering Tools (R4913)

Shield AIDallas, TX

Develops and maintains internal software tools to accelerate engineering workflows for cutting-edge aircraft, integrating EDA, CAD, and PLM systems. Requires 5+ years experience with Python, C++, JavaScript, SQL, Docker, CI/CD, and cloud platforms.

120k – 190k/yrOn-site5+ YOEDevOps / SRE
Bland AI

Senior Infrastructure Engineer

Bland AISan Francisco, CA

Builds and scales distributed systems for real-time voice processing, ML inference, and telephony integration using Kubernetes. Requires 5+ years experience with cloud infrastructure, real-time systems, and tools like Terraform and Datadog.

120k – 200k/yrOn-site5+ YOEDevOps / SRE
LiveKit

Senior Infrastructure Engineer

LiveKitUnited States

Builds and owns foundational infrastructure for globally distributed systems, implements SRE objectives in Golang, manages Kubernetes clusters, and leads incident response. Requires expertise in software engineering, systems administration, and multi-region operations.

120k – 250k/yrRemoteDevOps / SRE