Leads and builds a Managed Gateways Site Reliability Engineering team, initially contributing hands-on to enterprise implementations and reliability initiatives. The role requires engineering management experience, cloud-native distributed-systems expertise, Kubernetes, Golang, and strong observability and incident-management capabilities.
153k – 218k/yr
Remote5+ YOEEngineering Management
About the role
Responsibilities
Build Kong's Managed Gateways SRE team in Seattle, Washington, from the ground up by hiring, setting standards, and remaining hands-on with enterprise implementations as the team ramps.
Contribute directly to critical implementations and reliability work during the team's early stages.
Lead, mentor, and grow a high-performing team of Site Reliability Engineers supporting Managed Gateway offerings and enterprise customers across the Americas and Europe.
Architect and implement robust, scalable, fault-tolerant cloud-native systems using Kubernetes, Golang, and major cloud providers.
Own the operational lifecycle, including monitoring, alerting, incident response, blameless post-mortems, and continuous service improvement.
Improve developer experience and operational efficiency through automation, self-service tooling, and streamlined API gateway deployment and management workflows.
Define, track, and report on SLOs and SLIs.
Advocate for architectural practices that maintain performance and resilience as Managed Gateways scales.
Collaborate with Product, engineering, and support teams on roadmap decisions and operational readiness.
Requirements
Experience leading and managing Site Reliability Engineering or DevOps teams in a fast-paced, high-growth environment.
Expertise designing, deploying, and operating highly available distributed systems on AWS, Azure, or Google Cloud.
Extensive production experience with Kubernetes and container orchestration.
Proficiency in Golang or similar modern programming languages for infrastructure automation and service development.
Strong understanding of observability principles and tools such as Prometheus, Grafana, and OpenTelemetry.
Experience managing critical incidents, performing root cause analysis, and implementing preventative measures.
Nice-to-haves
Familiarity with API gateway technologies, service mesh, or network proxies.
Open-source contributions or participation in SRE/cloud-native communities.
Cloud or Kubernetes certifications, such as AWS Certified DevOps Engineer, CKA, or CKAD.
Experience at companies focused on developer tools, infrastructure software, or API management.
Leads a team of mechanical engineers developing systems for next-generation autonomous UAVs, supports design reviews, resource allocation, and cross-site collaboration while providing technical guidance and mentorship. Requires 7+ years in mechanical design, BS in Mechanical/Aerospace Engineering, and strong CAD/FEA skills.
153k – 229k/yrOn-site7+ YOEEngineering Management
Engineering Manager, Data Foundations
GitLabUnited States
Manage and grow a high-performing engineering team building GitLab's Data Insights Platform and classic search capabilities. Drive architecture for high-throughput distributed data systems across multiple deployment models while hiring, coaching, and delivering roadmap outcomes.
Lead engineering teams in the AXIS group to build observability, metrics, and AI-driven tools that measure and improve developer productivity and the software development lifecycle at MongoDB. Requires 8+ years software engineering experience and 4+ years managing engineers.
151k – 297k/yrRemote8+ YOEEngineering Management
Customer Support Engineering Manager
CrusoeSan Francisco, CA
Leads a team of support engineers resolving complex technical issues for enterprise AI/ML customers on Crusoe Cloud. Requires 6+ years technical experience in software/devops/support, 2+ years leadership, and deep cloud/GPU knowledge.
155k – 185k/yrOn-site6+ YOEEngineering Management
Engineering Manager, DevOps
Loop ReturnsColumbus, OH
Lead a team of DevOps and SRE engineers to own infrastructure, CI/CD pipelines, and progressive delivery on AWS and Kubernetes. Hands-on role (50% time) building self-service platforms, multi-region architectures, SLOs, and observability with Datadog while mentoring the team and partnering with product and data organizations. Requires 7+ years infrastructure experience and 2+ years leading DevOps/SRE teams.