DevOps Engineer
Build and operate the Kubernetes-based platform infrastructure, developer tooling, CI/CD systems, and agent infrastructure used across the engineering organization. The role requires production cloud experience, strong infrastructure-as-code skills, Python proficiency, and sound engineering judgment.
About the job
Responsibilities
- Own platform services end to end, including design, build, operation, improvement, and on-call support.
- Deliver infrastructure as code using Terraform, Helm, and ArgoCD, with GitOps as the default operating model.
- Build and maintain developer-experience tooling, including onboarding templates, CI/CD pipelines, local development environments, and self-service capabilities.
- Contribute to agent infrastructure, including sandboxed execution environments, credential brokering, guardrails, and internal MCP servers and skills.
- Improve observability and alerting so system health is visible and actionable.
- Participate in incident response and post-incident reviews, treating failures as systemic issues.
Requirements
- Hands-on experience running production Kubernetes with real users and consequences on GCP, Azure, or AWS.
- Working knowledge of Terraform, Helm, and GitOps delivery patterns such as ArgoCD or Flux.
- Proficiency in Python and ability to use other programming languages when needed.
- Demonstrable mastery of agentic coding tools, including scoping work, building effective context and tooling, and identifying incorrect output before production.
- Strong engineering judgment and ability to explain the rationale and limitations of practices such as immutable infrastructure, trunk-based development, and SLOs.
- Strong written communication for runbooks, design documents, and code reviews.
- Comfort operating in ambiguity and taking initiative.
Compensation and Benefits
- Compensation is adjusted for regional market conditions and cost of living.
- Offers include bonus and equity.
- Final compensation depends on location, experience, skills, knowledge, internal pay equity, and market conditions.
- Benefits and total compensation details are discussed during the hiring process.
Skills
Kubernetes, Terraform, Helm, Argo CD, Flux, GitOps, Python, GCP, Microsoft Azure, AWS, CI/CD, Observability, Mcp
Similar jobs
DevOps / SRE jobsBuild and maintain Cloudflare’s deployment platform, enabling progressive rollouts, health-mediated releases, and automated workflows at scale. The role requires at least four years of software development experience, backend and frontend experience, and comfort with rapid delivery and on-call support.
Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Site Reliability Engineers build and operate reliable, scalable production infrastructure across GitLab’s Infrastructure Platforms teams. The role requires strong software engineering and operations fundamentals, Kubernetes and infrastructure-as-code experience, cloud expertise, and comfort with automation, observability, and incident response.
Build and operate distributed infrastructure across compute, storage, networking, data, deployment, and reliability domains. The role requires 4+ years of backend or platform engineering experience, strong systems-language skills, and the ability to own complex production systems and lead cross-team technical initiatives.
Infrastructure engineer responsible for building and operating highly available cloud systems, automating operations, and improving reliability across a large-scale AI platform. Requires 5+ years of infrastructure or DevOps experience, production Kubernetes, cloud infrastructure, Terraform, and Python or Go.