Operates and automates production infrastructure for a high-scale AI inference service. The role requires production Kubernetes experience, Python or Go proficiency, observability expertise, and a focus on reliability, automation, and reducing operational toil.
Salary not listed
On-site5+ YOEDevOps / SRE
About the role
Responsibilities
Execute releases, capacity changes, and cluster upgrades while robust continuous delivery pipelines and self-service capabilities are developed.
Contribute to self-service continuous delivery pipelines using Kubernetes, Bazel, Prometheus, Grafana, InfluxDB, Python, and Go.
Build reusable automation and internal developer tools to minimize operational toil and cross-team friction.
Develop and extend telemetry, observability, and alerting solutions for reliability at scale.
Collaborate with Cluster Operations and development teams to identify high-impact automation opportunities.
Contribute to reliability practices, including SLOs, post-mortems, and capacity planning.
Requirements
2–4+ years of experience in site reliability engineering with a strong operations or automation focus.
Production Kubernetes experience.
Strong Python or Go skills for building tools and automation.
Proficiency with Prometheus, Grafana, and observability-driven workflows.
Ability to measure and communicate impact through reliability metrics, operational toil, and velocity gains.
Nice to Have
Hands-on GitOps expertise with Argo CD, Flux, or equivalent tools.
Experience building continuous delivery pipelines.
Experience with Bazel or similar build systems.
Familiarity with capacity planning and on-premises or multi-datacenter environments.
Compensation and Benefits
The role does not require 24/7 on-call rotations.
Immediate ownership of production systems, direct mentorship from experienced engineers, and collaboration with Staff SREs.
Cerebras offers a non-corporate work culture, startup vitality, opportunities to work on a breakthrough AI platform and AI supercomputer, and support for continuous learning, growth, and inclusion.
Build and operate secure, highly available cloud infrastructure for mission-critical government and space systems across AWS GovCloud and C2E environments. The role requires Kubernetes, Terraform, Python, observability, networking, compliance, and an active U.S. security clearance.
160k – 200k/yrHybrid3+ YOEDevOps / SRE
Network Engineer
OpenAISan Francisco, CA
Designs, operates, and improves secure enterprise networks spanning offices, campuses, cloud environments, and connectivity services. The role combines architecture, production operations, troubleshooting, observability, security, and infrastructure automation.
293k – 385k/yrHybridDevOps / SRE
Build & Release Engineer
Applied IntuitionSunnyvale, CA
Owns software release workflows, dependency updates, artifact management, CI/CD pipelines, and an internal release portal. The role requires at least three years of software development experience, strong coding skills, and hands-on expertise with Git, CI/CD, and artifact repositories.
Build secure, scalable infrastructure, data systems, compute tooling, and developer experiences for Anthropic’s Interpretability research team. The role partners closely with researchers, security, and platform teams and requires strong programming and infrastructure experience.
320k – 485k/yrHybridDevOps / SRE
Software Engineer - Continuous Delivery
BasetenSan Francisco, CA +1
Build and operate continuous delivery infrastructure for Kubernetes deployments across global regions, including progressive rollouts, automated health evaluation, and rollback systems. The role requires strong Go or Python skills, large-scale Kubernetes experience, and familiarity with GitOps tooling.