Skip to content
Cerebras SystemsCerebras SystemsUnited States

AI Inference Core - Senior SW Engineer for Platform & DevOps

Build and operate the platform layer behind Cerebras engineering infrastructure, including CI/CD, Kubernetes, deployment automation, cloud and on-premises systems, developer environments, and observability. The role requires 5+ years of infrastructure or software engineering experience and strong debugging and systems fundamentals.

Salary not listed
Hybrid5+ YOEDevOps / SRE

About the role

Responsibilities

  • Design, build, and maintain CI/CD systems supporting build, test, integration, qualification, and release workflows.
  • Build and operate Kubernetes-based platforms and services used by engineering teams.
  • Develop deployment systems, internal tools, and self-service workflows that make infrastructure changes repeatable, reviewable, and safe.
  • Improve infrastructure reliability, capacity, performance, cost efficiency, monitoring, and operational readiness.
  • Debug issues spanning CI pipelines, Kubernetes workloads, networking, storage, authentication, operating systems, and distributed applications.
  • Perform root-cause analysis and implement lasting fixes rather than relying on repeated manual intervention.
  • Partner with software, IT, security, networking, release, and developer-productivity teams to deliver scalable infrastructure solutions.

Requirements

  • 5+ years of professional experience in platform engineering, DevOps, infrastructure engineering, site reliability engineering, or software engineering.
  • Hands-on experience building or maintaining CI/CD pipelines and automated software-delivery workflows.
  • Experience deploying and operating services using Kubernetes and containerized environments.
  • Experience with a major cloud platform, preferably AWS, and programmatic infrastructure provisioning.
  • Strong understanding of Linux or Unix operating-system fundamentals.
  • Understanding of networking concepts such as DNS, routing, load balancing, proxies, ports, TLS, and service connectivity.
  • Proficiency in Python, Shell, or another language used to build infrastructure automation and operational tooling.
  • Experience with monitoring, logging, alerting, dashboards, and incident investigation.
  • Strong debugging and problem-solving skills across applications, infrastructure, networking, and operating systems.

Nice-to-haves

  • Experience with infrastructure-as-code tools, specifically Terraform.
  • Experience with Kubernetes controllers, operators, custom resources, Helm, Argo CD, or similar platform technologies.
  • Experience managing artifact repositories, package registries, build caches, or software-distribution infrastructure.
  • Familiarity with build systems, dependency management, and reproducible-build practices.
  • Experience supporting hybrid environments spanning cloud infrastructure, on-premises systems, and specialized hardware.
  • Experience with identity and access management, secrets, certificates, TLS, or mTLS.
  • Experience building internal developer platforms or self-service infrastructure products.
  • BS/MS in Computer Science or a related field, or equivalent practical experience.

Skills

PythonshellKubernetesAWSTerraformCI/CDDockerLinuxHelmargo cdDNStlsMonitoringloggingInfrastructure As Code

Similar roles

DevOps / SRE jobs
Snowflake

Senior Software Engineer - Snowpark Container Service

SnowflakeBellevue, WA +1

Senior engineer to design, build, and lead development of Snowpark Container Services, a Kubernetes-based container compute platform. Requires 7+ years building large-scale distributed systems and strong coding skills in Java, C++, or Go.

200k – 288k/yrHybrid7+ YOEDevOps / SRE
Astronomer

Senior Software Engineer, Infrastructure & Systems

AstronomerNew York, NY

Designs and operates control-plane systems that provision, scale, secure, and observe infrastructure running Airflow across multi-tenant and private-cloud environments. Requires 5+ years in infrastructure or systems engineering, strong Kubernetes and API expertise, and proficiency in Go or TypeScript.

200k – 300k/yrHybrid5+ YOEDevOps / SRE
Cerebras Systems

Cluster Operations Software Engineer

Cerebras SystemsSunnyvale, CA

Operates and develops software for large-scale AI compute clusters, improving reliability, capacity, monitoring, and incident response. Requires 6–8 years of complex infrastructure experience, strong Python and Go skills, distributed-systems expertise, and participation in 24/7 on-call support.

Salary not listedHybrid6+ YOEDevOps / SRE
Okta

Senior Manager, Site Reliability Engineering - Infrastructure Platform

OktaBellevue, WA

Leads Infrastructure Platform and Shared Services teams, overseeing Edge networking, Kubernetes platform, CI/CD, observability, and automation. Requires 6+ years technical leadership, AWS expertise, and strong Kubernetes/Terraform skills.

176k – 264k/yrHybrid6+ YOEDevOps / SRE
Clickhouse

Senior Cloud Software Engineer - Efficiency Engineering

ClickhouseUnited States

Build and operate scalable, fault-tolerant cloud infrastructure while leading efficiency initiatives across compute, storage, and networking. The role requires 5+ years of distributed-systems software development experience and expertise with public cloud, infrastructure as code, and cloud-native technologies.

133k – 232k/yrRemote5+ YOEDevOps / SRE