Skip to content
Cerebras SystemsCerebras SystemsSunnyvale, CA

Cluster Operations Software Engineer

Operates and develops software for large-scale AI compute clusters, improving reliability, capacity, monitoring, and incident response. Requires 6–8 years of complex infrastructure experience, strong Python and Go skills, distributed-systems expertise, and participation in 24/7 on-call support.

Salary not listed
Hybrid6+ YOEDevOps / SRE

About the role

Responsibilities

  • Deploy, configure, and debug container-based services using Docker.
  • Build and own software solutions for cluster operations, including monitoring platforms, workflow automation systems, operational dashboards, and reliability tooling.
  • Collaborate with cross-functional teams to translate operational requirements into scalable operations and maintenance products and platform capabilities.
  • Develop APIs, automation services, and integrations that improve operational visibility, incident response, and fleet management across global AI infrastructure.
  • Manage and operate multiple advanced AI compute infrastructure clusters.
  • Monitor and oversee cluster health, proactively identifying and resolving potential issues.
  • Maximize compute capacity through optimization and efficient resource allocation.
  • Provide 24/7 monitoring and support using automated tools and hands-on troubleshooting.
  • Handle engineering escalations and collaborate with other teams to resolve complex technical challenges.
  • Stay current with advancements in AI compute infrastructure and related technologies.

Requirements

  • 6–8 years of relevant experience managing and operating complex compute infrastructure, preferably in machine learning or high-performance computing.
  • Proficiency in Python and Go, including experience building operational platforms, workflow automation systems, and reliability tooling for large-scale infrastructure environments.
  • Expertise in distributed systems.
  • Deep understanding of Linux-based compute systems and command-line tools.
  • Extensive knowledge of Docker containers and orchestration platforms such as Kubernetes.
  • Ability to troubleshoot and resolve complex technical issues efficiently.
  • Experience with monitoring and alerting systems.
  • Proven ability to own challenges and drive them to completion.
  • Strong communication and collaboration skills.
  • Ability to work effectively in a fast-paced environment.
  • Willingness to participate in a 24/7 on-call rotation.

Preferred Skills

  • Experience operating and managing large-scale AI clusters.
  • Knowledge of Ethernet, RoCE, and TCP/IP.
  • Knowledge of cloud computing platforms such as AWS, Google Cloud, and Azure.

Benefits

  • Opportunity to build and operate a breakthrough AI platform beyond GPU constraints.
  • Opportunities to publish and open-source AI research.
  • Work on one of the world's fastest AI supercomputers.
  • Startup vitality with job stability.
  • A non-corporate work culture that respects individual beliefs.

Skills

PythonGoDistributed SystemsLinuxDockerKubernetesmonitoring and alertingethernetroceTCP/IPAWSGCPAzure

Similar roles

DevOps / SRE jobs
Snowflake

Senior Software Engineer - Snowpark Container Service

SnowflakeBellevue, WA +1

Senior engineer to design, build, and lead development of Snowpark Container Services, a Kubernetes-based container compute platform. Requires 7+ years building large-scale distributed systems and strong coding skills in Java, C++, or Go.

200k – 288k/yrHybrid7+ YOEDevOps / SRE
Astronomer

Senior Software Engineer, Infrastructure & Systems

AstronomerNew York, NY

Designs and operates control-plane systems that provision, scale, secure, and observe infrastructure running Airflow across multi-tenant and private-cloud environments. Requires 5+ years in infrastructure or systems engineering, strong Kubernetes and API expertise, and proficiency in Go or TypeScript.

200k – 300k/yrHybrid5+ YOEDevOps / SRE
Okta

Senior Manager, Site Reliability Engineering - Infrastructure Platform

OktaBellevue, WA

Leads Infrastructure Platform and Shared Services teams, overseeing Edge networking, Kubernetes platform, CI/CD, observability, and automation. Requires 6+ years technical leadership, AWS expertise, and strong Kubernetes/Terraform skills.

176k – 264k/yrHybrid6+ YOEDevOps / SRE
Clickhouse

Senior Cloud Software Engineer - Efficiency Engineering

ClickhouseUnited States

Build and operate scalable, fault-tolerant cloud infrastructure while leading efficiency initiatives across compute, storage, and networking. The role requires 5+ years of distributed-systems software development experience and expertise with public cloud, infrastructure as code, and cloud-native technologies.

133k – 232k/yrRemote5+ YOEDevOps / SRE
Temporal

Senior Software Engineer, Infrastructure Foundations

TemporalUnited States

Build and scale reliable cloud infrastructure systems, shape long-term architecture and roadmaps, and drive cross-functional alignment. The role requires 10+ years of coding experience, distributed-systems and concurrency expertise, deep infrastructure experience, and hands-on cloud-provider experience.

176k – 238k/yrRemote10+ YOEDevOps / SRE