Skip to content

Distributed Software Engineer

Build and operate distributed infrastructure software that automates, schedules, observes, and repairs large-scale AI compute clusters. The role requires 5+ years of infrastructure or distributed-systems experience, strong Go and Python skills, and deep Kubernetes expertise.

About the job

Responsibilities

  • Build declarative, CRD-driven automation for bare-metal networking, operating systems, and application software across clusters of Cerebras systems, servers, and switches.
  • Develop push-button cluster installation, upgrades, and security patching with canary gates and defined downtime budgets.
  • Build Kubernetes operators for large-scale inference workload scheduling, including resource locks, priority queues, network topology, and health-aware placement.
  • Develop gRPC control-plane services, authorization, admission webhooks, and quota policies for a multi-tenant fleet.
  • Build metrics and log pipelines with exporters for wafer-scale systems, servers, and network fabrics using Prometheus and Grafana.
  • Implement failure detection, highly available control planes, automated recovery, and CLIs, APIs, and MCP gateways for fleet management.

Requirements

  • 5+ years of experience building and operating production distributed systems or infrastructure software.
  • Production-quality Go and Python programming skills.
  • Deep Kubernetes experience, including controllers, operators, CRDs, reconciliation semantics, informer caches, admission webhooks, and RBAC.
  • Strong debugging skills across distributed systems, Linux, and networking.
  • Practical Prometheus and Grafana experience, including PromQL, exporter design, cardinality management, alerting, and SLOs.
  • Ability to work independently in a fast-moving, incompletely documented environment and drive cross-team initiatives.
  • Demonstrated use of AI tools in software engineering, with the ability to evaluate and verify generated output.

Nice to Have

  • Bare-metal or HPC fleet operations.
  • Scheduler internals.
  • RDMA/RoCE and eBPF networking.
  • Ceph or NVMe-oF.
  • etcd and high-availability upgrades.
  • Inference-serving stacks.

Skills

Go, Python, Kubernetes, Custom Resource Definitions, Kubernetes Operators, gRPC, Linux, Networking, Prometheus, Grafana, Promql, RBAC, Ebpf, Ceph, Rdma

Granica

Granica

Remote

Software Engineer, Infrastructure
No salary listedRemote5+ YOEDevOps / SRE

Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.

Baseten

Baseten

San Francisco, CA

Capacity Ops Engineer
$170k+/yrHybrid5+ YOEDevOps / SRE

Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.

Teleport

Teleport

United States

IT Security and Automation Engineer
$149k+/yrRemoteDevOps / SRE

Build IT workflow automation and security tooling using Go, Temporal, Kubernetes, Terraform, and shell scripting. The role also supports internal IT systems and integrations across endpoint management, identity, access, and security administration.

Crusoe

Crusoe

United States

Electrical Field Engineer - Data Center
$196k+/yrRemote5+ YOEDevOps / SRE

Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.

Beacon AI

Beacon AI

San Carlos, CA

Software Engineer, Cloud Infrastructure
$135k+/yrHybridDevOps / SRE

Build and operate AWS cloud, ML, LLM, RAG, and IoT infrastructure, including deployment platforms, data pipelines, vector search, observability, security, and cost controls. The role requires deep AWS experience and production experience with LLM-powered applications.