Skip to content
xAIxAI

Software Engineer

Build and optimize large-scale distributed systems powering xAI's massive supercomputing clusters for AI training. Requires strong systems programming in Rust/C++ and deep Kubernetes/Linux expertise.

About the job

Responsibilities

  • Design, build, and implement a large-scale distributed system that powers one of the world's largest supercomputing clusters.
  • Dive into the low-level stack to profile, debug, and optimize performance across diverse systems, including GPUs, Linux kernel, networking, and filesystems, to achieve peak efficiency.
  • Collaborate on hardware, software, and algorithm co-design to push the boundaries of AI training.
  • Maintain and innovate on our codebase to ensure scalability and reliability.
  • Develop tools to enhance team productivity and streamline workflows.

Basic Qualifications

  • Systems programming experience in C, C++, or Rust.
  • Computer systems fundamentals with a grasp of how computers execute code from transistors to high-level applications.
  • Hands-on expertise with Kubernetes (K8s), including cluster architecture, pod lifecycle, networking (CNI), storage (CSI), service mesh, and production-grade operations.

Preferred Skills and Experience

  • Collaborate in a fast-paced, open environment to design foundational systems.
  • Strong debugging skills across the full stack — from kernel and OS up through container orchestration layers.
  • Deep knowledge of operating systems internals (process scheduling, memory management, file systems, and synchronization primitives).
  • Proficiency in performance analysis, profiling, and low-level optimization techniques.
  • Solid understanding of computer networks and the TCP/IP stack.
  • Experience working with Linux kernel concepts or systems-level debugging tools (e.g., perf, gdb, strace, Wireshark).
  • Proficiency deploying and managing workloads using Kubernetes manifests, Helm, Operators, and GitOps workflows.
  • Solid understanding of containerization technologies (Docker, containerd, crio) and their interaction with the Linux kernel.
  • Experience with observability and monitoring in distributed systems (Prometheus, Grafana, VictoriaMetrics, OpenTelemetry, or similar).

Compensation and Benefits

  • $180,000 - $440,000 USD total compensation (base salary is just one part of our total rewards package at xAI, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks).

Skills

Rust, C++, C, Kubernetes, Linux, Docker, Prometheus, Grafana, TCP/IP, Helm

Benchling

Benchling

San Francisco, CA
Software Engineer, Platform
$173k+/yrHybrid4+ YOEDevOps / SRE

Build developer-experience tooling and release systems within Benchling’s Platform team, helping engineering teams develop, test, package, and ship high-quality software rapidly. The role requires 4+ years of software engineering experience, web framework expertise, strong problem-solving, and effective cross-functional communication.

Baseten

Baseten

San Francisco, CA

Capacity Ops Engineer
$170k+/yrHybrid5+ YOEDevOps / SRE

Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.

Mercor

Mercor

San Francisco, CA

Cloud Platform Engineer
$190k+/yrOn-siteDevOps / SRE

Own Mercor’s internal identity and cloud platform infrastructure as code, automating provisioning, access management, secrets, and employee lifecycle workflows. The role requires production Terraform, Okta, SCIM, and multi-cloud IAM experience, plus strong automation, incident response, and documentation skills.

Ramp

Ramp

New York, NY
TLM, Production Engineering
$168k+/yrHybrid3+ YOEDevOps / SRE

Production Engineer responsible for building and operating scalable infrastructure, driving reliability and architecture initiatives, and enabling product teams through platform tooling. Requires software engineering experience, distributed-systems expertise, cloud experience, and cross-team technical leadership.

Roboflow

Roboflow

New York, NY
Infrastructure Engineer
$165k+/yrRemoteDevOps / SRE

Infrastructure Engineer responsible for securing, scaling, and operating cloud infrastructure and machine-learning platforms across a distributed startup. The role requires Kubernetes, infrastructure-as-code, cloud operations, CI/CD, programming, observability, and security experience.