Skip to content
xAIxAIPalo Alto, CA

Software Engineer

Build and optimize large-scale distributed systems powering xAI's massive supercomputing clusters for AI training. Requires strong systems programming in Rust/C++ and deep Kubernetes/Linux expertise.

180k – 440k/yr
On-site5+ YOEDevOps / SRE

About the role

Responsibilities

  • Design, build, and implement a large-scale distributed system that powers one of the world's largest supercomputing clusters.
  • Dive into the low-level stack to profile, debug, and optimize performance across diverse systems, including GPUs, Linux kernel, networking, and filesystems, to achieve peak efficiency.
  • Collaborate on hardware, software, and algorithm co-design to push the boundaries of AI training.
  • Maintain and innovate on our codebase to ensure scalability and reliability.
  • Develop tools to enhance team productivity and streamline workflows.

Basic Qualifications

  • Systems programming experience in C, C++, or Rust.
  • Computer systems fundamentals with a grasp of how computers execute code from transistors to high-level applications.
  • Hands-on expertise with Kubernetes (K8s), including cluster architecture, pod lifecycle, networking (CNI), storage (CSI), service mesh, and production-grade operations.

Preferred Skills and Experience

  • Collaborate in a fast-paced, open environment to design foundational systems.
  • Strong debugging skills across the full stack — from kernel and OS up through container orchestration layers.
  • Deep knowledge of operating systems internals (process scheduling, memory management, file systems, and synchronization primitives).
  • Proficiency in performance analysis, profiling, and low-level optimization techniques.
  • Solid understanding of computer networks and the TCP/IP stack.
  • Experience working with Linux kernel concepts or systems-level debugging tools (e.g., perf, gdb, strace, Wireshark).
  • Proficiency deploying and managing workloads using Kubernetes manifests, Helm, Operators, and GitOps workflows.
  • Solid understanding of containerization technologies (Docker, containerd, crio) and their interaction with the Linux kernel.
  • Experience with observability and monitoring in distributed systems (Prometheus, Grafana, VictoriaMetrics, OpenTelemetry, or similar).

Compensation and Benefits

  • $180,000 - $440,000 USD total compensation (base salary is just one part of our total rewards package at xAI, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks).

Skills

RustC++CKubernetesLinuxDockerPrometheusGrafanaTCP/IPHelm

Similar roles

DevOps / SRE jobs
Scale AI

AI Infrastructure Engineer, Sandbox Platform

Scale AISan Francisco, CA +2

Build and evolve a secure, high-performance agent sandboxing platform for code execution. Combine deep systems expertise in isolation/virtualization with strong focus on developer experience, APIs, and internal partnerships. Requires 4+ years in high-performance systems software.

180k – 225k/yr
Hybrid4+ YOEDevOps / SRE
Unusual

Platform Engineer

UnusualNew York, NY

Build and scale data ingestion, storage, search, and deployment infrastructure across commercial, FedRAMP High, and IL5 environments. Own platform used by engineers and AI agents for government contracting data, with customer-facing work on integrations and security.

180k – 250k/yr
Hybrid5+ YOEDevOps / SRE
Govsignals

Platform Engineer

GovsignalsNew York, NY

Platform Engineer building and scaling data ingestion, storage, search, and deployment infrastructure across commercial, FedRAMP High, and IL5 environments. Requires 5+ years in backend/platform/infra roles with deep experience in Postgres, Kubernetes, and TypeScript/Python.

180k – 250k/yr
Hybrid5+ YOEDevOps / SRE
Fluidstack

Network Engineer, Design & Engineering

FluidstackNew York, NY +4

Design end-to-end datacenter network architectures for AI training and inference workloads. Own topology selection, fabric design, physical infrastructure integration, and produce deployable HLDs/LLDs across multiple GPU platforms and customer requirements.

180k – 300k/yr
On-site5+ YOEDevOps / SRE
Hightouch

Developer Productivity Engineer

HightouchUnited States

As a Senior Developer Productivity Engineer, you will own the build, test, and deployment processes for a 50+ person engineering team. You will improve monorepo productivity, drive excellence in testing, and support multi-cloud/multi-region infrastructure to enable fast and safe shipping.

180k – 320k/yr
Remote5+ YOEDevOps / SRE