# Software Engineer

**Company:** [xAI](https://hotfix.jobs/companies/xai)
**Location:** Palo Alto, CA
**Role:** DevOps / SRE
**Salary:** $180k – $440k/yr
**Experience:** 5+ years
**Skills:** Rust, C++, C, Kubernetes, Linux, Docker, Prometheus, Grafana, TCP/IP, Helm
**Posted:** 2026-07-20

> Build and optimize large-scale distributed systems powering xAI's massive supercomputing clusters for AI training. Requires strong systems programming in Rust/C++ and deep Kubernetes/Linux expertise.

## Job Description

## Responsibilities
- Design, build, and implement a large-scale distributed system that powers one of the world's largest supercomputing clusters.
- Dive into the low-level stack to profile, debug, and optimize performance across diverse systems, including GPUs, Linux kernel, networking, and filesystems, to achieve peak efficiency.
- Collaborate on hardware, software, and algorithm co-design to push the boundaries of AI training.
- Maintain and innovate on our codebase to ensure scalability and reliability.
- Develop tools to enhance team productivity and streamline workflows.

## Basic Qualifications
- Systems programming experience in C, C++, or Rust.
- Computer systems fundamentals with a grasp of how computers execute code from transistors to high-level applications.
- Hands-on expertise with Kubernetes (K8s), including cluster architecture, pod lifecycle, networking (CNI), storage (CSI), service mesh, and production-grade operations.

## Preferred Skills and Experience
- Collaborate in a fast-paced, open environment to design foundational systems.
- Strong debugging skills across the full stack — from kernel and OS up through container orchestration layers.
- Deep knowledge of operating systems internals (process scheduling, memory management, file systems, and synchronization primitives).
- Proficiency in performance analysis, profiling, and low-level optimization techniques.
- Solid understanding of computer networks and the TCP/IP stack.
- Experience working with Linux kernel concepts or systems-level debugging tools (e.g., perf, gdb, strace, Wireshark).
- Proficiency deploying and managing workloads using Kubernetes manifests, Helm, Operators, and GitOps workflows.
- Solid understanding of containerization technologies (Docker, containerd, crio) and their interaction with the Linux kernel.
- Experience with observability and monitoring in distributed systems (Prometheus, Grafana, VictoriaMetrics, OpenTelemetry, or similar).

## Compensation and Benefits
- $180,000 - $440,000 USD total compensation (base salary is just one part of our total rewards package at xAI, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks).

## Similar roles

- [AI Infrastructure Engineer, Sandbox Platform](https://hotfix.jobs/jobs/4235c3a2-e042-48e4-8e5e-d6872fd48ad9) - Scale AI - San Francisco, CA - $180k – $225k/yr
- [Platform Engineer](https://hotfix.jobs/jobs/4574139f-8958-4f47-b5ba-9a5414e41802) - Unusual - New York, NY - $180k – $250k/yr
- [Platform Engineer](https://hotfix.jobs/jobs/14231c8e-abad-4a41-94e3-d41fb2ede6f5) - Govsignals - New York, NY - $180k – $250k/yr
- [Network Engineer, Design & Engineering](https://hotfix.jobs/jobs/48ff3694-5232-4d66-b43c-a091b96c8902) - Fluidstack - New York, NY - $180k – $300k/yr
- [Developer Productivity Engineer](https://hotfix.jobs/jobs/9a60db26-80d5-4b2a-b613-56ae58d49794) - Hightouch - Remote - $180k – $320k/yr

**Apply:** https://hotfix.jobs/jobs/24949960-ef02-4961-8eb2-14082cd46e85
**Canonical:** https://hotfix.jobs/jobs/24949960-ef02-4961-8eb2-14082cd46e85