Software Engineer
Build and optimize large-scale distributed systems powering xAI's massive supercomputing clusters for AI training. Requires strong systems programming in Rust/C++ and deep Kubernetes/Linux expertise.
About the job
Responsibilities
- Design, build, and implement a large-scale distributed system that powers one of the world's largest supercomputing clusters.
- Dive into the low-level stack to profile, debug, and optimize performance across diverse systems, including GPUs, Linux kernel, networking, and filesystems, to achieve peak efficiency.
- Collaborate on hardware, software, and algorithm co-design to push the boundaries of AI training.
- Maintain and innovate on our codebase to ensure scalability and reliability.
- Develop tools to enhance team productivity and streamline workflows.
Basic Qualifications
- Systems programming experience in C, C++, or Rust.
- Computer systems fundamentals with a grasp of how computers execute code from transistors to high-level applications.
- Hands-on expertise with Kubernetes (K8s), including cluster architecture, pod lifecycle, networking (CNI), storage (CSI), service mesh, and production-grade operations.
Preferred Skills and Experience
- Collaborate in a fast-paced, open environment to design foundational systems.
- Strong debugging skills across the full stack — from kernel and OS up through container orchestration layers.
- Deep knowledge of operating systems internals (process scheduling, memory management, file systems, and synchronization primitives).
- Proficiency in performance analysis, profiling, and low-level optimization techniques.
- Solid understanding of computer networks and the TCP/IP stack.
- Experience working with Linux kernel concepts or systems-level debugging tools (e.g., perf, gdb, strace, Wireshark).
- Proficiency deploying and managing workloads using Kubernetes manifests, Helm, Operators, and GitOps workflows.
- Solid understanding of containerization technologies (Docker, containerd, crio) and their interaction with the Linux kernel.
- Experience with observability and monitoring in distributed systems (Prometheus, Grafana, VictoriaMetrics, OpenTelemetry, or similar).
Compensation and Benefits
- $180,000 - $440,000 USD total compensation (base salary is just one part of our total rewards package at xAI, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks).
Skills
Rust, C++, C, Kubernetes, Linux, Docker, Prometheus, Grafana, TCP/IP, Helm
Similar jobs
DevOps / SRE jobsBuild developer-experience tooling and release systems within Benchling’s Platform team, helping engineering teams develop, test, package, and ship high-quality software rapidly. The role requires 4+ years of software engineering experience, web framework expertise, strong problem-solving, and effective cross-functional communication.
Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.
Own Mercor’s internal identity and cloud platform infrastructure as code, automating provisioning, access management, secrets, and employee lifecycle workflows. The role requires production Terraform, Okta, SCIM, and multi-cloud IAM experience, plus strong automation, incident response, and documentation skills.
Production Engineer responsible for building and operating scalable infrastructure, driving reliability and architecture initiatives, and enabling product teams through platform tooling. Requires software engineering experience, distributed-systems expertise, cloud experience, and cross-team technical leadership.
Infrastructure Engineer responsible for securing, scaling, and operating cloud infrastructure and machine-learning platforms across a distributed startup. The role requires Kubernetes, infrastructure-as-code, cloud operations, CI/CD, programming, observability, and security experience.