Latest DevOps / SRE jobs at xAI
Job results
Designs and operates highly available OT infrastructure for AI supercomputer campuses, including power, cooling, facility controls, and industrial software platforms. Requires at least three years of OT or ICS administration experience, strong systems and networking skills, and onsite availability in the Memphis/Southaven area.
Designs, deploys, and operates high-performance networks powering AI supercomputer campuses, including training fabrics, storage, OT, and site networks. Requires substantial data-center networking experience, automation expertise, and readiness for on-call, hands-on infrastructure work.
Leads campus-scale site reliability for a data center environment, owning observability, incident command, postmortems, runbooks, and cross-functional reliability initiatives across infrastructure and facilities. Requires a bachelor's degree or equivalent experience and at least five years in SRE, systems engineering, or large-scale operations.
Designs and delivers large-scale data center connectivity, structured cabling, ISP circuits, and network upgrades for high-bandwidth infrastructure. Requires a bachelor’s degree or equivalent experience, at least five years in data center connectivity, and hands-on expertise with fiber, copper, testing, and low-voltage systems.
Build scalable software, automation, and frameworks for managing large AI network fabrics, including metrics, provisioning, monitoring, configuration, and remediation. The role requires deep networking expertise and a track record of designing reliable systems that orchestrate large device fleets.
Improves facility operations through process standardization, maintenance and construction workflow optimization, operational dashboards, data analysis, and automation. Requires an engineering bachelor's degree, at least one year of operations or process-improvement experience, and Excel, SQL, or Python skills.
Manages data center technicians and critical infrastructure supporting AI compute systems, including power, cooling, networking, hardware deployments, incidents, vendors, and capacity expansion. Requires 5+ years in data center operations and 3+ years managing technical teams.
Supervise data center technicians while overseeing server and network infrastructure installation, maintenance, troubleshooting, and operational improvement. The role requires 5+ years of relevant hardware and repair experience, technical leadership, Linux proficiency, and scripting experience.
Build and optimize large-scale distributed systems powering xAI's massive supercomputing clusters for AI training. Requires strong systems programming in Rust/C++ and deep Kubernetes/Linux expertise.
Build software, services, and frameworks for network management, automation, and monitoring of large-scale GPU supercomputing fabrics. Requires deep network protocol knowledge and experience orchestrating tens of thousands of devices.
Designs, builds, and operates secure, scalable infrastructure including Kubernetes clusters and GPU hardware for large-scale AI workloads in classified US government environments. Requires 5+ years experience, Top Secret clearance, and expertise in IaC tools like Terraform and Ansible.
Senior IT Systems Engineer leads design, implementation, and optimization of SaaS platforms like Okta and Google Workspace, advances IAM programs, drives automation, and troubleshoots complex issues in hybrid environments. Requires 8+ years experience, IAM expertise, and scripting proficiency.
Designs, builds, and optimizes high-speed copper and optical interconnects for large-scale AI/ML clusters. Requires 8+ years experience in high-speed networking, deep knowledge of SerDes, photonics, and Master's/PhD in EE/Photonics/Physics.
Designs, builds, and operates massive-scale compute clusters and custom container orchestration platforms for AI training and inference at exascale. Requires deep expertise in virtualization, containerization, systems programming in C++/Rust, and Linux kernel internals.
Maintains and troubleshoots server and network infrastructure in data centers, focusing on minimizing MTTD and MTTR. Handles racking, cabling, inventory, and on-call emergencies with 5+ years hardware experience required.
IT Systems Engineer builds, manages, and supports Windows/Linux infrastructure, VMware virtualization, and Puppet automation for corporate systems. Requires 3-5 years experience in systems engineering, troubleshooting, scripting, and on-call support in a fast-paced environment.