Engineer, Supercomputing & Distributed Systems
Builds and operates supercomputing infrastructure for AI research including 1000+ GPU Kubernetes clusters, distributed data pipelines processing petabytes, and fault-tolerant training systems. Requires strong distributed systems intuition and experience with Python, PyTorch, and large-scale infrastructure.
About the job
Responsibilities
Distributed data systems
- Design multi-stage pipelines that turn petabytes of raw data into clean, annotated datasets
- Run classification models on billions of images
- Deploy and combine LLMs to caption massive multimedia data
GPU infrastructure
- Manage distributed training and inference on 1000+ GPU Kubernetes clusters
- Solve orchestration and scaling for large-scale GPU job processing
- Scale workloads and research between clusters in multiple datacenters
Distributed training
- Profile and optimize dataloaders streaming thousands of images per second
- Profile and debug InfiniBand networking on huge training runs
- Build fault tolerance systems for large-scale pretraining
- Collaborate with researchers on evolving RL infrastructure
Applied ML pipelines
- Find clean scenes in millions of videos using distributed shot-boundary detection
- Customize and train models to filter billions of images for questions like "is this a screenshot?"
- Build the systems that bridge raw cluster capacity and research output
Requirements
- Intuition for distributed systems and great mental model of how systems interact under different conditions
- Work heavily with Python, Kubernetes, Torch, and data tools like DuckDB, Arrow
Strong candidates may have experience with:
- Python, PyArrow, DuckDB, SQL, massive relational databases, PyTorch, Pandas, NumPy
- Kubernetes
- Designing and implementing large-scale ETL systems
- Fundamental knowledge of containerization, operating systems, file-systems, and networking
- Distributed systems design
- Distributed training systems (NCCL, InfiniBand, RDMA)
- Streaming and event processing systems (Kafka, Pulsar, or similar)
- PyTorch internals, custom dataloaders, and training infrastructure
Skills
Python, Kubernetes, PyTorch, Duckdb, Pyarrow, InfiniBand, Nccl, Rdma, SQL, ETL, Arrow, pandas, NumPy, Kafka, Pulsar
Similar jobs
DevOps / SRE jobsDesigns and operates foundational developer-infrastructure services for CI, builds, deployments, and testing. The role requires senior-level systems engineering, end-to-end service ownership, and cross-functional technical leadership.
Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.
Build IT workflow automation and security tooling using Go, Temporal, Kubernetes, Terraform, and shell scripting. The role also supports internal IT systems and integrations across endpoint management, identity, access, and security administration.
Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.