Software Engineer, Supercomputing
Designs, builds, and operates GPU supercomputing environments for large-scale AI training and inference. Automates cluster management, extends orchestration systems, and optimizes performance metrics in collaboration with researchers.
About the job
What You’ll Do
- Operate and automate large GPU clusters including provisioning, imaging, and capacity planning.
- Write software that abstracts cluster management and presents a unified interface for training and inference.
- Extend scheduling/orchestration (Kubernetes, Slurm, or similar) for topology‑aware placement, preemption, quotas, and fair‑share multi‑tenancy.
- Monitor and improve operational metrics of speed, reliability, and error recovery.
- Build reliable storage and artifact paths for datasets, checkpoints, and logs with clear retention and lineage.
- Partner with researchers to unblock scale runs and advise on parallelism and performance trade‑offs.
Skills and Qualifications
Minimum qualifications:
- Bachelor’s degree or equivalent experience in computer science, engineering, or similar.
- Proficiency in at least one backend language (Python or Rust).
- Experience operating large‑scale clusters and container orchestration systems (e.g. Kubernetes or Slurm).
- Comfort operating across the stack and owning projects end-to-end.
- Thrive in a highly collaborative environment involving many, different cross-functional partners and subject matter experts.
- A bias for action with a mindset to take initiative to work across different stacks and different teams where you spot the opportunity to make sure something ships.
Preferred qualifications:
- Strong systems background: Linux, networking, and infrastructure‑as-code.
- Familiarity with CUDA/NCCL and performance profiling for distributed training/inference.
- Prior work supporting large‑scale model training or inference environments.
- Understanding of deep learning frameworks (e.g., PyTorch, TensorFlow, JAX) and their underlying system architectures.
- Track record of working in fast-paced environments balancing care with urgency.
Compensation
Depending on background, skills and experience, the expected annual salary range for this position is $350,000 - $475,000 USD.
Skills
Kubernetes, Slurm, Python, Rust, Linux, CUDA, Nccl, PyTorch, TensorFlow, JAX
Similar jobs
DevOps / SRE jobsSite Reliability Engineer drives end-to-end reliability for AI fine-tuning platform Tinker, including SLOs, monitoring, incident response, and multi-tenant GPU scheduling. Requires distributed systems experience, software proficiency for reliability, and production incident handling.
Build and operate an AI-first CI/CD and agent-operations platform for Salesforce and custom GTM applications. The role focuses on governed releases, approval workflows, observability, rollback, sandboxing, and SOX-compliant auditability.
Build secure, scalable infrastructure, data systems, compute tooling, and developer experiences for Anthropic’s Interpretability research team. The role partners closely with researchers, security, and platform teams and requires strong programming and infrastructure experience.
Build and operate scalable build systems, CI pipelines, and developer infrastructure for consumer-device software. The role requires 5+ years of engineering experience, expertise with Bazel or comparable build systems, and experience improving CI reliability and performance at scale.
Designs, operates, and improves secure enterprise networks spanning offices, campuses, cloud environments, and connectivity services. The role combines architecture, production operations, troubleshooting, observability, security, and infrastructure automation.