Member of Technical Staff, Systems Infrastructure
Build and operate large-scale scheduling, storage, caching, and networking infrastructure for AI training and inference. The role targets PhD researchers graduating by December 2026 with systems research depth and strong programming and performance-measurement skills.
About the job
Responsibilities
- Design, build, and operate core infrastructure systems for large-scale training and inference.
- Build benchmarks, traces, and simulators to evaluate scheduling, caching, and network performance.
- Identify and eliminate bottlenecks across the stack, from kernels and drivers to scheduler policies.
- Turn systems research ideas into production systems that withstand real workloads, failures, and customers.
- Collaborate with research and inference teams to align infrastructure and model design.
- Track emerging hardware, interconnects, and systems research to inform technical direction.
Areas of Focus
- Scheduling and resource management: GPU job scheduling across heterogeneous hardware, multi-tenant isolation, fair sharing, preemption, topology- and locality-aware placement, autoscaling, fleet utilization, and capacity planning.
- Distributed storage and caching: Storage and caching for model weights, checkpoints, datasets, and KV cache; tiering across memory, local NVMe, and object storage; cache admission and eviction, consistency, cold-start, and weight-loading performance.
- Datacenter networking: RDMA/RoCE and InfiniBand, collective communication performance, congestion control, topology design, load balancing, tail latency, and reliability at fleet scale.
Requirements
- PhD completed within the last six months or expected by December 2026 in Computer Science, Computer Engineering, Electrical Engineering, or a related field.
- Research background in distributed systems, operating systems, scheduling and resource management, storage systems, computer networks, computer architecture, or high-performance computing.
- Depth in at least one focus area, demonstrated through a dissertation, publications, or systems built.
- Strong systems programming skills in C/C++, Rust, Go, or a similar language, plus working proficiency in Python.
- Experience building and evaluating real systems, not only simulations, and rigorously measuring performance.
- Ability to communicate systems designs and results clearly to varied audiences.
Preferred Qualifications
- First-authored publications at top-tier systems or networking venues.
- Infrastructure, cloud, or HPC internships, or meaningful open-source systems contributions.
- Experience with GPU clusters, CUDA, ROCm, Triton, or heterogeneous compute environments.
- Kubernetes and cloud infrastructure experience, or experience operating multi-tenant production clusters.
- Performance profiling and debugging experience in distributed environments.
- Research or engineering achievements demonstrated through grants, fellowships, patents, or systems competitions.
Skills
C++, Rust, Go, Python, Kubernetes, CUDA, Rocm, Triton, Rdma, InfiniBand, Nccl, Rccl, Distributed Systems, Gpu Scheduling, Nvme
Similar jobs
DevOps / SRE jobsBuild and operate robust infrastructure, support enterprise deployments, and improve on-premises delivery for a rapidly scaling AI code review platform. The role requires networking expertise, cloud and container experience, and at least one year of infrastructure or software engineering experience.
Production Engineer builds and operates large-scale systems, focusing on automation, monitoring, infrastructure management, and resilient operations. Requires 2+ years in SRE/DevOps, expertise in Linux, AWS, Kubernetes, and programming in Python or Golang.
Backend engineers build and scale Airtable's infrastructure across teams like Base, Compute, Data, Storage, and Traffic. Requires 2-8 years experience in distributed systems, databases; CS degree; hybrid work in SF, NYC, Seattle, or LA areas.
Build Mercury’s secure, observable infrastructure platform across AWS, networking, containers, and developer tooling. The role requires strong Linux fundamentals, cloud-native experience, technical writing ability, and software development skills, with opportunities to support AI-agent infrastructure.
Infrastructure and site reliability intern building and operating on-premises backend infrastructure for a semiconductor fabrication environment. The role emphasizes systems programming, Linux, networking, reliability, observability, automation, and performance engineering.