Software Engineer, ML Infrastructure
Builds large-scale ML infrastructure including GPU clusters, training frameworks, and workload schedulers. Requires strong systems engineering in Python, Rust, Golang, with Kubernetes and distributed systems experience.
About the job
ML Infrastructure Engineer
The ML Infrastructure team builds large-scale compute, storage, and software infrastructure to support Cursor’s work building the world’s best agentic coding model. We’re looking for strong engineers who are interested in building high-performance infrastructure and the software to support it. This role works closely with ML researchers and engineers to enable their work through improvements to our training framework, systems reliability/performance, and developer experience.
What you might do
- Collaborate with ML researchers to improve the throughput and reliability of training
- Work with OEMs, cloud service providers, and others to plan and build cutting-edge GPU infrastructure
- Improve the density and scalability of compute environments to enable increasingly large RL workloads
- Create software and systems to automate building, monitoring, and running GPU clusters
- Build workload scheduling and data movement systems to support Cursor’s growing training footprint
What we’re looking for
- A strong background in systems and infrastructure-focused software engineering, particularly in Python, Typescript, Rust, and Golang
- Experience with distributed storage and networking infrastructure, particularly on Linux systems across cloud and bare metal environments
- Exposure to large-scale systems and their unique challenges, ideally across thousands of nodes with significant resource footprints
- Production use of infrastructure-as-code and configuration management, across hosts and Kubernetes
Nice to have
- Operational exposure to Nvidia GPUs with Infiniband or RoCE, particularly with Blackwell and Hopper-class hardware
- Exposure to Ray, Slurm, or other common compute and runtime schedulers
Skills
Python, TypeScript, Rust, Go, Kubernetes, Linux, Nvidia Gpus, InfiniBand, Ray, Slurm
Similar jobs
ML Engineering jobsBuild and optimize a high-scale LLM inference engine spanning accelerator programming, host-device coordination, and distributed systems. The role requires strong systems programming, performance analysis, and an understanding of LLM inference across compute, memory, and interconnects.
Build and operate production AI agents, automation workflows, and integrations that improve complex business processes. The role requires 5+ years of software engineering experience, modern LLM and agent-framework expertise, systems integration skills, and strong cross-functional collaboration.
Build and operate the engineering systems that support post-training research, including reinforcement learning infrastructure, sandboxed execution, data pipelines, and agent scaffolding. The role requires strong Python and systems engineering skills, project ownership, and a relevant bachelor’s degree or equivalent experience.
Build research infrastructure and tooling that enables AI models to design silicon, including reinforcement learning environments, EDA integrations, evaluations, and experiment workflows. The role requires strong software engineering fundamentals and comfort working across research, tooling, and chip-design systems.
Build production AI capabilities for automated slide and document generation, working across LLM applications, data analysis, and content generation. The role requires 3+ years in machine learning and NLP, advanced Python, and experience with LLM frameworks and production systems.