Training, Process Management Engineer
Build and optimize Rust- and Python-based distributed systems that orchestrate, monitor, and debug machine-learning workloads across extremely large supercomputing clusters. The role emphasizes performance, correctness, reliability, observability, and fault tolerance.
About the job
Responsibilities
- Work across Python and Rust stacks.
- Design, build, and maintain software to orchestrate and monitor machine learning workloads on large supercomputers.
- Profile and optimize the software stack for computation orchestration at frontier scale.
- Improve reliability, observability, and fault tolerance for long-running jobs.
- Debug complex distributed-systems issues across large clusters.
- Respond to changing ML-system requirements to enable researchers.
Requirements
- Experience developing distributed systems.
- Strong software engineering skills and proficiency in Python and Rust or another systems programming language such as C++.
- Solid Linux knowledge.
- Experience with systems-level debugging, performance analysis, and memory profiling.
- Experience developing asynchronous and concurrent systems.
- Interest in performance, correctness, reliability, and large-scale system behavior.
- Comfort working in high-ownership environments with light process and strong engineering agency.
Benefits
- Hybrid work model with three days per week in the office.
- Relocation assistance for new employees.
Skills
Python, Rust, C++, Linux, Distributed Systems, Asynchronous Systems, Concurrent Systems, Performance Analysis, Memory Profiling, Fault Tolerance
Similar jobs
Backend Engineering jobsBuild and operate Internet-scale HTTP and TLS infrastructure, migrate services to a Rust-based proxy, and improve protocol performance. The role requires systems programming experience, strong reliability and security practices, and interest in open-source standards.
Build and evolve shared platform infrastructure, data systems, and developer tooling for AI products. The role requires 5+ years of backend engineering experience, distributed systems expertise, and familiarity with cloud, orchestration, databases, and CI/CD technologies.
Develop and maintain Spotify’s native desktop application across macOS and Windows, working on C++, platform integrations, delivery systems, backend services, and release tooling. The role requires strong production C++ experience, systems problem-solving, and collaboration across distributed teams.
Design and lead secure, scalable wallet infrastructure spanning custody, key management, signing, authorization, and recovery. The role requires deep expertise in wallet or cryptographic security infrastructure, distributed systems architecture, and modern blockchain account models.
Build and operate Rust- and Go-based distributed microservices powering Cloudflare’s application security products. The role requires 3+ years of professional experience, systems-level programming skills, and experience with production-scale systems, databases, and container orchestration.