ML Platform Engineer
Builds and scales ML compute platform on Kubernetes with Argo Workflows and Ray for distributed training, orchestration, and resource governance. Optimizes performance, debugs issues, and integrates tooling for ML teams at scale. Requires deep Kubernetes, systems, and programming expertise.
About the job
What you will do
- Build and scale our ML compute platform on Kubernetes, using Argo Workflows for training, evaluation, and data processing orchestration
- Design and implement core platform capabilities, including a Ray-based internal SDK for distributed execution, and multi-tenant resource governance — scheduling, priorities, quotas, and policy enforcement across GPU, CPU, memory, and IO
- Improve end-to-end training throughput and platform efficiency by optimizing data access patterns, caching, and removing bottlenecks in storage, network, and resource contention
- Work directly with ML teams to debug complex workload issues, drive root-cause analysis, and turn recurring problems into platform-level fixes
- Evaluate, integrate and extend open-source tooling (Argo Workflows, Ray, Kubernetes ecosystem) to meet evolving platform needs
What you will need
- Strong proficiency in Python or Go; C++ is a plus
- Track record of designing and building scalable, maintainable systems and services
- Experience operating production services end-to-end: APIs, reliability practices, observability
- Deep knowledge of Kubernetes: how scheduling, resource management, controllers, and pod lifecycle actually behave under pressure
- Solid Linux and systems debugging skills: performance investigation, networking, storage/IO
- Ability to troubleshoot complex production issues across logs, metrics, and traces and drive them to resolution
Nice to have
- Experience with Argo Workflows, Ray, MLflow, or comparable distributed ML tooling
- Hands-on experience building or operating large-scale ML training systems: GPU scheduling, distributed training, training data pipelines
- Track record of optimizing resource usage and performance in distributed environments
Skills
Kubernetes, Python, Go, Argo Workflows, Ray, Linux, MLflow, Observability, Distributed Training, Gpu Scheduling
Similar jobs
ML Engineering jobsBuild and operate the engineering systems that support post-training research, including reinforcement learning infrastructure, sandboxed execution, data pipelines, and agent scaffolding. The role requires strong Python and systems engineering skills, project ownership, and a relevant bachelor’s degree or equivalent experience.
Build research infrastructure and tooling that enables AI models to design silicon, including reinforcement learning environments, EDA integrations, evaluations, and experiment workflows. The role requires strong software engineering fundamentals and comfort working across research, tooling, and chip-design systems.
Build production AI capabilities for automated slide and document generation, working across LLM applications, data analysis, and content generation. The role requires 3+ years in machine learning and NLP, advanced Python, and experience with LLM frameworks and production systems.
Build and operate large-scale ranking and retrieval systems that power search relevance, including hybrid lexical/vector search, embeddings, query understanding, and permission-aware retrieval. Requires a bachelor's degree and 5+ years of ML engineering experience in ranking or information retrieval.
Develop and deploy machine learning models for biomedical research and AI products, collaborating with scientific, engineering, and product teams. Requires an advanced quantitative degree, substantial ML experience, Python proficiency, and experience bringing models into production or research applications.