ML Infrastructure Engineer
Own and scale ML infrastructure for robotics AI: backend services, GPU orchestration, storage, and internal developer platforms across cloud and on-prem.
About the job
Responsibilities
- Own the architecture, implementation, reliability, and evolution of Maven's machine learning infrastructure.
- Build backend services and platforms for managing data, artifacts, jobs, logs, metadata, and compute resources across cloud and on-premise environments.
- Design scalable systems for workload orchestration, storage, observability, security, and infrastructure automation.
- Build intuitive internal tools and abstractions that make complex infrastructure easy for engineers to use.
- Lead technical and commercial discussions with cloud and ML compute providers, including capacity planning, performance, reliability, and cost.
Requirements
- Significant experience designing, building, and operating production backend, distributed, or compute infrastructure.
- A track record of independently owning complex infrastructure projects from architecture through deployment and ongoing operation.
- Strong programming ability in Python, Go, Rust, C++, or a similar backend or systems language.
- Experience operating GPU compute infrastructure and orchestrating distributed workloads using Kubernetes, Ray, ZenML, or similar systems.
- Experience designing and operating storage systems, observability platforms, infrastructure-as-code, and secure access controls.
- Experience managing large-scale GPU fleets or hybrid cloud and on-premise compute environments.
- Experience building internal developer platforms, CLIs, SDKs, or other self-service infrastructure tools.
- Strong technical judgment, leadership, and communication skills, with the ability to drive decisions across teams and external partners.
- Self-starter attitude with the ability to identify priorities and deliver durable solutions in a fast-paced startup environment.
Nice-to-Haves
- Familiarity with GPU architecture, accelerator-aware software design, and profiling compute-intensive workloads.
- Exposure to infrastructure supporting large-scale robot learning workloads, including policy training, simulation, and multimodal data pipelines.
- Familiarity with SOC 2 controls, security practices, and audit readiness.
Skills
Python, Go, Rust, C++, Kubernetes, Ray, Zenml, Gpu Compute, Infrastructure As Code, Observability
Similar jobs
DevOps / SRE jobsBuild and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.
Build IT workflow automation and security tooling using Go, Temporal, Kubernetes, Terraform, and shell scripting. The role also supports internal IT systems and integrations across endpoint management, identity, access, and security administration.
Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.
Build and operate AWS cloud, ML, LLM, RAG, and IoT infrastructure, including deployment platforms, data pipelines, vector search, observability, security, and cost controls. The role requires deep AWS experience and production experience with LLM-powered applications.