Skip to content

ML Infrastructure Engineer

Own and scale ML infrastructure for robotics AI: backend services, GPU orchestration, storage, and internal developer platforms across cloud and on-prem.

About the job

Responsibilities

  • Own the architecture, implementation, reliability, and evolution of Maven's machine learning infrastructure.
  • Build backend services and platforms for managing data, artifacts, jobs, logs, metadata, and compute resources across cloud and on-premise environments.
  • Design scalable systems for workload orchestration, storage, observability, security, and infrastructure automation.
  • Build intuitive internal tools and abstractions that make complex infrastructure easy for engineers to use.
  • Lead technical and commercial discussions with cloud and ML compute providers, including capacity planning, performance, reliability, and cost.

Requirements

  • Significant experience designing, building, and operating production backend, distributed, or compute infrastructure.
  • A track record of independently owning complex infrastructure projects from architecture through deployment and ongoing operation.
  • Strong programming ability in Python, Go, Rust, C++, or a similar backend or systems language.
  • Experience operating GPU compute infrastructure and orchestrating distributed workloads using Kubernetes, Ray, ZenML, or similar systems.
  • Experience designing and operating storage systems, observability platforms, infrastructure-as-code, and secure access controls.
  • Experience managing large-scale GPU fleets or hybrid cloud and on-premise compute environments.
  • Experience building internal developer platforms, CLIs, SDKs, or other self-service infrastructure tools.
  • Strong technical judgment, leadership, and communication skills, with the ability to drive decisions across teams and external partners.
  • Self-starter attitude with the ability to identify priorities and deliver durable solutions in a fast-paced startup environment.

Nice-to-Haves

  • Familiarity with GPU architecture, accelerator-aware software design, and profiling compute-intensive workloads.
  • Exposure to infrastructure supporting large-scale robot learning workloads, including policy training, simulation, and multimodal data pipelines.
  • Familiarity with SOC 2 controls, security practices, and audit readiness.

Skills

Python, Go, Rust, C++, Kubernetes, Ray, Zenml, Gpu Compute, Infrastructure As Code, Observability

Granica

Granica

Remote

Software Engineer, Infrastructure
No salary listedRemote5+ YOEDevOps / SRE

Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.

Baseten

Baseten

San Francisco, CA

Capacity Ops Engineer
$170k+/yrHybrid5+ YOEDevOps / SRE

Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.

Teleport

Teleport

United States

IT Security and Automation Engineer
$149k+/yrRemoteDevOps / SRE

Build IT workflow automation and security tooling using Go, Temporal, Kubernetes, Terraform, and shell scripting. The role also supports internal IT systems and integrations across endpoint management, identity, access, and security administration.

Crusoe

Crusoe

United States

Electrical Field Engineer - Data Center
$196k+/yrRemote5+ YOEDevOps / SRE

Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.

Beacon AI

Beacon AI

San Carlos, CA

Software Engineer, Cloud Infrastructure
$135k+/yrHybridDevOps / SRE

Build and operate AWS cloud, ML, LLM, RAG, and IoT infrastructure, including deployment platforms, data pipelines, vector search, observability, security, and cost controls. The role requires deep AWS experience and production experience with LLM-powered applications.