Skip to content
OpenAIOpenAI

Machine Learning Engineer, Distributed Data Systems

Designs and scales distributed data infrastructure for large-scale multimodal AI training and evaluation. Collaborates with researchers to build reliable, high-performance systems in a fast-paced environment.

About the job

In this role, you will:

  • Design, build, and maintain data infrastructure systems such as distributed compute, data orchestration, distributed storage, streaming infrastructure, machine learning infrastructure while ensuring scalability, reliability, and security.
  • Ensure our data platform can scale by orders of magnitude while remaining reliable and efficient.
  • Partner with researchers to deeply understand requirements and translate them into production-ready systems.
  • Harden, optimize, and maintain critical data infrastructure systems that power multimodal training and evaluation.

You might thrive in this role if you:

  • Have strong experience with distributed systems and large-scale infrastructure with a strong interest in data.
  • Are detail-oriented and bring rigor to building and maintaining reliable systems.
  • Demonstrate excellent software engineering fundamentals and organizational skills.
  • Are comfortable with ambiguity and rapid change.

Skills

Distributed Systems, Machine Learning Infrastructure, Data Orchestration, Distributed Storage, Streaming Infrastructure, Distributed Compute, Scalable Data Platforms

OpenAI

OpenAI

San Francisco, CA

Software Engineer, Trainium
$295k+/yrHybrid3+ YOEML Engineering

Build and optimize OpenAI’s inference stack for AWS Trainium across high-performance kernels, compilers, runtimes, and model execution. The role requires systems programming and accelerator experience, with opportunities to solve end-to-end performance problems for frontier-scale AI models.

Anthropic

Anthropic

San Francisco, CA
Applied AI Engineer, Beneficial Deployments
$280k+/yrHybridML Engineering

Build and deploy LLM-powered tools, agents, and ecosystem infrastructure with life sciences research institutions. The role requires deep scientific or biomedical research experience, production software development expertise, and the ability to translate partner workflows into scalable AI systems.

OpenAI

OpenAI

San Francisco, CA

Software Engineer, AI for Chip Design
$266k+/yrHybridML Engineering

Build research infrastructure and tooling that enables AI models to design silicon, including reinforcement learning environments, EDA integrations, evaluations, and experiment workflows. The role requires strong software engineering fundamentals and comfort working across research, tooling, and chip-design systems.

OpenAI

OpenAI

San Francisco, CA

Software Engineer, Model Runtime
$266k+/yrHybridML Engineering

Build and optimize the production LLM inference runtime for frontier models on OpenAI’s custom silicon. The role spans scheduling, distributed execution, memory and KV-cache management, performance tooling, and hardware-software co-design.

Hyperbound

Hyperbound

San Francisco, CA

Machine Learning Engineer
$260k+/yrOn-siteML Engineering

Build and operate machine learning models for sales roleplay, scoring, and coaching products, owning the lifecycle from fine-tuning and evaluation through production and on-device deployment. The role emphasizes open-source models, latency and privacy optimization, and rigorous model testing.