Skip to content
CursorCursor

Software Engineer, ML Platform

Build and operate distributed infrastructure and ML platform primitives that help researchers and product engineers move from experiments to trusted runs on large-scale compute. The role requires production systems experience, strong infrastructure skills, and close collaboration with ML research teams.

About the job

Responsibilities

  • Design, build, and operate core platform systems used daily by ML researchers and product engineers.
  • Partner with research teams to turn recurring pain points into durable infrastructure.
  • Own reliability, performance, and developer experience for assigned systems.
  • Ship iteratively in a high-ownership environment, measure impact, and continuously improve platform quality.

Requirements

  • Strong background in systems or infrastructure software engineering.
  • Experience building platforms that other engineers depend on.
  • Experience owning production distributed systems at meaningful scale, such as ingestion systems, data pipelines, or scheduling and orchestration platforms.
  • Comfortable working across Linux, cloud and/or bare-metal environments, and modern orchestration systems such as Kubernetes, Ray, or equivalent.
  • Enjoy collaborating closely with ML researchers and product engineers.
  • Able to thrive in a high-ownership environment with a short feedback loop.

Nice-to-haves

  • Event ingestion, product analytics pipelines, OpenTelemetry, tracing, and reliable data APIs.
  • Data frameworks, Spark, Flink, Ray, and ML dataset or training-data infrastructure.
  • Experiment and run monitoring, debugging and evaluation tooling, and agent-friendly observability.
  • GPU and cluster scheduling, job queues, node health, and research-compute developer experience.

Compensation

  • Compensation details were not provided.

Skills

Distributed Systems, Infrastructure, Linux, Kubernetes, Ray, Cloud Computing, Gpu Scheduling, Data Pipelines, Spark, Apache Flink, OpenTelemetry, Tracing, Job Queues, Observability

Thinking Machines Lab

Thinking Machines Lab

San Francisco, CA

Research Software Engineer, Post Training
$350k+/yrHybridML Engineering

Build and operate the engineering systems that support post-training research, including reinforcement learning infrastructure, sandboxed execution, data pipelines, and agent scaffolding. The role requires strong Python and systems engineering skills, project ownership, and a relevant bachelor’s degree or equivalent experience.

OpenAI

OpenAI

San Francisco, CA

Software Engineer, AI for Chip Design
$266k+/yrHybridML Engineering

Build research infrastructure and tooling that enables AI models to design silicon, including reinforcement learning environments, EDA integrations, evaluations, and experiment workflows. The role requires strong software engineering fundamentals and comfort working across research, tooling, and chip-design systems.

Rollstack

Rollstack

United States
AI Software Engineer
No salary listedRemote3+ YOEML Engineering

Build production AI capabilities for automated slide and document generation, working across LLM applications, data analysis, and content generation. The role requires 3+ years in machine learning and NLP, advanced Python, and experience with LLM frameworks and production systems.

ClickUp

ClickUp

United States

Machine Learning Engineer, Ranking & Retrieval
$200k+/yrRemote5+ YOEML Engineering

Build and operate large-scale ranking and retrieval systems that power search relevance, including hybrid lexical/vector search, embeddings, query understanding, and permission-aware retrieval. Requires a bachelor's degree and 5+ years of ML engineering experience in ranking or information retrieval.

PathAI

PathAI

Boston, MA
Machine Learning Engineer III
$131k+/yrOn-site5+ YOEML Engineering

Develop and deploy machine learning models for biomedical research and AI products, collaborating with scientific, engineering, and product teams. Requires an advanced quantitative degree, substantial ML experience, Python proficiency, and experience bringing models into production or research applications.