# Software Engineer, ML Platform

**Company:** [Cursor](https://hotfix.jobs/companies/cursor)
**Location:** San Francisco, CA, New York, NY
**Role:** ML Engineering
**Skills:** Distributed Systems, Infrastructure, Linux, Kubernetes, Ray, Cloud Computing, Gpu Scheduling, Data Pipelines, Spark, Apache Flink, OpenTelemetry, Tracing, Job Queues, Observability
**Posted:** 2026-08-31

> Build and operate distributed infrastructure and ML platform primitives that help researchers and product engineers move from experiments to trusted runs on large-scale compute. The role requires production systems experience, strong infrastructure skills, and close collaboration with ML research teams.

## Job Description

## Responsibilities
- Design, build, and operate core platform systems used daily by ML researchers and product engineers.
- Partner with research teams to turn recurring pain points into durable infrastructure.
- Own reliability, performance, and developer experience for assigned systems.
- Ship iteratively in a high-ownership environment, measure impact, and continuously improve platform quality.

## Requirements
- Strong background in systems or infrastructure software engineering.
- Experience building platforms that other engineers depend on.
- Experience owning production distributed systems at meaningful scale, such as ingestion systems, data pipelines, or scheduling and orchestration platforms.
- Comfortable working across Linux, cloud and/or bare-metal environments, and modern orchestration systems such as Kubernetes, Ray, or equivalent.
- Enjoy collaborating closely with ML researchers and product engineers.
- Able to thrive in a high-ownership environment with a short feedback loop.

## Nice-to-haves
- Event ingestion, product analytics pipelines, OpenTelemetry, tracing, and reliable data APIs.
- Data frameworks, Spark, Flink, Ray, and ML dataset or training-data infrastructure.
- Experiment and run monitoring, debugging and evaluation tooling, and agent-friendly observability.
- GPU and cluster scheduling, job queues, node health, and research-compute developer experience.

## Compensation
- Compensation details were not provided.

## Similar jobs

- [Research Software Engineer, Post Training](https://hotfix.jobs/jobs/168e3c8f-8577-4482-bf94-91b3d11744ba) - Thinking Machines Lab - San Francisco, CA - $350k – $475k/yr
- [Software Engineer, AI for Chip Design](https://hotfix.jobs/jobs/abfb017d-1ce2-4d24-901d-16084eb7b3bc) - OpenAI - San Francisco, CA - $266k – $468k/yr
- [AI Software Engineer](https://hotfix.jobs/jobs/c9a0e889-a36e-488f-87d0-e9c27ab63bf0) - Rollstack - Remote
- [Machine Learning Engineer, Ranking & Retrieval](https://hotfix.jobs/jobs/adb5411d-8bfc-48b8-a8e7-15cbf5ea21b2) - ClickUp - Remote - $200k – $250k/yr
- [Machine Learning Engineer III](https://hotfix.jobs/jobs/2359be26-5005-4fa7-94c9-8a86066a6bb5) - PathAI - Boston, MA - $131k – $200k/yr

**Apply:** https://hotfix.jobs/jobs/241cad4b-3e2e-4530-af0b-50499737030f
**Canonical:** https://hotfix.jobs/jobs/241cad4b-3e2e-4530-af0b-50499737030f