Software Engineer, ML Developer Experience
Build developer-facing tooling, platform services, and ML infrastructure across Ray and Anyscale, spanning CLI, SDK, APIs, workspaces, observability, and production serving. Requires 5+ years of production software experience, strong systems fundamentals, and familiarity with machine learning tooling.
About the job
Responsibilities
- Build developer tooling and MLOps capabilities on Ray for developers and coding agents.
- Develop CLI, SDK, API, and MCP surfaces with self-discovery, structured errors, dry-run support, and consistent platform behavior.
- Improve the path from local code to distributed execution across environments, dependencies, images, authentication, submission, and debugging.
- Build tools for data preparation, fine-tuning, post-training, evaluation, production serving, dataset management, experiment tracking, and lineage.
- Develop model registration, deployment workflows, performance benchmarking, and LLM service metrics.
- Provide observability across interfaces to diagnose failures across jobs, tasks, actors, nodes, and GPUs.
- Design integrations using OpenAI-compatible APIs, OpenTelemetry, and stable Jobs and Services interfaces.
- Design and operate highly available backend services and platform architecture across serverless and bring-your-own-cloud environments.
- Work with users, field teams, distributed systems experts, and machine learning experts to scope, ship, and iterate on products.
Requirements
- Bachelor's degree in Computer Science, Engineering, or equivalent practical experience.
- 5+ years of experience writing high-quality production code.
- Strong foundation in algorithms, data structures, and system design.
- Experience with modern machine learning tooling such as PyTorch, MLflow, and data catalogs.
- Hands-on experience building and operating highly available production services.
- Strong product judgment and experience shipping developer-facing tools.
Nice-to-haves
- Experience building and maintaining open-source projects.
- Experience building and operating machine learning infrastructure in production.
- Experience building and operating highly available serving systems.
- Experience using Ray.
Skills
Ray, Python, PyTorch, MLflow, Data Catalogs, MLOps, OpenTelemetry, Openai Apis, Distributed Systems, System Design, Algorithms, Data Structures, Kubernetes, Cloud Infrastructure, Machine Learning Infrastructure
Similar jobs
Fullstack Engineering jobsBuild and own customer-facing product features across Descript’s editor, media, growth, monetization, and AI surfaces. The role requires 5+ years of full-stack product engineering experience, strong product judgment, and expertise with modern web technologies and experimentation.
Build and lead full-stack, AI-powered software for government services using the Palantir stack and modern web technologies. The role requires 5+ years of engineering experience, active Top Secret or TS/SCI clearance, and the ability to work hybrid from New York City or Washington, DC.
Build and improve how AI agents discover, interpret, and use Firecrawl by shipping agent-facing product experiences and rigorous A/B tests. The role requires 5+ years of experience, strong product engineering ability, behavioral data fluency, and independent execution.
Own and evolve Firecrawl’s open-source repositories, maintaining the core self-hosting experience and building new context primitives such as a document parsing engine. The role requires 5+ years of experience, a substantial open-source track record, and strong software engineering judgment.
Build and ship developer-facing products that enable AI agents to navigate and act across the web. The role requires 3+ years of experience delivering customer-used products, strong API and developer-experience instincts, and comfort owning ambiguous problems end to end.