Skip to content
UnusualUnusual

Software Engineer, ML Serving

Build and scale real-time TTS serving infrastructure for voice AI models, from GPU inference engines to production APIs. Requires hands-on experience with multinode ML serving frameworks, distributed inference, and cloud/SRE practices.

About the job

What You'll Own

  • Architecture and implementation of Rime's TTS serving infrastructure, from GPU-backed inference engines to the API surface.
  • Model optimization from a single-node to disaggregated fleet serving.
  • Compatibility with different NVIDIA hardwares from Hopper to Blackwell and beyond for on-prem and cloud deployments.
  • Continuous integration and deployment workflows for the model serving pipeline.
  • Site reliability: on-call rotation, monitoring, alerting, and observability across the serving stack.
  • Resource provision, cost management across our GPU fleet.

What We're Looking For

  • Hands-on experience with real-time multinode ML serving infrastructure — ML serving framework experience: NVIDIA Dynamo/Triton, vLLM, SGLang, or equivalent.
  • Experience with distributed or disaggregated model serving (Tensor Parallel, Pipeline Parallel, or equivalent).
  • Strong cloud infrastructure fundamentals: Linux internals, networking, containerization (Docker, Kubernetes).
  • IaC experience — Terraform, Packer, or comparable.
  • On-call is part of the job. You treat production reliability as a shared responsibility.

Nice to Have

  • Experience with multinode training (DDP, FSDP, etc.).
  • Experience with gRPC or other bidirectional binary streaming protocols.
  • Experience with audio streaming and related technologies (WebRTC, WebSockets, etc.).
  • Experience with a multilingual monorepo where you pick the best language out of merit more than personal experience.
  • Experience with multi-cloud infrastructures (AWS, GCP, OCI, etc.).
  • Comfort with configuration management tooling (Ansible, Chef, Puppet, or similar).
  • SRE, DevOps, or platform engineering background at a startup.
  • Experience at an early-stage company.

Skills

Ml Serving, Nvidia Triton, vLLM, Sglang, Tensor Parallel, Pipeline Parallel, Linux, Docker, Kubernetes, Terraform, gRPC, Webrtc, AWS, GCP

Thinking Machines Lab

Thinking Machines Lab

San Francisco, CA

Research Software Engineer, Post Training
$350k+/yrHybridML Engineering

Build and operate the engineering systems that support post-training research, including reinforcement learning infrastructure, sandboxed execution, data pipelines, and agent scaffolding. The role requires strong Python and systems engineering skills, project ownership, and a relevant bachelor’s degree or equivalent experience.

OpenAI

OpenAI

San Francisco, CA

Software Engineer, AI for Chip Design
$266k+/yrHybridML Engineering

Build research infrastructure and tooling that enables AI models to design silicon, including reinforcement learning environments, EDA integrations, evaluations, and experiment workflows. The role requires strong software engineering fundamentals and comfort working across research, tooling, and chip-design systems.

Rollstack

Rollstack

United States
AI Software Engineer
No salary listedRemote3+ YOEML Engineering

Build production AI capabilities for automated slide and document generation, working across LLM applications, data analysis, and content generation. The role requires 3+ years in machine learning and NLP, advanced Python, and experience with LLM frameworks and production systems.

ClickUp

ClickUp

United States

Machine Learning Engineer, Ranking & Retrieval
$200k+/yrRemote5+ YOEML Engineering

Build and operate large-scale ranking and retrieval systems that power search relevance, including hybrid lexical/vector search, embeddings, query understanding, and permission-aware retrieval. Requires a bachelor's degree and 5+ years of ML engineering experience in ranking or information retrieval.

PathAI

PathAI

Boston, MA
Machine Learning Engineer III
$131k+/yrOn-site5+ YOEML Engineering

Develop and deploy machine learning models for biomedical research and AI products, collaborating with scientific, engineering, and product teams. Requires an advanced quantitative degree, substantial ML experience, Python proficiency, and experience bringing models into production or research applications.