Software Engineer, ML Serving

Own the serving infrastructure connecting ML inference engines to production, building real-time TTS systems on GPU fleets with distributed model serving and cloud infrastructure.

San Francisco, CAML EngineeringOnsite

Apply

About the role

What You'll Own

Architecture and implementation of Rime's TTS serving infrastructure, from GPU-backed inference engines to the API surface.
Model optimization from a single-node to disaggregated fleet serving.
Compatibility with different NVIDIA hardwares from Hopper to Blackwell and beyond for on-prem and cloud deployments.
Continuous integration and deployment workflows for the model serving pipeline.
Site reliability: on-call rotation, monitoring, alerting, and observability across the serving stack.
Resource provision, cost management across our GPU fleet.

What We're Looking For

Hands-on experience with real-time multinode ML serving infrastructure — ML serving framework experience: NVIDIA Dynamo/Triton, vLLM, SGLang, or equivalent.
Experience with distributed or disaggregated model serving (Tensor Parallel, Pipeline Parallel, or equivalent).
Strong cloud infrastructure fundamentals: Linux internals, networking, containerization (Docker, Kubernetes).
IaC experience — Terraform, Packer, or comparable.
On-call is part of the job. You treat production reliability as a shared responsibility.

Nice to Have

Experience with multinode training (DDP, FSDP, etc.).
Experience with gRPC or other bidirectional binary streaming protocols.
Experience with audio streaming and related technologies (WebRTC, WebSockets, etc.).
Experience with a multilingual monorepo where you pick the best language out of merit more than personal experience.
Experience with multi-cloud infrastructures (AWS, GCP, OCI, etc.).
Comfort with configuration management tooling (Ansible, Chef, Puppet, or similar).
SRE, DevOps, or platform engineering background at a startup.
Experience at an early-stage company.

Skills

Nvidia TritonvLLMSglangTensor ParallelPipeline ParallelDockerKubernetesTerraformLinuxgRPC

Similar roles

ML Engineering jobs

Mirage

ML Engineer, Agentic Systems

ML Engineer building and improving agentic systems powered by LLMs for multimodal video understanding, reasoning, and creative editing tasks at an AI-native video platform. Requires strong production ML experience with transformers, fine-tuning, and experimental rigor.

175k – 275kNew York, NYML EngineeringOn-siteLLMsPython

Mirage

Software Engineer, Agents

Design and build agentic systems for AI-native video creation, integrating LLMs and evaluation frameworks to power creative workflows. Requires 5+ years building ML/agentic systems in production.

175k – 275kNew York, NYML EngineeringOn-site5+ YOERAGLLMs

Pindrop

Research Scientist II

Research Scientist II building and improving fraud risk models and scam detection systems using audio, behavioral, and metadata signals. Requires an advanced degree and 3+ years of applied ML experience with Python and modern ML frameworks.

160k – 185kUnited StatesML EngineeringRemote3+ YOELLMsKeras

Harvey

Research Engineer, Post-Training

Research engineer focused on post-training LLMs and agents for legal work. Requires hands-on experience training open-weight models and strong Python/research engineering skills.

231k – 340kSan Francisco, CAML EngineeringHybridSftRLHF

AI Fund

AI Engineer

Build full-stack AI prototypes and agentic systems to pressure-test venture ideas. Requires 3+ years building production AI applications with strong frontend/backend fluency and frontier coding agent expertise.

150k – 190kMountain View, CAML EngineeringOn-site3+ YOESQLAPIs