Skip to content
CantinaCantinaUnited States

Machine Learning Engineer, Ops

Build and scale inference infrastructure for generative audio models including TTS, voice conversion, and ASR. Design high-performance, low-latency serving systems using Kubernetes, CI/CD, and GPU optimization to bridge research and production.

125k – 165k/yr
RemoteML Engineering

About the role

What You’ll Do

  • Design and maintain inference infrastructure for generative audio model architectures.
  • Implement and manage high-performance inference engines.
  • Orchestrate service deployments using Kubernetes (K8S), implementing advanced autoscaling paradigms to handle varying traffic loads efficiently.
  • Develop and automate robust CI/CD pipelines to streamline the testing and deployment of model artifacts and inference configurations.
  • Monitor production systems, establishing observability practices to track latency, resource utilization, and overall model performance.
  • Collaborate closely with research teams to optimize model serving paths and evaluate various inference strategies.
  • Optimize inference performance for both streaming and batch applications.

What You’ll Bring

  • Deep understanding of modern audio model architectures (e.g., TTS, ASR) and their specific inference requirements.
  • Strong hands-on experience with Kubernetes (K8S), container orchestration, and implementing autoscaling strategies for production workloads.
  • Solid background in MLOps, including CI/CD automation and managing scalable cloud infrastructure.
  • Proficiency in software engineering principles and experience with Python or Go for infrastructure tooling and backend services.
  • Experience with GPU-accelerated inference and performance profiling techniques.
  • Familiarity with high-performance inference engines (e.g., Triton Inference Server, vLLM-Omni) is a plus.

Compensation

The anticipated annual base salary range for this role is between $125,000-$165,000. When determining compensation, a number of factors will be considered, including skills, experience, job scope, location, and competitive compensation market data.

Benefits for U.S.-based roles

  • Competitive salary and generous company equity
  • Medical, dental, and vision insurance – 99.99% of premiums covered by Cantina
  • 42 days of paid time off, including: 15 PTO days, 10 sick days, 15 company holidays, 2 floating holidays
  • Generous parental leave & fertility support
  • 401(k) retirement savings plan
  • Lifestyle spending account – $500/month to use however you’d like
  • Complimentary lunch and snacks for in-office employees
  • One Medical membership, and more!

Skills

KubernetesMLOpsCI/CDPythonGottsasrGPUtriton inference servervLLMautoscalingObservability

Similar roles

ML Engineering jobs
Applied Intuition

Software Engineer - Axion Data Engine and ML Ops

Applied IntuitionSunnyvale, CA

Build and optimize ML training pipelines and data engines for perception models on edge devices and in the cloud. Integrate foundation models for automated labeling and work with DoD customers. Requires 5+ years experience with ML infrastructure, microservices, and data systems plus US citizenship.

125k – 220k/yr
On-site5+ YOEML Engineering
Applied Intuition

ML Perception Software Engineer

Applied IntuitionSunnyvale, CA

Develops ML perception algorithms and 4D world representations for autonomous vehicle stacks. Tests on real vehicles and collaborates with research teams. Requires 3+ years experience, C++/Python proficiency, and ML deployment expertise.

125k – 222k/yr
On-site3+ YOEML Engineering
Applied Intuition

Software Engineer - Prediction and Behavior ML

Applied IntuitionSunnyvale, CA

Develops ML-first behavior prediction modules to forecast road user motions and interactions for autonomous systems. Requires 3+ years experience with deep learning end-to-end cycles, C++/Python fluency, and collaboration with perception/planning teams.

125k – 222k/yr
On-site3+ YOEML Engineering
Applied Intuition

Sensor Sim - ML Engineer

Applied IntuitionSunnyvale, CA

Develops and deploys generative ML techniques for production-grade sensor simulation (Lidar, Radar, Cameras) in autonomous systems. Collaborates with research, rendering, and physics teams; requires 5+ years ML experience, Bachelor's in CS, and expertise in large models and 3D geometry.

125k – 222k/yr
On-site5+ YOEML Engineering
Applied Intuition

Software Engineer - Mapping and Localization

Applied IntuitionSunnyvale, CA

As a Software Engineer specializing in Mapping and Localization, you will design and implement state-of-the-art systems for autonomous vehicles and mobile robots. This role involves building scalable solutions, fusing live data with maps, and continuously improving system performance using data-driven methods.

125k – 222k/yr
On-site3+ YOEML Engineering