# Machine Learning Engineer, Ops

**Company:** [Cantina](https://hotfix.jobs/companies/cantina)
**Location:** Remote
**Role:** ML Engineering
**Salary:** $125k – $165k/yr
**Skills:** Kubernetes, MLOps, CI/CD, Python, Go, tts, asr, GPU, triton inference server, vLLM, autoscaling, Observability
**Posted:** 2026-07-27

> Build and scale inference infrastructure for generative audio models including TTS, voice conversion, and ASR. Design high-performance, low-latency serving systems using Kubernetes, CI/CD, and GPU optimization to bridge research and production.

## Job Description

## What You’ll Do
- Design and maintain inference infrastructure for generative audio model architectures.
- Implement and manage high-performance inference engines.
- Orchestrate service deployments using Kubernetes (K8S), implementing advanced autoscaling paradigms to handle varying traffic loads efficiently.
- Develop and automate robust CI/CD pipelines to streamline the testing and deployment of model artifacts and inference configurations.
- Monitor production systems, establishing observability practices to track latency, resource utilization, and overall model performance.
- Collaborate closely with research teams to optimize model serving paths and evaluate various inference strategies.
- Optimize inference performance for both streaming and batch applications.

## What You’ll Bring
- Deep understanding of modern audio model architectures (e.g., TTS, ASR) and their specific inference requirements.
- Strong hands-on experience with Kubernetes (K8S), container orchestration, and implementing autoscaling strategies for production workloads.
- Solid background in MLOps, including CI/CD automation and managing scalable cloud infrastructure.
- Proficiency in software engineering principles and experience with Python or Go for infrastructure tooling and backend services.
- Experience with GPU-accelerated inference and performance profiling techniques.
- Familiarity with high-performance inference engines (e.g., Triton Inference Server, vLLM-Omni) is a plus.

## Compensation
The anticipated annual base salary range for this role is between $125,000-$165,000. When determining compensation, a number of factors will be considered, including skills, experience, job scope, location, and competitive compensation market data.

## Benefits for U.S.-based roles
- Competitive salary and generous company equity
- Medical, dental, and vision insurance – 99.99% of premiums covered by Cantina
- 42 days of paid time off, including: 15 PTO days, 10 sick days, 15 company holidays, 2 floating holidays
- Generous parental leave & fertility support
- 401(k) retirement savings plan
- Lifestyle spending account – $500/month to use however you’d like
- Complimentary lunch and snacks for in-office employees
- One Medical membership, and more!

## Similar roles

- [Software Engineer - Axion Data Engine and ML Ops](https://hotfix.jobs/jobs/766f7d67-fc35-42a7-8d7c-67ce1e1c238b) - Applied Intuition - Sunnyvale, CA - $125k – $220k/yr
- [ML Perception Software Engineer](https://hotfix.jobs/jobs/13c384e6-7a36-4b3b-ab53-5a84170b4d04) - Applied Intuition - Sunnyvale, CA - $125k – $222k/yr
- [Software Engineer - Prediction and Behavior ML](https://hotfix.jobs/jobs/d8d6dc59-bef9-4aea-87d3-9bde2cf94bd1) - Applied Intuition - Sunnyvale, CA - $125k – $222k/yr
- [Sensor Sim - ML Engineer](https://hotfix.jobs/jobs/e4a76af5-32ff-4fbe-8821-ce434ed4f18d) - Applied Intuition - Sunnyvale, CA - $125k – $222k/yr
- [Software Engineer - Mapping and Localization](https://hotfix.jobs/jobs/1b7dc1f0-76b6-492f-b6bd-0c6e427cbce4) - Applied Intuition - Sunnyvale, CA - $125k – $222k/yr

**Apply:** https://hotfix.jobs/jobs/1bf55d60-89ba-4ce3-8b09-abd7681839a7
**Canonical:** https://hotfix.jobs/jobs/1bf55d60-89ba-4ce3-8b09-abd7681839a7