Skip to content
CantinaCantina

Machine Learning Engineer, Ops

Build and scale inference infrastructure for generative audio models including TTS, voice conversion, and ASR. Design high-performance, low-latency serving systems using Kubernetes, CI/CD, and GPU optimization to bridge research and production.

About the job

What You’ll Do

  • Design and maintain inference infrastructure for generative audio model architectures.
  • Implement and manage high-performance inference engines.
  • Orchestrate service deployments using Kubernetes (K8S), implementing advanced autoscaling paradigms to handle varying traffic loads efficiently.
  • Develop and automate robust CI/CD pipelines to streamline the testing and deployment of model artifacts and inference configurations.
  • Monitor production systems, establishing observability practices to track latency, resource utilization, and overall model performance.
  • Collaborate closely with research teams to optimize model serving paths and evaluate various inference strategies.
  • Optimize inference performance for both streaming and batch applications.

What You’ll Bring

  • Deep understanding of modern audio model architectures (e.g., TTS, ASR) and their specific inference requirements.
  • Strong hands-on experience with Kubernetes (K8S), container orchestration, and implementing autoscaling strategies for production workloads.
  • Solid background in MLOps, including CI/CD automation and managing scalable cloud infrastructure.
  • Proficiency in software engineering principles and experience with Python or Go for infrastructure tooling and backend services.
  • Experience with GPU-accelerated inference and performance profiling techniques.
  • Familiarity with high-performance inference engines (e.g., Triton Inference Server, vLLM-Omni) is a plus.

Compensation

The anticipated annual base salary range for this role is between $125,000-$165,000. When determining compensation, a number of factors will be considered, including skills, experience, job scope, location, and competitive compensation market data.

Benefits for U.S.-based roles

  • Competitive salary and generous company equity
  • Medical, dental, and vision insurance – 99.99% of premiums covered by Cantina
  • 42 days of paid time off, including: 15 PTO days, 10 sick days, 15 company holidays, 2 floating holidays
  • Generous parental leave & fertility support
  • 401(k) retirement savings plan
  • Lifestyle spending account – $500/month to use however you’d like
  • Complimentary lunch and snacks for in-office employees
  • One Medical membership, and more!

Skills

Kubernetes, MLOps, CI/CD, Python, Go, Tts, Asr, GPU, Triton Inference Server, vLLM, Autoscaling, Observability

Build

Build

New York, NY
AI Engineer (Assistant)
$125k+/yrOn-siteML Engineering

Build and ship production agentic AI workflows for complex real estate and built-world processes. The role combines product engineering, applied AI, customer collaboration, workflow orchestration, evaluation, and reliable user-facing experiences.

Build

Build

New York, NY

AI Engineer - Workflows
$120k+/yrOn-siteML Engineering

Build modular AI operations and evaluation systems that power complex real estate workflows. The role focuses on improving output quality, defining correctness with domain experts, and reducing human review while maintaining high standards.

Build

Build

New York, NY
AI Engineer - Assistant Experience
$120k+/yrOn-siteML Engineering

Build and operate Dougie, an agentic AI system that executes workflows, evaluates its own performance, retains institutional context, and improves in production. The role requires experience deploying unattended agentic systems and engineering reliable memory, retrieval, orchestration, and feedback loops.

Black Forest Labs

Black Forest Labs

Freiburg, Germany

Member of Technical Staff - Image / Video Generation
€130k+/yrHybridML Engineering

Trains and fine-tunes large-scale diffusion transformer models for image and video generation, conducts rigorous ablation studies, and optimizes distributed training. Requires hands-on diffusion-model experience, strong PyTorch and transformer expertise, and understanding of generative-model evaluation.

Mintlify

Mintlify

San Francisco, CA

Applied AI Engineer
$130k+/yrOn-site4+ YOEML Engineering

Build and own customer-facing AI products from experimentation through production, including reliable agents, evaluation systems, APIs, interfaces, and infrastructure. Requires at least four years of software development experience and deep production experience with language-model systems.