Skip to content

Member of Technical Staff - Model Serving / API Backend Engineer

Build and optimize high-performance ML inference services and APIs that turn frontier research models (FLUX, Stable Diffusion) into production systems serving millions of requests. Requires experience scaling ML serving infrastructure, GPU optimization, and production backend systems.

About the job

What You’ll Work On

  • Turn research checkpoints into production-ready inference services
  • Design and maintain high-performance APIs serving millions of requests
  • Optimize inference latency and throughput across GPU infrastructure
  • Build scalable serving architectures that handle unpredictable traffic
  • Improve reliability, monitoring, and observability across model-serving systems
  • Prototype and ship demos that showcase new capabilities in days, not weeks
  • Collaborate closely with researchers to move from idea to live endpoint rapidly

Tools & Context

  • Python, FastAPI, async systems
  • GPU infrastructure, CUDA, inference optimization
  • Docker and Kubernetes
  • Redis, Postgres, distributed task queues
  • Cloud platforms (AWS, GCP, or Azure)
  • Observability stacks (metrics, logging, tracing)

What We’re Looking For

  • Strong judgment around performance, reliability, and cost tradeoffs
  • Experience scaling APIs or ML systems under load
  • Comfort working in fast-moving, research-adjacent environments
  • Ownership from system design through debugging and deployment

Role-specific experience we value:

  • Building and operating ML inference services in production
  • Designing scalable API architectures with async processing
  • Optimizing GPU workloads (batching, quantization, compilation, CUDA)
  • Managing distributed systems and task queues under variable load
  • Implementing monitoring and observability for production ML systems
  • Debugging performance bottlenecks across model, infrastructure, and network layers

Bonus experience includes:

  • Real-time or low-latency inference systems
  • TensorRT, reduced precision, layer fusion, or model compilation techniques
  • Frontend demo tooling (Streamlit, Gradio, React)
  • CI/CD and automated testing for ML systems
  • Security best practices for API and model serving

Compensation

Base Annual Salary: $180,000–$300,000 USD

Skills

Python, FastAPI, CUDA, Docker, Kubernetes, Redis, Postgres, AWS, GCP, Azure

Mirage

Mirage

New York, NY

Software Engineer, Agents
$175k+/yrOn-site5+ YOEML Engineering

Build and deploy agentic systems that power AI-driven creative video workflows. The role requires 5+ years of experience, production ML or agentic pipeline development, context engineering, and expertise in evaluation and agent infrastructure.

Mirage

Mirage

New York, NY

Research Engineer, Agentic Systems
$175k+/yrOn-siteML Engineering

Build and advance agentic machine-learning systems for multimodal creative tasks, with a focus on video understanding, reasoning, control, and tool use. The role requires strong production ML or agent-pipeline experience and deep knowledge of modern LLM techniques.

Taste Labs

Taste Labs

San Francisco, CA

AI Engineer, RL
$175k+/yrOn-siteML Engineering

Build evaluation methods, RL environments, agent tooling, and scalable infrastructure that make subjective qualities such as design and taste measurable for frontier AI models. The role requires experience with evaluations, RL environments, ML or post-training, plus strong backend engineering skills.

Mirage

Mirage

New York, NY

Research Engineer, Generative Video
$175k+/yrOn-site5+ YOEML Engineering

Build and scale generative video and multimodal models, optimizing training and inference for efficiency, throughput, and ultra-low latency. The role requires deep learning systems expertise, strong PyTorch/CUDA experience, and the ability to move research models into production.

Earnin

Earnin

Mountain View, CA

AI Builder
$189k+/yrHybrid3+ YOEML Engineering

Build production-grade AI agents, evaluation infrastructure, and developer tooling that make AI-assisted engineering faster, safer, and reusable across teams. The role requires software engineering experience, platform or internal developer-product experience, and hands-on expertise with LLM integration and orchestration.