# Member of Technical Staff - Model Serving / API Backend Engineer

**Company:** [Black Forest Labs](https://hotfix.jobs/companies/black-forest-labs)
**Location:** San Francisco, CA
**Role:** ML Engineering
**Salary:** $180k – $300k/yr
**Experience:** 5+ years
**Skills:** Python, FastAPI, CUDA, Docker, Kubernetes, Redis, Postgres, AWS, GCP, Azure
**Posted:** 2026-07-19

> Build and optimize high-performance ML inference services and APIs that turn frontier research models (FLUX, Stable Diffusion) into production systems serving millions of requests. Requires experience scaling ML serving infrastructure, GPU optimization, and production backend systems.

## Job Description

## What You’ll Work On
- Turn research checkpoints into production-ready inference services
- Design and maintain high-performance APIs serving millions of requests
- Optimize inference latency and throughput across GPU infrastructure
- Build scalable serving architectures that handle unpredictable traffic
- Improve reliability, monitoring, and observability across model-serving systems
- Prototype and ship demos that showcase new capabilities in days, not weeks
- Collaborate closely with researchers to move from idea to live endpoint rapidly

## Tools & Context
- Python, FastAPI, async systems
- GPU infrastructure, CUDA, inference optimization
- Docker and Kubernetes
- Redis, Postgres, distributed task queues
- Cloud platforms (AWS, GCP, or Azure)
- Observability stacks (metrics, logging, tracing)

## What We’re Looking For
- Strong judgment around performance, reliability, and cost tradeoffs
- Experience scaling APIs or ML systems under load
- Comfort working in fast-moving, research-adjacent environments
- Ownership from system design through debugging and deployment

**Role-specific experience we value:**
- Building and operating ML inference services in production
- Designing scalable API architectures with async processing
- Optimizing GPU workloads (batching, quantization, compilation, CUDA)
- Managing distributed systems and task queues under variable load
- Implementing monitoring and observability for production ML systems
- Debugging performance bottlenecks across model, infrastructure, and network layers

**Bonus experience includes:**
- Real-time or low-latency inference systems
- TensorRT, reduced precision, layer fusion, or model compilation techniques
- Frontend demo tooling (Streamlit, Gradio, React)
- CI/CD and automated testing for ML systems
- Security best practices for API and model serving

## Compensation
Base Annual Salary: $180,000–$300,000 USD

## Similar roles

- [Member of Technical Staff](https://hotfix.jobs/jobs/dd7c8bc4-741a-45d5-91a8-e6b1c4ad67e0) - xAI - Palo Alto, CA - $180k – $440k/yr
- [Member of Technical Staff, Applied Research](https://hotfix.jobs/jobs/d19460df-bf56-4f87-aa95-5e7bdf40798f) - LlamaIndex - San Francisco, CA - $180k – $250k/yr
- [Member of Technical Staff, Inference](https://hotfix.jobs/jobs/70280e8d-9e99-4be1-a488-b13edc2d4080) - xAI - Palo Alto, CA - $180k – $440k/yr
- [Member of Technical Staff](https://hotfix.jobs/jobs/544a5b36-3a56-45fa-89fe-0873120ece22) - Black Forest Labs - San Francisco, CA - $180k – $290k/yr
- [Member of Technical Staff - X Search](https://hotfix.jobs/jobs/2480c779-d772-4053-b432-64b3e48d4cd3) - xAI - Palo Alto, CA - $180k – $440k/yr

**Apply:** https://hotfix.jobs/jobs/959837fc-b68e-4e19-9d1c-6dd997000671
**Canonical:** https://hotfix.jobs/jobs/959837fc-b68e-4e19-9d1c-6dd997000671