Skip to content
Black Forest LabsBlack Forest LabsSan Francisco, CA

Member of Technical Staff - Model Serving / API Backend Engineer

Build and optimize high-performance ML inference services and APIs that turn frontier research models (FLUX, Stable Diffusion) into production systems serving millions of requests. Requires experience scaling ML serving infrastructure, GPU optimization, and production backend systems.

180k – 300k/yr
Hybrid5+ YOEML Engineering

About the role

What You’ll Work On

  • Turn research checkpoints into production-ready inference services
  • Design and maintain high-performance APIs serving millions of requests
  • Optimize inference latency and throughput across GPU infrastructure
  • Build scalable serving architectures that handle unpredictable traffic
  • Improve reliability, monitoring, and observability across model-serving systems
  • Prototype and ship demos that showcase new capabilities in days, not weeks
  • Collaborate closely with researchers to move from idea to live endpoint rapidly

Tools & Context

  • Python, FastAPI, async systems
  • GPU infrastructure, CUDA, inference optimization
  • Docker and Kubernetes
  • Redis, Postgres, distributed task queues
  • Cloud platforms (AWS, GCP, or Azure)
  • Observability stacks (metrics, logging, tracing)

What We’re Looking For

  • Strong judgment around performance, reliability, and cost tradeoffs
  • Experience scaling APIs or ML systems under load
  • Comfort working in fast-moving, research-adjacent environments
  • Ownership from system design through debugging and deployment

Role-specific experience we value:

  • Building and operating ML inference services in production
  • Designing scalable API architectures with async processing
  • Optimizing GPU workloads (batching, quantization, compilation, CUDA)
  • Managing distributed systems and task queues under variable load
  • Implementing monitoring and observability for production ML systems
  • Debugging performance bottlenecks across model, infrastructure, and network layers

Bonus experience includes:

  • Real-time or low-latency inference systems
  • TensorRT, reduced precision, layer fusion, or model compilation techniques
  • Frontend demo tooling (Streamlit, Gradio, React)
  • CI/CD and automated testing for ML systems
  • Security best practices for API and model serving

Compensation

Base Annual Salary: $180,000–$300,000 USD

Skills

PythonFastAPICUDADockerKubernetesRedisPostgresAWSGCPAzure

Similar roles

ML Engineering jobs
xAI

Member of Technical Staff

xAIPalo Alto, CA

Build and optimize the RL training framework and infrastructure for large-scale workloads at SpaceXAI, from ablations to production runs. Requires experience with distributed systems and proficiency in Python, JAX, Rust, or C++.

180k – 440k/yr
On-site5+ YOEML Engineering
LlamaIndex

Member of Technical Staff, Applied Research

LlamaIndexSan Francisco, CA

Build and productionize vision-language models for document understanding at LlamaIndex. Focus on training, fine-tuning, synthetic data, benchmarking, and turning research prototypes into accurate, low-latency production systems for real-world PDFs, tables, and enterprise docs. Requires 3+ years ML engineering/applied research experience with strong PyTorch skills.

180k – 250k/yr
Hybrid3+ YOEML Engineering
xAI

Member of Technical Staff, Inference

xAIPalo Alto, CA

Build and optimize the high-performance inference platform serving Grok at massive scale. Design distributed serving infrastructure, low-level GPU optimizations, quantization, speculative decoding, and CI/CD for production reliability and low latency.

180k – 440k/yr
On-site5+ YOEML Engineering
Black Forest Labs

Member of Technical Staff

Black Forest LabsSan Francisco, CA

Research Engineer embedded in large-scale multimodal model training at Black Forest Labs (creators of FLUX). Own performance, stability, and low-precision optimizations for production training runs; profile, debug distributed systems, implement GPU kernels, and partner with researchers to ship frontier generative models.

180k – 290k/yr
Hybrid7+ YOEML Engineering
xAI

Member of Technical Staff - X Search

xAIPalo Alto, CA

Develops and operates large-scale search engine infrastructure, including retrieval algorithms, indexing, and ML ranking models integrated with Grok AI. Requires experience with search systems, vector databases, and production ML in Python, Go, or Rust.

180k – 440k/yr
On-siteML Engineering