What You’ll Work On
- Turn research checkpoints into production-ready inference services
- Design and maintain high-performance APIs serving millions of requests
- Optimize inference latency and throughput across GPU infrastructure
- Build scalable serving architectures that handle unpredictable traffic
- Improve reliability, monitoring, and observability across model-serving systems
- Prototype and ship demos that showcase new capabilities in days, not weeks
- Collaborate closely with researchers to move from idea to live endpoint rapidly
Tools & Context
- Python, FastAPI, async systems
- GPU infrastructure, CUDA, inference optimization
- Docker and Kubernetes
- Redis, Postgres, distributed task queues
- Cloud platforms (AWS, GCP, or Azure)
- Observability stacks (metrics, logging, tracing)
What We’re Looking For
- Strong judgment around performance, reliability, and cost tradeoffs
- Experience scaling APIs or ML systems under load
- Comfort working in fast-moving, research-adjacent environments
- Ownership from system design through debugging and deployment
Role-specific experience we value:
- Building and operating ML inference services in production
- Designing scalable API architectures with async processing
- Optimizing GPU workloads (batching, quantization, compilation, CUDA)
- Managing distributed systems and task queues under variable load
- Implementing monitoring and observability for production ML systems
- Debugging performance bottlenecks across model, infrastructure, and network layers
Bonus experience includes:
- Real-time or low-latency inference systems
- TensorRT, reduced precision, layer fusion, or model compilation techniques
- Frontend demo tooling (Streamlit, Gradio, React)
- CI/CD and automated testing for ML systems
- Security best practices for API and model serving
Compensation
Base Annual Salary: $180,000–$300,000 USD