Research Engineer - Inference
Deploy and optimize frontier AI models for fast, reliable, real-time production serving at scale. The role requires production ML serving experience, GPU programming and inference optimization expertise, and the ability to diagnose bottlenecks across the serving stack.
About the job
Responsibilities
- Deploy state-of-the-art AI models to production, owning the path from research checkpoints to serving infrastructure.
- Optimize inference performance across latency, throughput, and cost using quantization, distillation, KV-cache optimization, batching strategies, and custom kernels.
- Build and tune high-performance serving systems for real-time, streaming workloads.
- Create tooling and infrastructure that enables researchers to ship models to production quickly and safely.
- Profile, diagnose, and eliminate bottlenecks across model architecture, kernels, serving infrastructure, and orchestration.
Requirements
- Experience deploying and serving machine-learning models in production, ideally for latency-sensitive or real-time applications.
- Strong engineering skills in GPU programming and inference optimization.
- Ability to independently profile and measure serving-stack performance.
- Demonstrated ability to solve difficult engineering problems through projects, designs, or open-source contributions.
Compensation and Benefits
- Annual professional development stipend.
- Annual stipend for social travel with colleagues.
- Monthly coworking stipend for employees outside major hubs.
- Annual company offsite.
Skills
Gpu Programming, Inference Optimization, CUDA, Triton, TensorRT, vLLM, Sglang, Quantization, Knowledge Distillation, Kv Cache Optimization, Custom Kernels, Model Serving
Similar jobs
ML Engineering jobsBuild and operate the engineering systems that support post-training research, including reinforcement learning infrastructure, sandboxed execution, data pipelines, and agent scaffolding. The role requires strong Python and systems engineering skills, project ownership, and a relevant bachelor’s degree or equivalent experience.
Build research infrastructure and tooling that enables AI models to design silicon, including reinforcement learning environments, EDA integrations, evaluations, and experiment workflows. The role requires strong software engineering fundamentals and comfort working across research, tooling, and chip-design systems.
Build production AI capabilities for automated slide and document generation, working across LLM applications, data analysis, and content generation. The role requires 3+ years in machine learning and NLP, advanced Python, and experience with LLM frameworks and production systems.
Build and operate large-scale ranking and retrieval systems that power search relevance, including hybrid lexical/vector search, embeddings, query understanding, and permission-aware retrieval. Requires a bachelor's degree and 5+ years of ML engineering experience in ranking or information retrieval.
Develop and deploy machine learning models for biomedical research and AI products, collaborating with scientific, engineering, and product teams. Requires an advanced quantitative degree, substantial ML experience, Python proficiency, and experience bringing models into production or research applications.