AI Systems Research and Development Engineer – LLM Inference Systems & Optimization
Develop and optimize production-scale LLM inference systems across distributed runtimes, GPU kernels, scheduling, and model-system co-design. The role requires a bachelor’s degree and at least five years of experience in inference, distributed AI, GPU systems, or high-performance computing.
About the job
Responsibilities
- Design and develop high-performance LLM inference systems across distributed serving, runtime systems, GPU execution, and performance-critical kernels.
- Improve inference latency, generation speed, throughput, memory efficiency, scalability, and cost.
- Develop techniques including speculative and parallel decoding, prefill/decode disaggregation, adaptive parallelism, continuous batching, scheduling, KV-cache management, quantization, and communication optimization.
- Build adaptive inference systems that optimize execution for new model architectures, hardware, workloads, and deployment environments.
- Apply AI-native approaches to profiling, bottleneck identification, configuration search, code generation, experimentation, debugging, and performance tuning.
- Identify high-impact systems problems, prototype solutions, and drive successful ideas from research through production.
- Design distributed inference strategies across GPUs and nodes, including tensor, sequence, pipeline, data, and expert parallelism.
- Develop multi-model serving, dynamic resource management, model loading and swapping, and workload-aware scheduling.
- Analyze and optimize GPU kernels and operators for attention, mixture-of-experts, communication, and other performance-critical components.
- Explore model-system co-design and post-training techniques for efficient inference.
- Profile and benchmark end-to-end workloads across compute, memory, communication, networking, scheduling, and model execution.
- Collaborate with research, infrastructure, and product teams to deploy innovations in production.
- Open-source and publish innovations through technical blogs and leading systems and machine learning conferences.
Requirements
- Bachelor’s degree in Computer Science, Electrical Engineering, or a related field.
- 5+ years of experience in LLM inference systems, distributed AI systems, GPU systems, or high-performance computing.
- Strong understanding of modern LLM inference architectures and large-scale model-serving tradeoffs.
- Hands-on experience with LLM inference and serving frameworks such as vLLM, SGLang, TensorRT-LLM, or similar systems.
- Experience designing, extending, or optimizing inference runtimes, including scheduling, batching, KV-cache management, distributed execution, parallelism, speculative decoding, or disaggregated serving.
- Strong understanding of GPU architectures and experience with CUDA, Triton, or similar GPU programming environments.
- Experience with performance-oriented libraries and frameworks such as CUTLASS, cuBLAS, cuDNN, or related technologies.
- Experience profiling and diagnosing end-to-end system performance using Nsight Systems, Nsight Compute, or equivalent tools.
- Ability to independently identify important problems, define technical questions, and drive solutions through ambiguity.
- Ability to work across model, runtime, distributed-system, and hardware layers and reason about end-to-end performance tradeoffs.
- Excellent communication and cross-functional collaboration skills.
Nice-to-haves
- Master’s degree or PhD.
- Experience using AI-native engineering approaches to accelerate software development, experimentation, debugging, optimization, or system adaptation.
Compensation
- Annual salary range: $236,000–$330,000.
Skills
Llm Inference, Distributed Systems, Gpu Architecture, CUDA, Triton, vLLM, Sglang, Tensorrt-Llm, Cutlass, Cublas, Cudnn, Nsight Systems, Nsight Compute, Kubernetes
Similar jobs
AI Research jobsConducts causal inference research for financial market prediction and portfolio optimization, developing and validating models from research through live trading. Requires Ph.D.-level coursework, strong causal inference and statistics expertise, mathematical ability, and production Python skills.
Build and ship agentic AI product experiences, internal automation, and customer-facing features across the stack. The role requires 5+ years of software engineering experience, hands-on experience with AI or LLM-powered products, Python proficiency, and strong autonomy.
Applied research scientists develop deep-learning and generative media systems for video, audio, and multimodal editing features that ship to millions of users. The role requires strong PyTorch or TensorFlow skills, rapid experimentation, and evidence of impactful research or production machine-learning work.
Designs, validates, and publishes rigorous evaluations and benchmarks for frontier AI systems across agentic, coding, safety, and expert-domain applications. The role requires strong research publication experience, scientific writing, experimental rigor, and the ability to deliver reproducible evaluation systems.
Leads the design, measurement, publication, and adoption of APEX benchmarks evaluating frontier models on economically valuable professional work. The role requires rigorous research judgment, strong coding and statistical skills, and excellent communication across technical, commercial, and research audiences.