# AI Systems Research and Development Engineer – LLM Inference Systems & Optimization

**Company:** [Snowflake](https://hotfix.jobs/companies/snowflake)
**Location:** Bellevue, WA
**Role:** AI Research
**Salary:** $236k – $330k/yr
**Experience:** 5+ years
**Skills:** Llm Inference, Distributed Systems, Gpu Architecture, CUDA, Triton, vLLM, Sglang, Tensorrt-Llm, Cutlass, Cublas, Cudnn, Nsight Systems, Nsight Compute, Kubernetes
**Posted:** 2026-09-09

> Develop and optimize production-scale LLM inference systems across distributed runtimes, GPU kernels, scheduling, and model-system co-design. The role requires a bachelor’s degree and at least five years of experience in inference, distributed AI, GPU systems, or high-performance computing.

## Job Description

## Responsibilities
- Design and develop high-performance LLM inference systems across distributed serving, runtime systems, GPU execution, and performance-critical kernels.
- Improve inference latency, generation speed, throughput, memory efficiency, scalability, and cost.
- Develop techniques including speculative and parallel decoding, prefill/decode disaggregation, adaptive parallelism, continuous batching, scheduling, KV-cache management, quantization, and communication optimization.
- Build adaptive inference systems that optimize execution for new model architectures, hardware, workloads, and deployment environments.
- Apply AI-native approaches to profiling, bottleneck identification, configuration search, code generation, experimentation, debugging, and performance tuning.
- Identify high-impact systems problems, prototype solutions, and drive successful ideas from research through production.
- Design distributed inference strategies across GPUs and nodes, including tensor, sequence, pipeline, data, and expert parallelism.
- Develop multi-model serving, dynamic resource management, model loading and swapping, and workload-aware scheduling.
- Analyze and optimize GPU kernels and operators for attention, mixture-of-experts, communication, and other performance-critical components.
- Explore model-system co-design and post-training techniques for efficient inference.
- Profile and benchmark end-to-end workloads across compute, memory, communication, networking, scheduling, and model execution.
- Collaborate with research, infrastructure, and product teams to deploy innovations in production.
- Open-source and publish innovations through technical blogs and leading systems and machine learning conferences.

## Requirements
- Bachelor’s degree in Computer Science, Electrical Engineering, or a related field.
- 5+ years of experience in LLM inference systems, distributed AI systems, GPU systems, or high-performance computing.
- Strong understanding of modern LLM inference architectures and large-scale model-serving tradeoffs.
- Hands-on experience with LLM inference and serving frameworks such as vLLM, SGLang, TensorRT-LLM, or similar systems.
- Experience designing, extending, or optimizing inference runtimes, including scheduling, batching, KV-cache management, distributed execution, parallelism, speculative decoding, or disaggregated serving.
- Strong understanding of GPU architectures and experience with CUDA, Triton, or similar GPU programming environments.
- Experience with performance-oriented libraries and frameworks such as CUTLASS, cuBLAS, cuDNN, or related technologies.
- Experience profiling and diagnosing end-to-end system performance using Nsight Systems, Nsight Compute, or equivalent tools.
- Ability to independently identify important problems, define technical questions, and drive solutions through ambiguity.
- Ability to work across model, runtime, distributed-system, and hardware layers and reason about end-to-end performance tradeoffs.
- Excellent communication and cross-functional collaboration skills.

## Nice-to-haves
- Master’s degree or PhD.
- Experience using AI-native engineering approaches to accelerate software development, experimentation, debugging, optimization, or system adaptation.

## Compensation
- Annual salary range: $236,000–$330,000.

## Similar jobs

- [Member of Research Staff, Causal Inference](https://hotfix.jobs/jobs/89a9c13a-f535-4360-a252-b7b876682de1) - The Voleon Group - New York, NY - $250k – $275k/yr
- [AI Engineer](https://hotfix.jobs/jobs/821a43a1-d7f0-4e9c-82eb-c41190d3bab9) - Baseten - San Francisco, CA - $220k – $260k/yr
- [Applied Research Scientist, AI Research](https://hotfix.jobs/jobs/bb123228-e1d3-4cd0-a44c-fb5ab09b4766) - Descript - San Francisco, CA - $262k – $299k/yr
- [Research Scientist – Frontier Evaluations](https://hotfix.jobs/jobs/7f457896-ca9c-4429-82a4-fd6190219c60) - AfterQuery - San Francisco, CA - $210k – $450k/yr
- [Research Scientist, APEX Benchmarks](https://hotfix.jobs/jobs/77bc636d-de61-47a6-9ca0-db586f4d2d22) - Mercor - San Francisco, CA - $200k – $400k/yr

**Apply:** https://hotfix.jobs/jobs/6de1a7ac-346f-4206-b1ff-39baef6d5226
**Canonical:** https://hotfix.jobs/jobs/6de1a7ac-346f-4206-b1ff-39baef6d5226