Software Engineer - Model Performance
Software Engineer optimizes ML model inference performance using techniques like quantization and speculative decoding. Requires backend experience with PyTorch, TensorRT, CUDA, and deep GPU knowledge for LLMs.
About the job
Responsibilities
- Implement, refine, and productionize cutting-edge techniques (quantization, speculative decoding, kv cache reuse, chunked prefill and LoRA) for ML model inference and infrastructure.
- Deep dive into underlying codebases of TensorRT, PyTorch, TensorRT-LLM, vllm, sglang, CUDA, and other libraries to debug ML performance issues.
- Apply and scale optimization techniques across a wide range of ML models, particularly large language models.
- Collaborate with a diverse team to design and implement innovative solutions.
- Own projects from idea to production.
Requirements
- Bachelor's, Master's, or Ph.D. degree in Computer Science, Engineering, Mathematics, or related field.
- Experience with one or more general-purpose programming languages, such as Python or C++.
- Familiarity with LLM optimization techniques (e.g., quantization, speculative decoding, continuous batching).
- Strong familiarity with ML libraries, especially PyTorch, TensorRT, or TensorRT-LLM.
- Demonstrated interest and experience in LLMs.
- Deep understanding of GPU architecture.
Bonus
- Proficiency in enhancing the performance of software systems, particularly in the context of large language models (LLMs).
- Experience with CUDA or similar technologies.
- Deep understanding of software engineering principles and a proven track record of developing and deploying AI/ML inference solutions.
- Experience with Docker and Kubernetes.
Skills
PyTorch, TensorRT, Tensorrt-Llm, CUDA, Python, C++, vLLM, Sglang, Docker, Kubernetes
Similar jobs
ML Engineering jobsBuild and deploy agentic systems that power AI-driven creative video workflows. The role requires 5+ years of experience, production ML or agentic pipeline development, context engineering, and expertise in evaluation and agent infrastructure.
Build and advance agentic machine-learning systems for multimodal creative tasks, with a focus on video understanding, reasoning, control, and tool use. The role requires strong production ML or agent-pipeline experience and deep knowledge of modern LLM techniques.
Build evaluation methods, RL environments, agent tooling, and scalable infrastructure that make subjective qualities such as design and taste measurable for frontier AI models. The role requires experience with evaluations, RL environments, ML or post-training, plus strong backend engineering skills.
Build and scale generative video and multimodal models, optimizing training and inference for efficiency, throughput, and ultra-low latency. The role requires deep learning systems expertise, strong PyTorch/CUDA experience, and the ability to move research models into production.
Build production-grade AI agents, evaluation infrastructure, and developer tooling that make AI-assisted engineering faster, safer, and reusable across teams. The role requires software engineering experience, platform or internal developer-product experience, and hands-on expertise with LLM integration and orchestration.