Senior AI Engineer
Leads production deployment and optimization of large language models on GPU hardware, focusing on quantization, inference engines, serving infrastructure, and performance benchmarking. Requires a bachelor's degree and 5+ years of software engineering experience in ML infrastructure, LLM inference, or model optimization.
About the job
Responsibilities
- Design and build local large language model (LLM) serving environments on GPU hardware, selecting configurations based on VRAM, memory bandwidth, and workload.
- Install and maintain inference stacks including NVIDIA drivers, CUDA, cuDNN, and inference engines across single- and multi-GPU deployments.
- Implement tensor and pipeline parallelism.
- Optimize LLMs for latency, throughput, memory efficiency, and cost.
- Apply quantization from FP32 to FP16/BF16, FP8, FP4, INT8, and INT4 using GPTQ, AWQ, SmoothQuant, post-training quantization, and quantization-aware training.
- Apply pruning, distillation, sparsity, mixed-precision strategies, and calibration to minimize quality loss.
- Deploy and tune vLLM, TensorRT-LLM, TGI, and llama.cpp, including KV-cache optimization, PagedAttention, continuous batching, speculative decoding, FlashAttention, and FP8 tensor cores.
- Build benchmarks for time to first token, inter-token latency, throughput, GPU utilization, and cost per million tokens.
- Deploy production model-serving infrastructure with autoscaling, load balancing, observability, quality regression gates, and A/B testing.
- Research and implement novel inference optimization and model-compression techniques.
Requirements
- Bachelor's degree in Computer Science, Software Engineering, or a related field.
- 5+ years of software engineering experience focused on ML infrastructure, LLM inference, or model optimization.
- Hands-on experience deploying and serving LLMs on GPU hardware in production.
- Strong understanding of quantization and model compression.
- Experience with high-throughput inference engines.
- Understanding of GPU architecture, CUDA, and the memory-bandwidth-bound nature of LLM inference.
- Familiarity with KV cache, continuous batching, PagedAttention, and speculative decoding.
- Understanding of software development methodologies such as Agile and Scrum.
- Strong communication, collaboration, problem-solving, and independent working skills.
- Proficiency in Python; familiarity with C++ and CUDA is a plus.
Nice-to-haves
- Experience writing or tuning custom CUDA or Triton kernels.
- Experience with multi-GPU and distributed inference.
- AWS or Azure architecture certifications.
- Experience with CI/CD and cloud production deployment on Azure, AWS, or GCP, including GPU-backed instances.
Skills
Python, C++, CUDA, Triton, vLLM, Tensorrt-Llm, Tgi, Llama.Cpp, Quantization, Gptq, Awq, Smoothquant, Kubernetes, AWS, Azure
Similar jobs
ML Engineering jobsSets the technical direction for production machine learning across a payments platform, building and scaling models for risk, authorization, disputes, and forecasting. Requires 8+ years of ML engineering experience, including production model ownership and strong technical leadership.
Build the technical foundation for a new business vertical, creating reusable infrastructure and leading early customer engagements from scoping through delivery. The role requires 3+ years of engineering experience, strong Python and SQL skills, backend/data expertise, and comfort operating in ambiguity.
Build and integrate AI/LLM capabilities for a customer and partner portal, including RAG, agents, retrieval systems, and evaluation pipelines. The role requires 3+ years of AI/ML engineering experience, strong Python skills, and production experience serving and evaluating LLMs.
Build and ship locally hosted language-model capabilities for a cybersecurity product, owning training data, fine-tuning, evaluation, security, and constrained-hardware inference. The role requires strong Python, LLM serving and grounding experience, with Rust and cybersecurity knowledge valued.