Machine Learning Infrastructure Engineer- Model Inference
Builds and optimizes scalable ML inference infrastructure using Kubernetes and GPU resources to deploy production AI models with low latency. Collaborates with ML research and product teams on model serving, orchestration, and compute efficiency.
About the job
What You'll Do
- Design, deploy and maintain scalable Kubernetes clusters for AI model inference and training
- Develop, optimize, and maintain ML model serving infrastructure, ensuring high-performance and low-latency.
- Collaborate with ML and product teams to scale backend infrastructure for AI-driven products, focusing on model deployment, throughput optimization, and compute efficiency.
- Optimize compute-heavy workflows and enhance GPU utilization for ML workloads.
- Build a robust model API orchestration system
- Collaborate with leadership to define and implement strategies for scaling infrastructure as the company grows, ensuring long-term efficiency and performance.
What You’ll Bring
- Strong experience in building and deploying machine learning models in production environments.
- Deep understanding of container orchestration and distributed systems architecture
- Expertise in Kubernetes administration, including custom resource definitions, operators, and cluster management
- Experience developing APIs and managing distributed systems for both batch and real-time workloads
- Excellent communication skills, with the ability to interface between research and product engineering
Ideally, You Have
- Expertise with model serving frameworks such as NVIDIA Triton Server, VLLM, TRT-LLM and so on.
- Expertise with ML toolchains such as PyTorch, Tensorflow or distributed training and inference libraries.
- Familiarity with GPU cluster management and CUDA optimization
- Knowledge of infrastructure as code (Terraform, Ansible) and GitOps practices
- Experience with container registries, image optimization, and multi-stage builds for ML workloads
- Experience orchestrating across ASR models or LLM models for building various GenAI applications
Skills
Kubernetes, PyTorch, TensorFlow, Nvidia Triton Server, vLLM, Trt-Llm, Terraform, Ansible, CUDA, GPU, Distributed Systems, APIs
Similar jobs
ML Engineering jobsBuild and operate production machine-learning systems for content safety, from messy customer data through classification, evaluation, and inference. The role requires 5+ years of ML engineering experience, strong Python and MLOps skills, and sound judgment across classical models and LLMs.
Build AI agent harnesses, models, and product capabilities that enable agents to perform complex work across digital environments. The role combines applied AI research and software engineering, requiring Python proficiency, strong product judgment, and experience with agent tooling, reinforcement learning, or browser technologies.
Builds the platform, verifiers, environments, and grading infrastructure used to evaluate enterprise AI agents at scale. The role combines strong software engineering with expertise in agent runtimes, evaluation design, benchmarks, and production failure analysis.
Optimizes distributed machine learning training and high-throughput offline inference across large accelerator clusters. The role focuses on profiling, scaling efficiency, cluster goodput, GPU performance, and cost-effective processing of autonomy data.
Build and operate production machine-learning systems for search ranking, relevance, extraction quality, and LLM-driven features. The role requires production ML ownership, ranking or relevance expertise, large-scale data experience, Python, and rigorous experimentation skills.