Senior Software Engineer, AI / ML Inference Platform
Senior software engineer building shared AI/ML platform systems across GPU training, model lifecycle management, and production inference. The role requires seven-plus years of production engineering experience, strong backend or infrastructure skills, Kubernetes and cloud expertise, and practical understanding of ML systems and performance.
About the job
Responsibilities
- Design, build, and improve platform capabilities spanning model training, evaluation, artifact management, release, production inference, and operational feedback.
- Build and operate reliable, reproducible GPU training infrastructure.
- Improve cluster scheduling, workload isolation, capacity management, storage, networking, observability, and accelerator utilization.
- Develop low-latency, high-throughput, highly available inference services.
- Integrate and adapt model-training frameworks and inference runtimes for automation, observability, security, and operational control.
- Optimize GPU workloads across compute, memory, storage, networking, batching, concurrency, and scheduling.
- Partner with ASR and NLP scientists on scalable production designs and model lifecycle concerns.
- Version, trace, validate, compare, promote, deploy, and roll back models and artifacts across environments.
- Build safe release processes using automated checks, shadow traffic, staged rollouts, candidate-versus-incumbent comparisons, and rollback.
- Build benchmarking and evaluation infrastructure covering model quality, latency, throughput, reliability, resource utilization, and cost.
- Strengthen telemetry, logging, tracing, dashboards, alerting, and diagnostic tooling.
- Build self-service workflows and standards for AI teams.
- Lead technical projects, contribute to architecture, and mentor engineers.
Requirements
- Seven or more years of professional software engineering experience owning backend, infrastructure, distributed, or ML platform systems in production.
- Experience building or operating systems supporting model training, model inference, or the connecting lifecycle.
- Proficiency in Python, Go, or another backend-oriented language.
- Hands-on experience with Linux, containers, Kubernetes, cloud infrastructure, CI/CD, deployment automation, and production operations.
- Experience operating GPU workloads and optimizing utilization, memory, storage, networking, scheduling, and performance.
- Working knowledge of datasets, experiments, distributed execution, checkpoints, reproducibility, and model artifacts.
- Understanding of dataset quality, evaluation design, experimental validity, error analysis, model-quality metrics, and production model behavior.
- Ability to make trade-offs among model quality, latency, throughput, reliability, capacity, and cost.
- Strong judgment around observability, repeatability, release safety, failure containment, rollback, and resilience.
- Ability to lead ambiguous projects, communicate across disciplines, mentor engineers, and influence technical decisions.
Nice-to-haves
- Experience with ASR, speech processing, NLP, large language models, or production model-backed systems.
- GPU-based or distributed model training experience.
- Experience with vLLM, Triton, TGI, or comparable model-serving runtimes.
Skills
Python, Go, Linux, Kubernetes, GCP, Docker, CI/CD, Gpu Computing, Distributed Systems, Model Inference, Model Training, Nvidia Gpus, vLLM, Triton, Tgi
Similar jobs
ML Engineering jobsSets the technical direction for production machine learning across a payments platform, building and scaling models for risk, authorization, disputes, and forecasting. Requires 8+ years of ML engineering experience, including production model ownership and strong technical leadership.
Build the technical foundation for a new business vertical, creating reusable infrastructure and leading early customer engagements from scoping through delivery. The role requires 3+ years of engineering experience, strong Python and SQL skills, backend/data expertise, and comfort operating in ambiguity.
Build production agent systems that plan, use tools, recover from failures, and improve over time. The role requires 5+ years of production ML or backend experience, LLM or agent deployment experience, and expertise in evaluation, tracing, observability, and agent architecture.
Own inference-stack cost and performance by optimizing serving, caching, batching, quantization, decoding, routing, and GPU execution. The role requires 5+ years in ML systems, inference infrastructure, or performance engineering, plus strong Python and systems-language skills.