Software Engineer, ML Inference Platform
Builds production inference infrastructure for in-house AI models, including model serving, GPU optimization, deployment safety, benchmarking, and runtime reliability. Requires 6+ years of software engineering experience and strong backend, Kubernetes, Linux, and distributed-systems skills.
About the job
Responsibilities
- Build and improve model-serving pathways for low-latency, high-throughput, high-availability inference workloads.
- Operate and optimize containerized workloads on Kubernetes and Google Cloud, focusing on efficient use of NVIDIA GPUs, memory, storage, and networking.
- Integrate and adapt model-serving frameworks and runtimes such as vLLM, Triton, and TGI for internal deployment, observability, and release requirements.
- Enable shadow serving, canary rollouts, staged deployments, candidate-versus-incumbent comparisons, and fast rollback mechanisms for model-backed services.
- Build benchmarking and evaluation tooling to measure latency, throughput, cost, saturation behavior, and reliability under realistic production traffic.
- Improve packaging, versioning, promotion, deployment, and rollback of model and capability artifacts across environments.
- Strengthen runtime telemetry, structured logging, tracing, dashboards, and alerting for production model-serving behavior.
- Improve compute efficiency, GPU utilization, autoscaling behavior, and cost-performance tradeoffs across the inference platform.
Requirements
- 6+ years of professional software engineering experience shipping backend services, infrastructure systems, or production platforms.
- Proficiency in Python, Go, or another backend-oriented language.
- Experience building, operating, or optimizing high-throughput services, distributed systems, data or ML infrastructure, or runtime platforms where latency, reliability, and resource utilization matter.
- Hands-on experience with containers, Kubernetes, Linux environments, CI/CD, deployment automation, and production operations.
- Strong debugging, systems-thinking, observability, reproducibility, rollout-safety, and resilience skills.
- Ability to reason about bottlenecks across compute, memory, network, storage, batching, concurrency, and service-level objectives.
- Ability to collaborate with model developers, product engineers, infrastructure teams, and technical leadership.
Benefits
- Competitive salary and comprehensive benefits.
- Opportunities for professional growth and training.
- AI tools designed to amplify employee impact.
- Inclusive offices and a collaborative work environment.
Skills
Python, Go, Kubernetes, Linux, GCP, Nvidia Gpus, vLLM, Triton, Tgi, Docker, CI/CD, Distributed Systems, Model Serving, Observability, Autoscaling
Similar jobs
ML Engineering jobsSenior software engineer building shared AI/ML platform systems across GPU training, model lifecycle management, and production inference. The role requires seven-plus years of production engineering experience, strong backend or infrastructure skills, Kubernetes and cloud expertise, and practical understanding of ML systems and performance.
Sets the technical direction for production machine learning across a payments platform, building and scaling models for risk, authorization, disputes, and forecasting. Requires 8+ years of ML engineering experience, including production model ownership and strong technical leadership.
Build the technical foundation for a new business vertical, creating reusable infrastructure and leading early customer engagements from scoping through delivery. The role requires 3+ years of engineering experience, strong Python and SQL skills, backend/data expertise, and comfort operating in ambiguity.
Build production agent systems that plan, use tools, recover from failures, and improve over time. The role requires 5+ years of production ML or backend experience, LLM or agent deployment experience, and expertise in evaluation, tracing, observability, and agent architecture.
Own inference-stack cost and performance by optimizing serving, caching, batching, quantization, decoding, routing, and GPU execution. The role requires 5+ years in ML systems, inference infrastructure, or performance engineering, plus strong Python and systems-language skills.