# Software Engineer, ML Inference Platform

**Company:** [Dialpad](https://hotfix.jobs/companies/dialpad)
**Location:** Buenos Aires, Argentina
**Role:** ML Engineering
**Experience:** 6+ years
**Skills:** Python, Go, Kubernetes, Linux, GCP, Nvidia Gpus, vLLM, Triton, Tgi, Docker, CI/CD, Distributed Systems, Model Serving, Observability, Autoscaling
**Posted:** 2026-08-03

> Builds production inference infrastructure for in-house AI models, including model serving, GPU optimization, deployment safety, benchmarking, and runtime reliability. Requires 6+ years of software engineering experience and strong backend, Kubernetes, Linux, and distributed-systems skills.

## Job Description

## Responsibilities

- Build and improve model-serving pathways for low-latency, high-throughput, high-availability inference workloads.
- Operate and optimize containerized workloads on Kubernetes and Google Cloud, focusing on efficient use of NVIDIA GPUs, memory, storage, and networking.
- Integrate and adapt model-serving frameworks and runtimes such as vLLM, Triton, and TGI for internal deployment, observability, and release requirements.
- Enable shadow serving, canary rollouts, staged deployments, candidate-versus-incumbent comparisons, and fast rollback mechanisms for model-backed services.
- Build benchmarking and evaluation tooling to measure latency, throughput, cost, saturation behavior, and reliability under realistic production traffic.
- Improve packaging, versioning, promotion, deployment, and rollback of model and capability artifacts across environments.
- Strengthen runtime telemetry, structured logging, tracing, dashboards, and alerting for production model-serving behavior.
- Improve compute efficiency, GPU utilization, autoscaling behavior, and cost-performance tradeoffs across the inference platform.

## Requirements

- 6+ years of professional software engineering experience shipping backend services, infrastructure systems, or production platforms.
- Proficiency in Python, Go, or another backend-oriented language.
- Experience building, operating, or optimizing high-throughput services, distributed systems, data or ML infrastructure, or runtime platforms where latency, reliability, and resource utilization matter.
- Hands-on experience with containers, Kubernetes, Linux environments, CI/CD, deployment automation, and production operations.
- Strong debugging, systems-thinking, observability, reproducibility, rollout-safety, and resilience skills.
- Ability to reason about bottlenecks across compute, memory, network, storage, batching, concurrency, and service-level objectives.
- Ability to collaborate with model developers, product engineers, infrastructure teams, and technical leadership.

## Benefits

- Competitive salary and comprehensive benefits.
- Opportunities for professional growth and training.
- AI tools designed to amplify employee impact.
- Inclusive offices and a collaborative work environment.

## Similar jobs

- [Senior Software Engineer, AI / ML Inference Platform](https://hotfix.jobs/jobs/9cb9db74-743b-47e0-a06b-491e968f8ebe) - Dialpad - Buenos Aires, Argentina
- [Staff Machine Learning Engineer](https://hotfix.jobs/jobs/2cb6da8f-4a54-448d-a4b6-fc05f1412a12) - Payabli - Remote
- [AI Engineer - New Verticals](https://hotfix.jobs/jobs/4159c72e-536e-4211-969c-6bcb9ad805fd) - Protege - Remote
- [Agent Systems Engineer](https://hotfix.jobs/jobs/2b04a068-018b-4fe0-915f-4fdf80d4c582) - Adaption Labs - San Francisco, CA
- [Inference Performance Engineer](https://hotfix.jobs/jobs/7718c69a-6d49-445e-8c7c-840bd1b3f346) - Adaption Labs - San Francisco, CA

**Apply:** https://hotfix.jobs/jobs/807077b4-aa44-4459-be5e-667ad64d7407
**Canonical:** https://hotfix.jobs/jobs/807077b4-aa44-4459-be5e-667ad64d7407