# Senior Software Engineer, AI / ML Inference Platform

**Company:** [Dialpad](https://hotfix.jobs/companies/dialpad)
**Location:** Buenos Aires, Argentina
**Role:** ML Engineering
**Experience:** 7+ years
**Skills:** Python, Go, Linux, Kubernetes, GCP, Docker, CI/CD, Gpu Computing, Distributed Systems, Model Inference, Model Training, Nvidia Gpus, vLLM, Triton, Tgi
**Posted:** 2026-09-07

> Senior software engineer building shared AI/ML platform systems across GPU training, model lifecycle management, and production inference. The role requires seven-plus years of production engineering experience, strong backend or infrastructure skills, Kubernetes and cloud expertise, and practical understanding of ML systems and performance.

## Job Description

## Responsibilities
- Design, build, and improve platform capabilities spanning model training, evaluation, artifact management, release, production inference, and operational feedback.
- Build and operate reliable, reproducible GPU training infrastructure.
- Improve cluster scheduling, workload isolation, capacity management, storage, networking, observability, and accelerator utilization.
- Develop low-latency, high-throughput, highly available inference services.
- Integrate and adapt model-training frameworks and inference runtimes for automation, observability, security, and operational control.
- Optimize GPU workloads across compute, memory, storage, networking, batching, concurrency, and scheduling.
- Partner with ASR and NLP scientists on scalable production designs and model lifecycle concerns.
- Version, trace, validate, compare, promote, deploy, and roll back models and artifacts across environments.
- Build safe release processes using automated checks, shadow traffic, staged rollouts, candidate-versus-incumbent comparisons, and rollback.
- Build benchmarking and evaluation infrastructure covering model quality, latency, throughput, reliability, resource utilization, and cost.
- Strengthen telemetry, logging, tracing, dashboards, alerting, and diagnostic tooling.
- Build self-service workflows and standards for AI teams.
- Lead technical projects, contribute to architecture, and mentor engineers.

## Requirements
- Seven or more years of professional software engineering experience owning backend, infrastructure, distributed, or ML platform systems in production.
- Experience building or operating systems supporting model training, model inference, or the connecting lifecycle.
- Proficiency in Python, Go, or another backend-oriented language.
- Hands-on experience with Linux, containers, Kubernetes, cloud infrastructure, CI/CD, deployment automation, and production operations.
- Experience operating GPU workloads and optimizing utilization, memory, storage, networking, scheduling, and performance.
- Working knowledge of datasets, experiments, distributed execution, checkpoints, reproducibility, and model artifacts.
- Understanding of dataset quality, evaluation design, experimental validity, error analysis, model-quality metrics, and production model behavior.
- Ability to make trade-offs among model quality, latency, throughput, reliability, capacity, and cost.
- Strong judgment around observability, repeatability, release safety, failure containment, rollback, and resilience.
- Ability to lead ambiguous projects, communicate across disciplines, mentor engineers, and influence technical decisions.

## Nice-to-haves
- Experience with ASR, speech processing, NLP, large language models, or production model-backed systems.
- GPU-based or distributed model training experience.
- Experience with vLLM, Triton, TGI, or comparable model-serving runtimes.

## Similar jobs

- [Staff Machine Learning Engineer](https://hotfix.jobs/jobs/2cb6da8f-4a54-448d-a4b6-fc05f1412a12) - Payabli - Remote
- [AI Engineer - New Verticals](https://hotfix.jobs/jobs/4159c72e-536e-4211-969c-6bcb9ad805fd) - Protege - Remote
- [Agent Systems Engineer](https://hotfix.jobs/jobs/2b04a068-018b-4fe0-915f-4fdf80d4c582) - Adaption Labs - San Francisco, CA
- [Inference Performance Engineer](https://hotfix.jobs/jobs/7718c69a-6d49-445e-8c7c-840bd1b3f346) - Adaption Labs - San Francisco, CA

**Apply:** https://hotfix.jobs/jobs/9cb9db74-743b-47e0-a06b-491e968f8ebe
**Canonical:** https://hotfix.jobs/jobs/9cb9db74-743b-47e0-a06b-491e968f8ebe