Skip to content
DialpadDialpad

Software Engineer, ML Inference Platform

Builds production inference infrastructure for in-house AI models, including model serving, GPU optimization, deployment safety, benchmarking, and runtime reliability. Requires 6+ years of software engineering experience and strong backend, Kubernetes, Linux, and distributed-systems skills.

About the job

Responsibilities

  • Build and improve model-serving pathways for low-latency, high-throughput, high-availability inference workloads.
  • Operate and optimize containerized workloads on Kubernetes and Google Cloud, focusing on efficient use of NVIDIA GPUs, memory, storage, and networking.
  • Integrate and adapt model-serving frameworks and runtimes such as vLLM, Triton, and TGI for internal deployment, observability, and release requirements.
  • Enable shadow serving, canary rollouts, staged deployments, candidate-versus-incumbent comparisons, and fast rollback mechanisms for model-backed services.
  • Build benchmarking and evaluation tooling to measure latency, throughput, cost, saturation behavior, and reliability under realistic production traffic.
  • Improve packaging, versioning, promotion, deployment, and rollback of model and capability artifacts across environments.
  • Strengthen runtime telemetry, structured logging, tracing, dashboards, and alerting for production model-serving behavior.
  • Improve compute efficiency, GPU utilization, autoscaling behavior, and cost-performance tradeoffs across the inference platform.

Requirements

  • 6+ years of professional software engineering experience shipping backend services, infrastructure systems, or production platforms.
  • Proficiency in Python, Go, or another backend-oriented language.
  • Experience building, operating, or optimizing high-throughput services, distributed systems, data or ML infrastructure, or runtime platforms where latency, reliability, and resource utilization matter.
  • Hands-on experience with containers, Kubernetes, Linux environments, CI/CD, deployment automation, and production operations.
  • Strong debugging, systems-thinking, observability, reproducibility, rollout-safety, and resilience skills.
  • Ability to reason about bottlenecks across compute, memory, network, storage, batching, concurrency, and service-level objectives.
  • Ability to collaborate with model developers, product engineers, infrastructure teams, and technical leadership.

Benefits

  • Competitive salary and comprehensive benefits.
  • Opportunities for professional growth and training.
  • AI tools designed to amplify employee impact.
  • Inclusive offices and a collaborative work environment.

Skills

Python, Go, Kubernetes, Linux, GCP, Nvidia Gpus, vLLM, Triton, Tgi, Docker, CI/CD, Distributed Systems, Model Serving, Observability, Autoscaling

Dialpad

Dialpad

Buenos Aires, Argentina

Senior Software Engineer, AI / ML Inference Platform
No salary listedOn-site7+ YOEML Engineering

Senior software engineer building shared AI/ML platform systems across GPU training, model lifecycle management, and production inference. The role requires seven-plus years of production engineering experience, strong backend or infrastructure skills, Kubernetes and cloud expertise, and practical understanding of ML systems and performance.

Payabli

Payabli

Remote

Staff Machine Learning Engineer
No salary listedRemote8+ YOEML Engineering

Sets the technical direction for production machine learning across a payments platform, building and scaling models for risk, authorization, disputes, and forecasting. Requires 8+ years of ML engineering experience, including production model ownership and strong technical leadership.

Protege

Protege

Remote

AI Engineer - New Verticals
No salary listedRemote3+ YOEML Engineering

Build the technical foundation for a new business vertical, creating reusable infrastructure and leading early customer engagements from scoping through delivery. The role requires 3+ years of engineering experience, strong Python and SQL skills, backend/data expertise, and comfort operating in ambiguity.

Adaption Labs

Adaption Labs

San Francisco, CA

Agent Systems Engineer
No salary listedHybrid5+ YOEML Engineering

Build production agent systems that plan, use tools, recover from failures, and improve over time. The role requires 5+ years of production ML or backend experience, LLM or agent deployment experience, and expertise in evaluation, tracing, observability, and agent architecture.

Adaption Labs

Adaption Labs

San Francisco, CA

Inference Performance Engineer
No salary listedHybrid5+ YOEML Engineering

Own inference-stack cost and performance by optimizing serving, caching, batching, quantization, decoding, routing, and GPU execution. The role requires 5+ years in ML systems, inference infrastructure, or performance engineering, plus strong Python and systems-language skills.