Skip to content
DialpadDialpad

Senior Software Engineer, AI / ML Inference Platform

Senior software engineer building shared AI/ML platform systems across GPU training, model lifecycle management, and production inference. The role requires seven-plus years of production engineering experience, strong backend or infrastructure skills, Kubernetes and cloud expertise, and practical understanding of ML systems and performance.

About the job

Responsibilities

  • Design, build, and improve platform capabilities spanning model training, evaluation, artifact management, release, production inference, and operational feedback.
  • Build and operate reliable, reproducible GPU training infrastructure.
  • Improve cluster scheduling, workload isolation, capacity management, storage, networking, observability, and accelerator utilization.
  • Develop low-latency, high-throughput, highly available inference services.
  • Integrate and adapt model-training frameworks and inference runtimes for automation, observability, security, and operational control.
  • Optimize GPU workloads across compute, memory, storage, networking, batching, concurrency, and scheduling.
  • Partner with ASR and NLP scientists on scalable production designs and model lifecycle concerns.
  • Version, trace, validate, compare, promote, deploy, and roll back models and artifacts across environments.
  • Build safe release processes using automated checks, shadow traffic, staged rollouts, candidate-versus-incumbent comparisons, and rollback.
  • Build benchmarking and evaluation infrastructure covering model quality, latency, throughput, reliability, resource utilization, and cost.
  • Strengthen telemetry, logging, tracing, dashboards, alerting, and diagnostic tooling.
  • Build self-service workflows and standards for AI teams.
  • Lead technical projects, contribute to architecture, and mentor engineers.

Requirements

  • Seven or more years of professional software engineering experience owning backend, infrastructure, distributed, or ML platform systems in production.
  • Experience building or operating systems supporting model training, model inference, or the connecting lifecycle.
  • Proficiency in Python, Go, or another backend-oriented language.
  • Hands-on experience with Linux, containers, Kubernetes, cloud infrastructure, CI/CD, deployment automation, and production operations.
  • Experience operating GPU workloads and optimizing utilization, memory, storage, networking, scheduling, and performance.
  • Working knowledge of datasets, experiments, distributed execution, checkpoints, reproducibility, and model artifacts.
  • Understanding of dataset quality, evaluation design, experimental validity, error analysis, model-quality metrics, and production model behavior.
  • Ability to make trade-offs among model quality, latency, throughput, reliability, capacity, and cost.
  • Strong judgment around observability, repeatability, release safety, failure containment, rollback, and resilience.
  • Ability to lead ambiguous projects, communicate across disciplines, mentor engineers, and influence technical decisions.

Nice-to-haves

  • Experience with ASR, speech processing, NLP, large language models, or production model-backed systems.
  • GPU-based or distributed model training experience.
  • Experience with vLLM, Triton, TGI, or comparable model-serving runtimes.

Skills

Python, Go, Linux, Kubernetes, GCP, Docker, CI/CD, Gpu Computing, Distributed Systems, Model Inference, Model Training, Nvidia Gpus, vLLM, Triton, Tgi

Payabli

Payabli

Remote

Staff Machine Learning Engineer
No salary listedRemote8+ YOEML Engineering

Sets the technical direction for production machine learning across a payments platform, building and scaling models for risk, authorization, disputes, and forecasting. Requires 8+ years of ML engineering experience, including production model ownership and strong technical leadership.

Protege

Protege

Remote

AI Engineer - New Verticals
No salary listedRemote3+ YOEML Engineering

Build the technical foundation for a new business vertical, creating reusable infrastructure and leading early customer engagements from scoping through delivery. The role requires 3+ years of engineering experience, strong Python and SQL skills, backend/data expertise, and comfort operating in ambiguity.

Adaption Labs

Adaption Labs

San Francisco, CA

Agent Systems Engineer
No salary listedHybrid5+ YOEML Engineering

Build production agent systems that plan, use tools, recover from failures, and improve over time. The role requires 5+ years of production ML or backend experience, LLM or agent deployment experience, and expertise in evaluation, tracing, observability, and agent architecture.

Adaption Labs

Adaption Labs

San Francisco, CA

Inference Performance Engineer
No salary listedHybrid5+ YOEML Engineering

Own inference-stack cost and performance by optimizing serving, caching, batching, quantization, decoding, routing, and GPU execution. The role requires 5+ years in ML systems, inference infrastructure, or performance engineering, plus strong Python and systems-language skills.