Skip to content
KrakenKraken

Senior Software Engineer – AI Infrastructure

Build and operate high-performance Rust infrastructure powering production AI agents, including inference, orchestration, execution, reliability, and observability systems. The role requires 5+ years of high-scale production engineering experience and expertise in distributed systems, performance optimization, and ML infrastructure.

About the job

Responsibilities

  • Design and build the infrastructure layer powering AI agent systems in production.
  • Develop high-performance Rust services for model inference, orchestration, and execution.
  • Architect scalable systems supporting millions of users and high request throughput.
  • Build reliable ML infrastructure and MLOps patterns for model deployment, evaluation, and monitoring.
  • Define guardrails, observability, and failure handling for agent-driven workflows.
  • Optimize latency, throughput, and cost across inference and orchestration layers.
  • Partner with the Agent Systems team to translate experimental prototypes into hardened production systems.
  • Contribute to foundational infrastructure decisions in a high-scale, high-impact environment.

Requirements

  • 5+ years of experience building and operating high-scale production systems.
  • Strong proficiency in Rust and systems-level programming.
  • Deep understanding of distributed systems, reliability engineering, and performance optimization.
  • Experience operating services serving millions of users or supporting high-throughput workloads.
  • Familiarity with ML infrastructure, model serving, or MLOps in production environments.
  • Experience designing observability, monitoring, and failure-recovery systems.
  • Strong collaboration skills across infrastructure and applied engineering teams.
  • High ownership mindset in a high-stakes production environment.

Nice to Have

  • Experience building infrastructure for agent-based or LLM-powered systems.
  • Background in high-performance networking, asynchronous systems, or low-latency architectures.
  • Experience with container orchestration and cloud-native infrastructure.
  • Familiarity with evaluation frameworks and model performance monitoring at scale.
  • Experience working in fast-moving 0→1 or platform-building teams.

Skills

Rust, Distributed Systems, MLOps, Model Serving, Machine Learning Infrastructure, Observability, Monitoring, Performance Optimization, Reliability Engineering, Container Orchestration, Cloud-Native Infrastructure, High-Performance Networking, Asynchronous Systems, Llm Systems, Model Evaluation

Kinter

Kinter

Brazil

Senior Applied AI Software Engineer
No salary listedRemote8+ YOEML Engineering

Build and operate production AI agents for finance workflows, owning orchestration, evaluation, reliability, and human-in-the-loop safeguards. Requires 8+ years building SaaS products, 2+ years shipping agentic systems, and hands-on Node.js and TypeScript experience.

Dialpad

Dialpad

Buenos Aires, Argentina

Senior Software Engineer, AI / ML Inference Platform
No salary listedOn-site7+ YOEML Engineering

Senior software engineer building shared AI/ML platform systems across GPU training, model lifecycle management, and production inference. The role requires seven-plus years of production engineering experience, strong backend or infrastructure skills, Kubernetes and cloud expertise, and practical understanding of ML systems and performance.

Dialpad

Dialpad

Vancouver, Canada

Senior AI Engineer
CA$185k+/yrOn-site5+ YOEML Engineering

Senior AI engineer owning production systems for real-time speech models and AI voice agents. The role requires 5+ years of production software experience, Python, model serving and inference optimization, cloud and distributed systems expertise, and strong reliability and operations skills.

OPSWAT

OPSWAT

Budapest, Hungary

Senior Machine Learning System Builder
No salary listedOn-site5+ YOEML Engineering

Senior machine learning engineer owning cybersecurity detection capabilities end to end, from data pipelines and experimentation through production deployment and monitoring. Requires 3+ years of applied ML experience, strong Python and deep-learning skills, and demonstrated model delivery in customer-facing systems.

Instacart

Instacart

United States
Senior Machine Learning Engineer, Economist
$173k+/yrRemote5+ YOEML Engineering

Build and deploy machine learning systems that apply economic theory, econometrics, and causal inference to marketplace problems. The role requires advanced training in economics, strong Python and data skills, and production ML experience for senior-level hires.