# Senior Software Engineer – AI Infrastructure

**Company:** [Kraken](https://hotfix.jobs/companies/kraken)
**Location:** Remote
**Role:** ML Engineering
**Experience:** 5+ years
**Skills:** Rust, Distributed Systems, MLOps, Model Serving, Machine Learning Infrastructure, Observability, Monitoring, Performance Optimization, Reliability Engineering, Container Orchestration, Cloud-Native Infrastructure, High-Performance Networking, Asynchronous Systems, Llm Systems, Model Evaluation
**Posted:** 2026-06-02

> Build and operate high-performance Rust infrastructure powering production AI agents, including inference, orchestration, execution, reliability, and observability systems. The role requires 5+ years of high-scale production engineering experience and expertise in distributed systems, performance optimization, and ML infrastructure.

## Job Description

## Responsibilities
- Design and build the infrastructure layer powering AI agent systems in production.
- Develop high-performance Rust services for model inference, orchestration, and execution.
- Architect scalable systems supporting millions of users and high request throughput.
- Build reliable ML infrastructure and MLOps patterns for model deployment, evaluation, and monitoring.
- Define guardrails, observability, and failure handling for agent-driven workflows.
- Optimize latency, throughput, and cost across inference and orchestration layers.
- Partner with the Agent Systems team to translate experimental prototypes into hardened production systems.
- Contribute to foundational infrastructure decisions in a high-scale, high-impact environment.

## Requirements
- 5+ years of experience building and operating high-scale production systems.
- Strong proficiency in Rust and systems-level programming.
- Deep understanding of distributed systems, reliability engineering, and performance optimization.
- Experience operating services serving millions of users or supporting high-throughput workloads.
- Familiarity with ML infrastructure, model serving, or MLOps in production environments.
- Experience designing observability, monitoring, and failure-recovery systems.
- Strong collaboration skills across infrastructure and applied engineering teams.
- High ownership mindset in a high-stakes production environment.

## Nice to Have
- Experience building infrastructure for agent-based or LLM-powered systems.
- Background in high-performance networking, asynchronous systems, or low-latency architectures.
- Experience with container orchestration and cloud-native infrastructure.
- Familiarity with evaluation frameworks and model performance monitoring at scale.
- Experience working in fast-moving 0→1 or platform-building teams.

## Similar jobs

- [Senior Applied AI Software Engineer](https://hotfix.jobs/jobs/edd384fa-194d-4ebb-9c16-de26dbcf5680) - Kinter - Remote
- [Senior Software Engineer, AI / ML Inference Platform](https://hotfix.jobs/jobs/9cb9db74-743b-47e0-a06b-491e968f8ebe) - Dialpad - Buenos Aires, Argentina
- [Senior AI Engineer](https://hotfix.jobs/jobs/8b741bc4-bcf4-41fa-a738-e05371dd7267) - Dialpad - Vancouver, Canada - CA$185k – CA$214k/yr
- [Senior Machine Learning System Builder](https://hotfix.jobs/jobs/c13f31b4-43d9-45ca-9eed-572cd079738d) - OPSWAT - Budapest, Hungary
- [Senior Machine Learning Engineer, Economist](https://hotfix.jobs/jobs/be808d21-6efa-4349-9efe-087188c17002) - Instacart - Remote - $173k – $219k/yr

**Apply:** https://hotfix.jobs/jobs/849e1ac0-1c45-4773-9732-10067460ebc7
**Canonical:** https://hotfix.jobs/jobs/849e1ac0-1c45-4773-9732-10067460ebc7