# Senior AI Compute Infrastructure Engineer

**Company:** [Kraken](https://hotfix.jobs/companies/kraken)
**Location:** Remote
**Role:** DevOps / SRE
**Experience:** 5+ years
**Skills:** Gpu Clusters, Kubernetes, Linux, Python, Docker, Networking, Storage, Distributed Systems, vLLM, Triton Inference Server, TensorRT, CUDA, Nccl, Ray, Deepspeed
**Posted:** 2026-05-07

> This senior engineer will operate and optimize GPU and accelerator infrastructure for AI training, inference, evaluation, and experimentation. The role requires 5+ years of infrastructure experience, production GPU cluster operations, strong systems fundamentals, and expertise in serving, observability, reliability, and compute-cost optimization.

## Job Description

## Responsibilities
- Own and operate GPU and accelerator clusters for training, inference, evaluation, and experimentation, including drivers, runtimes, kernels, device plugins, node configuration, scheduling primitives, and workload isolation.
- Design infrastructure for running models locally on GPUs where strategically and economically preferable.
- Build and improve scheduling, orchestration, placement, quota management, and utilization systems across heterogeneous accelerator environments.
- Optimize inference pipelines for latency, throughput, reliability, memory efficiency, and cost using serving stacks such as vLLM, Triton Inference Server, and TensorRT.
- Partner with ML engineers and researchers to remove bottlenecks across training, evaluation, batch inference, online inference, deployment, and production debugging.
- Build observability for GPU utilization, memory pressure, queue depth, saturation, token throughput, request latency, failed workloads, capacity pressure, and spend.
- Drive reliability, incident response, alerting, runbooks, and post-incident improvements for always-on AI compute infrastructure.
- Evaluate and integrate new hardware, cloud instance families, specialized accelerators, runtimes, schedulers, and serving frameworks.
- Build tooling that makes GPU usage visible, accountable, and easier for internal teams to consume.
- Contribute to architecture decisions balancing performance, cost efficiency, scalability, operational simplicity, and production safety.

## Requirements
- 5+ years of infrastructure engineering experience, including significant experience with GPU compute, ML infrastructure, distributed systems, high-performance computing, or large-scale production platforms.
- Production or production-like experience operating GPU clusters or accelerator-backed infrastructure, including scheduling, orchestration, utilization monitoring, and cost optimization.
- Strong systems engineering fundamentals across Linux, networking, storage, containers, Kubernetes, distributed runtimes, and production debugging.
- Experience with ML serving frameworks such as vLLM, Triton Inference Server, TensorRT, TorchServe, KServe, Ray Serve, or equivalent systems.
- Proficiency in Python for infrastructure automation, tooling, debugging, integration, and operational workflows.
- Understanding of performance tradeoffs involving batching, concurrency, memory usage, GPU utilization, model size, latency, throughput, availability, and cost.
- Experience optimizing compute costs while maintaining performance, reliability, and availability expectations.
- Experience building observable systems with metrics, logs, traces, dashboards, alerts, and incident workflows.
- Ability to work in high-stakes, always-on environments where uptime, throughput, correctness, and operational discipline are critical.
- Clear communication skills for explaining infrastructure tradeoffs to researchers, product teams, platform engineers, security stakeholders, and engineering leadership.

## Nice-to-haves
- Experience at a frontier AI lab, hyperscaler, high-frequency trading firm, research platform, or high-scale ML organization.
- Familiarity with custom silicon or specialized accelerators such as TPUs, AWS Trainium, and Gaudi.
- Background in capacity planning, procurement input, reserved capacity strategy, cloud accelerator economics, or GPU fleet cost management.
- Experience with distributed training frameworks such as DeepSpeed, Megatron-LM, FSDP, and Ray.
- Experience debugging CUDA, NCCL, kernels, drivers, runtimes, memory, networking, or low-level performance issues.
- Experience with Rust, C++, Go, CUDA, or other systems languages used for performance-critical infrastructure.
- Crypto, financial services, trading infrastructure, or security-sensitive production infrastructure experience.

## Similar jobs

- [Senior Software Engineer, Cloud Engineering](https://hotfix.jobs/jobs/e14b1da4-1208-4648-a785-30451d299dd7) - Mozilla - Remote - CA$95k – CA$139k/yr
- [Senior Software Engineer, Cloud Engineering](https://hotfix.jobs/jobs/d332c3cb-7ad5-4a0d-b2c2-af9f684bd506) - Mozilla - Remote
- [Senior DevSecOps Engineer](https://hotfix.jobs/jobs/9fad0d81-f196-425c-b1d6-cf7adc34ac02) - Shield AI - London, United Kingdom
- [Senior Production Engineer](https://hotfix.jobs/jobs/98eb4848-f8e8-4b98-95bd-c5f2b852c00f) - Clear Street - London, United Kingdom
- [Software Engineer, GPU Infrastructure](https://hotfix.jobs/jobs/0140c9f5-05ed-4ac4-8f3c-89978a1ea960) - Cohere

**Apply:** https://hotfix.jobs/jobs/de938ea8-5c22-4b6d-9a16-407e85d0e3a3
**Canonical:** https://hotfix.jobs/jobs/de938ea8-5c22-4b6d-9a16-407e85d0e3a3