Skip to content
KrakenKraken

Senior AI Compute Infrastructure Engineer

This senior engineer will operate and optimize GPU and accelerator infrastructure for AI training, inference, evaluation, and experimentation. The role requires 5+ years of infrastructure experience, production GPU cluster operations, strong systems fundamentals, and expertise in serving, observability, reliability, and compute-cost optimization.

About the job

Responsibilities

  • Own and operate GPU and accelerator clusters for training, inference, evaluation, and experimentation, including drivers, runtimes, kernels, device plugins, node configuration, scheduling primitives, and workload isolation.
  • Design infrastructure for running models locally on GPUs where strategically and economically preferable.
  • Build and improve scheduling, orchestration, placement, quota management, and utilization systems across heterogeneous accelerator environments.
  • Optimize inference pipelines for latency, throughput, reliability, memory efficiency, and cost using serving stacks such as vLLM, Triton Inference Server, and TensorRT.
  • Partner with ML engineers and researchers to remove bottlenecks across training, evaluation, batch inference, online inference, deployment, and production debugging.
  • Build observability for GPU utilization, memory pressure, queue depth, saturation, token throughput, request latency, failed workloads, capacity pressure, and spend.
  • Drive reliability, incident response, alerting, runbooks, and post-incident improvements for always-on AI compute infrastructure.
  • Evaluate and integrate new hardware, cloud instance families, specialized accelerators, runtimes, schedulers, and serving frameworks.
  • Build tooling that makes GPU usage visible, accountable, and easier for internal teams to consume.
  • Contribute to architecture decisions balancing performance, cost efficiency, scalability, operational simplicity, and production safety.

Requirements

  • 5+ years of infrastructure engineering experience, including significant experience with GPU compute, ML infrastructure, distributed systems, high-performance computing, or large-scale production platforms.
  • Production or production-like experience operating GPU clusters or accelerator-backed infrastructure, including scheduling, orchestration, utilization monitoring, and cost optimization.
  • Strong systems engineering fundamentals across Linux, networking, storage, containers, Kubernetes, distributed runtimes, and production debugging.
  • Experience with ML serving frameworks such as vLLM, Triton Inference Server, TensorRT, TorchServe, KServe, Ray Serve, or equivalent systems.
  • Proficiency in Python for infrastructure automation, tooling, debugging, integration, and operational workflows.
  • Understanding of performance tradeoffs involving batching, concurrency, memory usage, GPU utilization, model size, latency, throughput, availability, and cost.
  • Experience optimizing compute costs while maintaining performance, reliability, and availability expectations.
  • Experience building observable systems with metrics, logs, traces, dashboards, alerts, and incident workflows.
  • Ability to work in high-stakes, always-on environments where uptime, throughput, correctness, and operational discipline are critical.
  • Clear communication skills for explaining infrastructure tradeoffs to researchers, product teams, platform engineers, security stakeholders, and engineering leadership.

Nice-to-haves

  • Experience at a frontier AI lab, hyperscaler, high-frequency trading firm, research platform, or high-scale ML organization.
  • Familiarity with custom silicon or specialized accelerators such as TPUs, AWS Trainium, and Gaudi.
  • Background in capacity planning, procurement input, reserved capacity strategy, cloud accelerator economics, or GPU fleet cost management.
  • Experience with distributed training frameworks such as DeepSpeed, Megatron-LM, FSDP, and Ray.
  • Experience debugging CUDA, NCCL, kernels, drivers, runtimes, memory, networking, or low-level performance issues.
  • Experience with Rust, C++, Go, CUDA, or other systems languages used for performance-critical infrastructure.
  • Crypto, financial services, trading infrastructure, or security-sensitive production infrastructure experience.

Skills

Gpu Clusters, Kubernetes, Linux, Python, Docker, Networking, Storage, Distributed Systems, vLLM, Triton Inference Server, TensorRT, CUDA, Nccl, Ray, Deepspeed

Mozilla

Mozilla

Canada

Senior Software Engineer, Cloud Engineering
CA$95k+/yrRemote5+ YOEDevOps / SRE

Build and mature Mozilla’s internal developer infrastructure platform, including CI/CD, observability, Kubernetes optimization, environment bootstrapping, and cost optimization. The role requires 5+ years of software engineering experience, cloud-native expertise, and strong technical leadership.

Mozilla

Mozilla

Canada

Senior Software Engineer, Cloud Engineering
No salary listedRemote5+ YOEDevOps / SRE

Senior Software Engineer building and improving Mozilla’s internal developer infrastructure platform, including CI/CD, observability, Kubernetes, cloud optimization, and developer productivity workflows. Requires 5+ years of software engineering experience and expertise in cloud-native or platform engineering.

Shield AI

Shield AI

London, United Kingdom

Senior DevSecOps Engineer
No salary listedHybrid5+ YOEDevOps / SRE

Designs and operates secure development infrastructure and CI/CD pipelines for autonomous defence systems. Requires at least five years of DevOps or related experience, plus UK defence or regulated national-security experience and knowledge of Secure by Design and assurance practices.

Clear Street

Clear Street

London, United Kingdom

Senior Production Engineer
No salary listedOn-site5+ YOEDevOps / SRE

Own production reliability and operational excellence by supporting incidents while building automation, observability, self-healing, and diagnostic tooling. The role requires strong Python, cloud-native, Kubernetes, distributed-systems, and infrastructure-as-code experience.

Cohere

Cohere

United States
Software Engineer, GPU Infrastructure
No salary listedHybrid7+ YOEDevOps / SRE

Build and operate scalable GPU/TPU HPC infrastructure for training and serving frontier AI models. The role partners with AI researchers, optimizes distributed workloads across clouds, and requires expertise in Kubernetes, Python, Go, Linux, and high-performance networking.