# Staff Software Engineer, GPU Inference

**Company:** [Cerebras Systems](https://hotfix.jobs/companies/cerebras-systems)
**Location:** Toronto, Canada, Sunnyvale, CA
**Role:** ML Engineering
**Experience:** 8+ years
**Skills:** C++, Python, vLLM, PyTorch, rocm, hip, CUDA, Kubernetes, Linux, Docker, rccl, rdma, tensor parallelism, quantization, gpu profiling
**Posted:** 2026-07-28

> Build and optimize Cerebras’s production GPU inference stack across APIs, vLLM, PyTorch, ROCm, distributed systems, and AMD infrastructure. The role requires 8+ years of software engineering experience, strong C++ and Python skills, and deep expertise in GPU performance, reliability, and model serving.

## Job Description

## Responsibilities
- Design, build, deploy, and maintain the complete GPU prefill path across API services, model-serving workers, vLLM, PyTorch, ROCm, GPU nodes, networking, and rack-scale infrastructure.
- Establish deployment, upgrade, rollback, health-checking, capacity-management, and failure-recovery practices for an AMD GPU fleet.
- Build automation to make driver, firmware, runtime, model, and container compatibility explicit and reproducible.
- Define service-level indicators and objectives for GPU-backed inference; improve fault isolation, graceful degradation, automated recovery, incident response, and post-incident remediation.
- Optimize time to first token, request throughput, tokens per second per GPU, tail latency, GPU utilization, memory efficiency, and rack-level capacity.
- Tune scheduling, continuous batching, prefix caching, KV-cache management, tensor and expert parallelism, request admission, quantization, graph execution, and distributed communication.
- Diagnose failures and performance regressions across application code, vLLM, PyTorch, ROCm/HIP, collective communication libraries, kernels, drivers, firmware, networking, and hardware.
- Build validation and regression infrastructure for model quality, numerical accuracy, precision changes, quantization, determinism, and software/hardware compatibility.
- Develop benchmarks, workload-replay tools, profiling automation, release qualification, dashboards, and regression gates.

## Requirements
- 8+ years of software engineering experience, including substantial individual-contributor ownership of complex production systems.
- Experience building, operating, or optimizing production inference systems for large language models, multimodal models, or similarly demanding GPU workloads.
- Strong programming ability in C++ and Python, including multithreading, concurrency, memory management, and performance-sensitive software.
- Hands-on experience with vLLM, SGLang, TensorRT-LLM, Triton Inference Server, or an equivalent model-serving system.
- Strong understanding of GPU execution and performance, including asynchronous execution, memory movement, synchronization, kernel launches, communication overhead, and profiling.
- Experience debugging distributed systems across multiple layers.
- Experience with Linux, containers, Kubernetes or comparable orchestration systems, observability, CI/CD, and latency-sensitive production services.
- Ability to design rigorous benchmarks, interpret noisy performance results, identify bottlenecks, and translate findings into production improvements.
- Strong communication and technical leadership skills.
- Bachelor’s degree in Computer Science, Computer Engineering, Electrical Engineering, or a related discipline, or equivalent practical experience.

## Nice-to-haves
- AMD Instinct accelerators and the ROCm ecosystem, including HIP, RCCL, rocprofiler, AMD SMI, AITER, hipBLASLt, and Composable Kernel.
- Deep CUDA experience.
- Contributions to vLLM, SGLang, PyTorch, Triton, TensorRT-LLM, or another open-source ML systems project.
- Prefill-heavy or disaggregated prefill/decode inference architectures.
- KV-cache transfer, prefix caching, continuous batching, chunked prefill, request scheduling, and memory-aware admission control.
- Multi-GPU and multi-node inference, tensor parallelism, pipeline parallelism, expert parallelism, RDMA, collective communication, and failure handling.
- Mixture-of-Experts or multimodal models.
- GPU kernel optimization, operator fusion, graph capture, attention kernels, GEMM tuning, and communication/computation overlap.
- Reduced-precision inference and quantization formats including BF16, FP8, FP4, INT8, and INT4.
- Numerical-comparison, determinism, model-validation, or performance-regression test systems.
- Collaboration with accelerator vendors, framework maintainers, or open-source communities.

## Similar roles

- [Staff Machine Learning Engineer, Traffic Intelligence](https://hotfix.jobs/jobs/cee28c3a-eb64-42f7-a584-8d2b4a0980b8) - Airbnb - Remote - $212k – $265k/yr
- [Member of Technical Staff, Agentic Environments](https://hotfix.jobs/jobs/f3ceb248-2d3e-4b28-a5d1-10891969916e) - Cohere - Remote
- [Staff Machine Learning Infrastructure Engineer, Embedding Platform](https://hotfix.jobs/jobs/9b4387bf-ac16-48ba-810b-2e6a612e3617) - Reddit - Remote - $253k – $355k/yr
- [Senior Staff Machine Learning Engineer, Feed Relevance](https://hotfix.jobs/jobs/a2f9b178-c6cf-49bf-ba93-00688d1027fb) - Reddit - Remote - $266k – $372k/yr
- [Staff Applied Scientist](https://hotfix.jobs/jobs/d518c7a7-07d4-4f06-af79-9c478df769ca) - Garner Health - New York, NY - $300k – $390k/yr

**Apply:** https://hotfix.jobs/jobs/4447537a-cae9-4d19-a17b-2dd8f74005d1
**Canonical:** https://hotfix.jobs/jobs/4447537a-cae9-4d19-a17b-2dd8f74005d1