# Software Engineer, GPU Inference

**Company:** [Cerebras Systems](https://hotfix.jobs/companies/cerebras-systems)
**Location:** Unspecified
**Role:** ML Engineering
**Experience:** 5+ years
**Skills:** C++, Python, vLLM, PyTorch, rocm, hip, CUDA, Kubernetes, Linux, Docker, Distributed Systems, rccl, rdma, tensor parallelism, quantization
**Posted:** 2025-11-25

> Build and optimize Cerebras’s production GPU prefill and inference stack across APIs, serving runtimes, ROCm, distributed systems, and hardware. The role requires 5+ years of software engineering experience, strong C++ and Python skills, and hands-on experience operating high-performance model-serving systems.

## Job Description

## Responsibilities
- Productionize the GPU inference stack across custom inference APIs, model-serving workers, vLLM, PyTorch, ROCm, GPU nodes, networking, and rack-scale infrastructure.
- Establish deployment, upgrade, rollback, health-checking, capacity-management, and failure-recovery practices for AMD GPU fleets.
- Build automation to make driver, firmware, runtime, model, and container compatibility explicit and reproducible.
- Define service-level indicators and objectives for GPU-backed inference.
- Improve fault isolation, graceful degradation, automated recovery, incident response, and post-incident remediation.
- Optimize time to first token, request throughput, tokens per second per GPU, tail latency, GPU utilization, memory efficiency, and rack-level capacity.
- Tune scheduling, continuous batching, prefix caching, KV-cache management, tensor and expert parallelism, request admission, quantization, graph execution, and distributed communication.
- Diagnose failures and performance regressions across application code, vLLM, PyTorch, ROCm/HIP, collective communication libraries, kernels, drivers, firmware, networking, and hardware.
- Build validation and regression infrastructure for model quality, numerical accuracy, precision changes, quantization, determinism, and software/hardware compatibility.
- Develop benchmarks, workload replay tools, profiling automation, release qualification, dashboards, and regression gates.

## Requirements
- 5+ years of software engineering experience, including substantial individual-contributor ownership of complex production systems.
- Experience building, operating, or optimizing production inference systems for large language models, multimodal models, or comparable GPU workloads.
- Strong C++ and Python programming skills, including multithreading, concurrency, memory management, and performance-sensitive software.
- Experience with a high-performance model-serving framework such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, or an equivalent system.
- Strong understanding of GPU execution and performance, including asynchronous execution, memory movement, synchronization, kernel launches, communication overhead, and profiling.
- Experience debugging distributed systems across multiple layers.
- Experience with Linux, containers, Kubernetes or comparable orchestration, observability, CI/CD, and latency-sensitive production services.
- Ability to design rigorous benchmarks, interpret noisy performance results, identify bottlenecks, and translate findings into production improvements.
- Strong communication and technical leadership skills.
- Bachelor’s degree in Computer Science, Computer Engineering, Electrical Engineering, or a related discipline, or equivalent practical experience.

## Nice-to-haves
- AMD Instinct accelerators and the ROCm ecosystem, including HIP, RCCL, rocprofiler, AMD SMI, AITER, hipBLASLt, and Composable Kernel.
- Deep CUDA experience.
- Contributions to vLLM, SGLang, PyTorch, Triton, TensorRT-LLM, or another open-source ML systems project.
- Prefill-heavy or disaggregated prefill/decode inference architectures.
- KV-cache transfer, prefix caching, continuous batching, chunked prefill, request scheduling, and memory-aware admission control.
- Multi-GPU and multi-node inference, tensor parallelism, pipeline parallelism, expert parallelism, RDMA, collective communication, and failure handling.
- Mixture-of-Experts or multimodal model optimization.
- GPU kernel optimization, operator fusion, graph capture, attention kernels, GEMM tuning, and communication/computation overlap.
- Reduced-precision inference and quantization formats including BF16, FP8, FP4, INT8, or INT4.
- Numerical-comparison, determinism, model-validation, or performance-regression test systems.
- Collaboration with accelerator vendors, framework maintainers, or open-source communities.

## Benefits
- Work on a breakthrough AI platform and one of the fastest AI supercomputers in the world.
- Opportunities to publish and open-source AI research.
- Startup vitality with job stability.
- A non-corporate work culture that respects individual beliefs.
- Continuous learning, growth, and team support.

## Similar roles

- [Machine Learning Engineer](https://hotfix.jobs/jobs/4b3b83f5-9df1-4b0d-9e62-eb79d4706711) - Clay - New York, NY - $170k – $300k/yr
- [Applied AI/ML Engineer](https://hotfix.jobs/jobs/0c7d59b1-78f2-4eed-9c9a-3b25b3196166) - Confido Legal - New York, NY - $200k – $250k/yr
- [AI Platform Engineer](https://hotfix.jobs/jobs/300f336b-42cc-4c14-ad80-ea9261af8d6b) - Tulip - Somerville, MA
- [Applied Scientist III](https://hotfix.jobs/jobs/00c9850b-9666-4032-aa7c-5fd2253c5ef7) - Garner Health - New York, NY - $236k – $260k/yr
- [Software Engineer - Prediction and Planning ML](https://hotfix.jobs/jobs/21b9c778-e1ae-4695-b26d-fec68ea8a8cc) - Applied Intuition - Sunnyvale, CA - $151k – $258k/yr

**Apply:** https://hotfix.jobs/jobs/2ada735f-2b01-47af-8b55-ee5f33f4efb1
**Canonical:** https://hotfix.jobs/jobs/2ada735f-2b01-47af-8b55-ee5f33f4efb1