Skip to content
Cerebras SystemsCerebras SystemsToronto, Canada

Staff Software Engineer, GPU Inference

Build and optimize Cerebras’s production GPU inference stack across APIs, vLLM, PyTorch, ROCm, distributed systems, and AMD infrastructure. The role requires 8+ years of software engineering experience, strong C++ and Python skills, and deep expertise in GPU performance, reliability, and model serving.

Salary not listed
Hybrid8+ YOEML Engineering

About the role

Responsibilities

  • Design, build, deploy, and maintain the complete GPU prefill path across API services, model-serving workers, vLLM, PyTorch, ROCm, GPU nodes, networking, and rack-scale infrastructure.
  • Establish deployment, upgrade, rollback, health-checking, capacity-management, and failure-recovery practices for an AMD GPU fleet.
  • Build automation to make driver, firmware, runtime, model, and container compatibility explicit and reproducible.
  • Define service-level indicators and objectives for GPU-backed inference; improve fault isolation, graceful degradation, automated recovery, incident response, and post-incident remediation.
  • Optimize time to first token, request throughput, tokens per second per GPU, tail latency, GPU utilization, memory efficiency, and rack-level capacity.
  • Tune scheduling, continuous batching, prefix caching, KV-cache management, tensor and expert parallelism, request admission, quantization, graph execution, and distributed communication.
  • Diagnose failures and performance regressions across application code, vLLM, PyTorch, ROCm/HIP, collective communication libraries, kernels, drivers, firmware, networking, and hardware.
  • Build validation and regression infrastructure for model quality, numerical accuracy, precision changes, quantization, determinism, and software/hardware compatibility.
  • Develop benchmarks, workload-replay tools, profiling automation, release qualification, dashboards, and regression gates.

Requirements

  • 8+ years of software engineering experience, including substantial individual-contributor ownership of complex production systems.
  • Experience building, operating, or optimizing production inference systems for large language models, multimodal models, or similarly demanding GPU workloads.
  • Strong programming ability in C++ and Python, including multithreading, concurrency, memory management, and performance-sensitive software.
  • Hands-on experience with vLLM, SGLang, TensorRT-LLM, Triton Inference Server, or an equivalent model-serving system.
  • Strong understanding of GPU execution and performance, including asynchronous execution, memory movement, synchronization, kernel launches, communication overhead, and profiling.
  • Experience debugging distributed systems across multiple layers.
  • Experience with Linux, containers, Kubernetes or comparable orchestration systems, observability, CI/CD, and latency-sensitive production services.
  • Ability to design rigorous benchmarks, interpret noisy performance results, identify bottlenecks, and translate findings into production improvements.
  • Strong communication and technical leadership skills.
  • Bachelor’s degree in Computer Science, Computer Engineering, Electrical Engineering, or a related discipline, or equivalent practical experience.

Nice-to-haves

  • AMD Instinct accelerators and the ROCm ecosystem, including HIP, RCCL, rocprofiler, AMD SMI, AITER, hipBLASLt, and Composable Kernel.
  • Deep CUDA experience.
  • Contributions to vLLM, SGLang, PyTorch, Triton, TensorRT-LLM, or another open-source ML systems project.
  • Prefill-heavy or disaggregated prefill/decode inference architectures.
  • KV-cache transfer, prefix caching, continuous batching, chunked prefill, request scheduling, and memory-aware admission control.
  • Multi-GPU and multi-node inference, tensor parallelism, pipeline parallelism, expert parallelism, RDMA, collective communication, and failure handling.
  • Mixture-of-Experts or multimodal models.
  • GPU kernel optimization, operator fusion, graph capture, attention kernels, GEMM tuning, and communication/computation overlap.
  • Reduced-precision inference and quantization formats including BF16, FP8, FP4, INT8, and INT4.
  • Numerical-comparison, determinism, model-validation, or performance-regression test systems.
  • Collaboration with accelerator vendors, framework maintainers, or open-source communities.

Skills

C++PythonvLLMPyTorchrocmhipCUDAKubernetesLinuxDockerrcclrdmatensor parallelismquantizationgpu profiling

Similar roles

ML Engineering jobs
Airbnb

Staff Machine Learning Engineer, Traffic Intelligence

AirbnbUnited States

Architects and operates production machine-learning systems that classify web and API traffic, detect bots and scrapers, and support real-time mitigation at internet edge latency. The role requires 9+ years of applied ML experience in adversarial domains and strong expertise in evaluation, data pipelines, and large-scale systems.

212k – 265k/yrRemote9+ YOEML Engineering
Cohere

Member of Technical Staff, Agentic Environments

CohereNew York, NY

Build scalable software and tools for frontier model training, research experimentation, and production machine learning systems. The role requires strong Python and distributed-training expertise, experience with ML frameworks and infrastructure, and the ability to optimize and debug large language model systems.

Salary not listedRemoteML Engineering
Reddit

Staff Machine Learning Infrastructure Engineer, Embedding Platform

RedditUnited States

Leads the technical direction of large-scale ML infrastructure for embedding, recommendation, and personalization systems. The role requires 8+ years of ML engineering experience, expertise in deep learning and distributed training, and strong leadership across research, infrastructure, and production deployment.

253k – 355k/yrRemote8+ YOEML Engineering
Reddit

Senior Staff Machine Learning Engineer, Feed Relevance

RedditUnited States

Leads the technical direction and development of large-scale, GenAI-powered recommendation and feed-ranking systems. Requires 10+ years of industry experience in relevance-driven products, deep expertise in machine learning and recommendations, and strong organizational influence and mentoring skills.

266k – 372k/yrRemote10+ YOEML Engineering
Garner Health

Staff Applied Scientist

Garner HealthNew York, NY

Leads end-to-end development of production algorithmic systems for healthcare, spanning machine learning, optimization, and LLM applications. The player-coach role requires 6+ years of industry experience, strong problem-solving and metrics judgment, and technical leadership of a small team.

300k – 390k/yrHybrid7+ YOEML Engineering