Skip to content
Cerebras SystemsCerebras SystemsUnited States

Software Engineer, GPU Inference

Build and optimize Cerebras’s production GPU prefill and inference stack across APIs, serving runtimes, ROCm, distributed systems, and hardware. The role requires 5+ years of software engineering experience, strong C++ and Python skills, and hands-on experience operating high-performance model-serving systems.

Salary not listed
On-site5+ YOEML Engineering

About the role

Responsibilities

  • Productionize the GPU inference stack across custom inference APIs, model-serving workers, vLLM, PyTorch, ROCm, GPU nodes, networking, and rack-scale infrastructure.
  • Establish deployment, upgrade, rollback, health-checking, capacity-management, and failure-recovery practices for AMD GPU fleets.
  • Build automation to make driver, firmware, runtime, model, and container compatibility explicit and reproducible.
  • Define service-level indicators and objectives for GPU-backed inference.
  • Improve fault isolation, graceful degradation, automated recovery, incident response, and post-incident remediation.
  • Optimize time to first token, request throughput, tokens per second per GPU, tail latency, GPU utilization, memory efficiency, and rack-level capacity.
  • Tune scheduling, continuous batching, prefix caching, KV-cache management, tensor and expert parallelism, request admission, quantization, graph execution, and distributed communication.
  • Diagnose failures and performance regressions across application code, vLLM, PyTorch, ROCm/HIP, collective communication libraries, kernels, drivers, firmware, networking, and hardware.
  • Build validation and regression infrastructure for model quality, numerical accuracy, precision changes, quantization, determinism, and software/hardware compatibility.
  • Develop benchmarks, workload replay tools, profiling automation, release qualification, dashboards, and regression gates.

Requirements

  • 5+ years of software engineering experience, including substantial individual-contributor ownership of complex production systems.
  • Experience building, operating, or optimizing production inference systems for large language models, multimodal models, or comparable GPU workloads.
  • Strong C++ and Python programming skills, including multithreading, concurrency, memory management, and performance-sensitive software.
  • Experience with a high-performance model-serving framework such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, or an equivalent system.
  • Strong understanding of GPU execution and performance, including asynchronous execution, memory movement, synchronization, kernel launches, communication overhead, and profiling.
  • Experience debugging distributed systems across multiple layers.
  • Experience with Linux, containers, Kubernetes or comparable orchestration, observability, CI/CD, and latency-sensitive production services.
  • Ability to design rigorous benchmarks, interpret noisy performance results, identify bottlenecks, and translate findings into production improvements.
  • Strong communication and technical leadership skills.
  • Bachelor’s degree in Computer Science, Computer Engineering, Electrical Engineering, or a related discipline, or equivalent practical experience.

Nice-to-haves

  • AMD Instinct accelerators and the ROCm ecosystem, including HIP, RCCL, rocprofiler, AMD SMI, AITER, hipBLASLt, and Composable Kernel.
  • Deep CUDA experience.
  • Contributions to vLLM, SGLang, PyTorch, Triton, TensorRT-LLM, or another open-source ML systems project.
  • Prefill-heavy or disaggregated prefill/decode inference architectures.
  • KV-cache transfer, prefix caching, continuous batching, chunked prefill, request scheduling, and memory-aware admission control.
  • Multi-GPU and multi-node inference, tensor parallelism, pipeline parallelism, expert parallelism, RDMA, collective communication, and failure handling.
  • Mixture-of-Experts or multimodal model optimization.
  • GPU kernel optimization, operator fusion, graph capture, attention kernels, GEMM tuning, and communication/computation overlap.
  • Reduced-precision inference and quantization formats including BF16, FP8, FP4, INT8, or INT4.
  • Numerical-comparison, determinism, model-validation, or performance-regression test systems.
  • Collaboration with accelerator vendors, framework maintainers, or open-source communities.

Benefits

  • Work on a breakthrough AI platform and one of the fastest AI supercomputers in the world.
  • Opportunities to publish and open-source AI research.
  • Startup vitality with job stability.
  • A non-corporate work culture that respects individual beliefs.
  • Continuous learning, growth, and team support.

Skills

C++PythonvLLMPyTorchrocmhipCUDAKubernetesLinuxDockerDistributed Systemsrcclrdmatensor parallelismquantization

Similar roles

ML Engineering jobs
Clay

Machine Learning Engineer

ClayNew York, NY +1

Build and ship production machine-learning systems that learn from customer data and behavior, including recommendations, LLM-powered features, evaluation systems, and ML infrastructure. The role requires 5+ years of ML engineering or ML-heavy software engineering experience and strong production systems expertise.

170k – 300k/yrHybrid5+ YOEML Engineering
Confido Legal

Applied AI/ML Engineer

Confido LegalNew York, NY

Build and productionize applied AI/ML systems for document understanding, agentic workflows, and demand forecasting using rich, messy enterprise data. The role requires 3+ years of production AI/ML experience, strong evaluation and monitoring practices, and a STEM master’s degree.

200k – 250k/yrOn-site5+ YOEML Engineering
Tulip

AI Platform Engineer

TulipSomerville, MA

Build and deploy production LLM agents and the internal agentic AI platform, partnering with business teams to solve operational problems and drive adoption. The role requires 5+ years of software experience, full-stack development skills, and hands-on experience with LLMs, RAG, and agent frameworks.

Salary not listedHybrid5+ YOEML Engineering
Garner Health

Applied Scientist III

Garner HealthNew York, NY

Build and deploy algorithmic systems for high-impact healthcare problems, choosing among machine learning, optimization, heuristics, and hybrid approaches. The role requires 4+ years of relevant industry experience, strong applied problem-solving and evaluation skills, and fluency in modern ML tooling.

236k – 260k/yrOn-site4+ YOEML Engineering
Applied Intuition

Software Engineer - Prediction and Planning ML

Applied IntuitionSunnyvale, CA

Develop and deploy ML-first behavior prediction and planning systems for autonomous vehicles, forecasting the motion and interactions of road users. Requires a bachelor's degree, deep learning lifecycle expertise, and at least three years of production software experience with C++ or Python.

151k – 258k/yrOn-site3+ YOEML Engineering