Build and optimize Cerebras’s production GPU inference stack across APIs, vLLM, PyTorch, ROCm, distributed systems, and AMD infrastructure. The role requires 8+ years of software engineering experience, strong C++ and Python skills, and deep expertise in GPU performance, reliability, and model serving.
Salary not listed
Hybrid8+ YOEML Engineering
About the role
Responsibilities
Design, build, deploy, and maintain the complete GPU prefill path across API services, model-serving workers, vLLM, PyTorch, ROCm, GPU nodes, networking, and rack-scale infrastructure.
Establish deployment, upgrade, rollback, health-checking, capacity-management, and failure-recovery practices for an AMD GPU fleet.
Build automation to make driver, firmware, runtime, model, and container compatibility explicit and reproducible.
Define service-level indicators and objectives for GPU-backed inference; improve fault isolation, graceful degradation, automated recovery, incident response, and post-incident remediation.
Optimize time to first token, request throughput, tokens per second per GPU, tail latency, GPU utilization, memory efficiency, and rack-level capacity.
Tune scheduling, continuous batching, prefix caching, KV-cache management, tensor and expert parallelism, request admission, quantization, graph execution, and distributed communication.
Diagnose failures and performance regressions across application code, vLLM, PyTorch, ROCm/HIP, collective communication libraries, kernels, drivers, firmware, networking, and hardware.
Build validation and regression infrastructure for model quality, numerical accuracy, precision changes, quantization, determinism, and software/hardware compatibility.
8+ years of software engineering experience, including substantial individual-contributor ownership of complex production systems.
Experience building, operating, or optimizing production inference systems for large language models, multimodal models, or similarly demanding GPU workloads.
Strong programming ability in C++ and Python, including multithreading, concurrency, memory management, and performance-sensitive software.
Hands-on experience with vLLM, SGLang, TensorRT-LLM, Triton Inference Server, or an equivalent model-serving system.
Strong understanding of GPU execution and performance, including asynchronous execution, memory movement, synchronization, kernel launches, communication overhead, and profiling.
Experience debugging distributed systems across multiple layers.
Experience with Linux, containers, Kubernetes or comparable orchestration systems, observability, CI/CD, and latency-sensitive production services.
Ability to design rigorous benchmarks, interpret noisy performance results, identify bottlenecks, and translate findings into production improvements.
Strong communication and technical leadership skills.
Bachelor’s degree in Computer Science, Computer Engineering, Electrical Engineering, or a related discipline, or equivalent practical experience.
Nice-to-haves
AMD Instinct accelerators and the ROCm ecosystem, including HIP, RCCL, rocprofiler, AMD SMI, AITER, hipBLASLt, and Composable Kernel.
Deep CUDA experience.
Contributions to vLLM, SGLang, PyTorch, Triton, TensorRT-LLM, or another open-source ML systems project.
Prefill-heavy or disaggregated prefill/decode inference architectures.
Architects and operates production machine-learning systems that classify web and API traffic, detect bots and scrapers, and support real-time mitigation at internet edge latency. The role requires 9+ years of applied ML experience in adversarial domains and strong expertise in evaluation, data pipelines, and large-scale systems.
212k – 265k/yrRemote9+ YOEML Engineering
Member of Technical Staff, Agentic Environments
CohereNew York, NY
Build scalable software and tools for frontier model training, research experimentation, and production machine learning systems. The role requires strong Python and distributed-training expertise, experience with ML frameworks and infrastructure, and the ability to optimize and debug large language model systems.
Leads the technical direction of large-scale ML infrastructure for embedding, recommendation, and personalization systems. The role requires 8+ years of ML engineering experience, expertise in deep learning and distributed training, and strong leadership across research, infrastructure, and production deployment.
Leads the technical direction and development of large-scale, GenAI-powered recommendation and feed-ranking systems. Requires 10+ years of industry experience in relevance-driven products, deep expertise in machine learning and recommendations, and strong organizational influence and mentoring skills.
266k – 372k/yrRemote10+ YOEML Engineering
Staff Applied Scientist
Garner HealthNew York, NY
Leads end-to-end development of production algorithmic systems for healthcare, spanning machine learning, optimization, and LLM applications. The player-coach role requires 6+ years of industry experience, strong problem-solving and metrics judgment, and technical leadership of a small team.