Build and optimize Cerebras’s production GPU prefill and inference stack across APIs, serving runtimes, ROCm, distributed systems, and hardware. The role requires 5+ years of software engineering experience, strong C++ and Python skills, and hands-on experience operating high-performance model-serving systems.
Salary not listed
On-site5+ YOEML Engineering
About the role
Responsibilities
Productionize the GPU inference stack across custom inference APIs, model-serving workers, vLLM, PyTorch, ROCm, GPU nodes, networking, and rack-scale infrastructure.
Establish deployment, upgrade, rollback, health-checking, capacity-management, and failure-recovery practices for AMD GPU fleets.
Build automation to make driver, firmware, runtime, model, and container compatibility explicit and reproducible.
Define service-level indicators and objectives for GPU-backed inference.
Optimize time to first token, request throughput, tokens per second per GPU, tail latency, GPU utilization, memory efficiency, and rack-level capacity.
Tune scheduling, continuous batching, prefix caching, KV-cache management, tensor and expert parallelism, request admission, quantization, graph execution, and distributed communication.
Diagnose failures and performance regressions across application code, vLLM, PyTorch, ROCm/HIP, collective communication libraries, kernels, drivers, firmware, networking, and hardware.
Build validation and regression infrastructure for model quality, numerical accuracy, precision changes, quantization, determinism, and software/hardware compatibility.
5+ years of software engineering experience, including substantial individual-contributor ownership of complex production systems.
Experience building, operating, or optimizing production inference systems for large language models, multimodal models, or comparable GPU workloads.
Strong C++ and Python programming skills, including multithreading, concurrency, memory management, and performance-sensitive software.
Experience with a high-performance model-serving framework such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, or an equivalent system.
Strong understanding of GPU execution and performance, including asynchronous execution, memory movement, synchronization, kernel launches, communication overhead, and profiling.
Experience debugging distributed systems across multiple layers.
Experience with Linux, containers, Kubernetes or comparable orchestration, observability, CI/CD, and latency-sensitive production services.
Ability to design rigorous benchmarks, interpret noisy performance results, identify bottlenecks, and translate findings into production improvements.
Strong communication and technical leadership skills.
Bachelor’s degree in Computer Science, Computer Engineering, Electrical Engineering, or a related discipline, or equivalent practical experience.
Nice-to-haves
AMD Instinct accelerators and the ROCm ecosystem, including HIP, RCCL, rocprofiler, AMD SMI, AITER, hipBLASLt, and Composable Kernel.
Deep CUDA experience.
Contributions to vLLM, SGLang, PyTorch, Triton, TensorRT-LLM, or another open-source ML systems project.
Prefill-heavy or disaggregated prefill/decode inference architectures.
Build and ship production machine-learning systems that learn from customer data and behavior, including recommendations, LLM-powered features, evaluation systems, and ML infrastructure. The role requires 5+ years of ML engineering or ML-heavy software engineering experience and strong production systems expertise.
170k – 300k/yrHybrid5+ YOEML Engineering
Applied AI/ML Engineer
Confido LegalNew York, NY
Build and productionize applied AI/ML systems for document understanding, agentic workflows, and demand forecasting using rich, messy enterprise data. The role requires 3+ years of production AI/ML experience, strong evaluation and monitoring practices, and a STEM master’s degree.
200k – 250k/yrOn-site5+ YOEML Engineering
AI Platform Engineer
TulipSomerville, MA
Build and deploy production LLM agents and the internal agentic AI platform, partnering with business teams to solve operational problems and drive adoption. The role requires 5+ years of software experience, full-stack development skills, and hands-on experience with LLMs, RAG, and agent frameworks.
Salary not listedHybrid5+ YOEML Engineering
Applied Scientist III
Garner HealthNew York, NY
Build and deploy algorithmic systems for high-impact healthcare problems, choosing among machine learning, optimization, heuristics, and hybrid approaches. The role requires 4+ years of relevant industry experience, strong applied problem-solving and evaluation skills, and fluency in modern ML tooling.
236k – 260k/yrOn-site4+ YOEML Engineering
Software Engineer - Prediction and Planning ML
Applied IntuitionSunnyvale, CA
Develop and deploy ML-first behavior prediction and planning systems for autonomous vehicles, forecasting the motion and interactions of road users. Requires a bachelor's degree, deep learning lifecycle expertise, and at least three years of production software experience with C++ or Python.