Build and optimize Cerebras’s production GPU inference stack across APIs, vLLM, PyTorch, ROCm, distributed systems, and AMD infrastructure. The role requires 8+ years of software engineering experience, strong C++ and Python skills, and deep expertise in GPU performance, reliability, and model serving.
Salary not listedHybrid8+ YOEML Engineering
CoDesign & NextGen Performance Engineer
Cerebras SystemsSunnyvale, CA
Characterize, analyze, and optimize performance of state-of-the-art AI models on Cerebras' wafer-scale hardware. Build performance models, optimize kernels and compilers, debug runtime behavior, and develop visualization tools to influence next-gen AI architecture.
Salary not listedOn-site3+ YOEML Engineering
AI Engineer, Model Quality and Performance
Cerebras SystemsSunnyvale, CA
Own model quality and performance for Cerebras inference by building AI agent-driven eval suites, automating benchmarking, and creating customer-specific tooling. Requires strong AI agent experience and tooling intuition.
Salary not listedOn-siteML Engineering
Senior Performance Engineer, Inference
Cerebras SystemsSunnyvale, CA
Senior Performance Engineer benchmarks Cerebras inference performance against competitors for real workloads, measuring metrics like tokens/second and TCO, while modeling competitor pricing to inform sales strategies. Requires 5+ years in ML systems and deep expertise in open-source inference stacks and transformer optimizations.
Salary not listedOn-site5+ YOEML Engineering
Principal ML Investigator
Cerebras SystemsSunnyvale, CA
Leads new ML research team focusing on post-training/RL, dataset optimization, LLM pretraining, sparsity, and domain-specific agents. Adapts algorithms for Cerebras hardware, builds teams, and collaborates on hardware/software design. Requires PhD and ML leadership experience.
Salary not listedOn-siteML Engineering
Advanced Technology: R&D Engineer - AI/ML, HPC
Cerebras SystemsSunnyvale, CA
Designs and implements AI/ML and scientific computing workloads on wafer-scale hardware to set performance benchmarks. Leads algorithm-hardware co-design, builds performance models, and contributes to technology roadmap with publications in top venues. PhD preferred in CS, Engineering, or related field.
Salary not listedOn-siteML Engineering
Applied Machine Learning Research Scientist
Cerebras SystemsUnited States
Build and optimize scalable machine learning systems for LLM pretraining, fine-tuning, alignment, and evaluation. The role requires 4+ years of ML systems experience, strong Python and PyTorch skills, and familiarity with transformers; experience with LLMs, reinforcement learning, and distributed training is preferred.
Salary not listedOn-site4+ YOEML Engineering
ML Software Tool Development Engineer
Cerebras SystemsUnited States
Build system-level debugging, validation, observability, and anomaly-analysis tooling for Cerebras’s AI hardware and software stack. The role requires strong C++ and Python skills, experience debugging complex hardware/software systems, and familiarity with compilers, runtimes, or high-performance computing.
Salary not listedOn-siteML Engineering
Software Engineer, GPU Inference
Cerebras SystemsUnited States
Build and optimize Cerebras’s production GPU prefill and inference stack across APIs, serving runtimes, ROCm, distributed systems, and hardware. The role requires 5+ years of software engineering experience, strong C++ and Python skills, and hands-on experience operating high-performance model-serving systems.
Salary not listedOn-site5+ YOEML Engineering
Search
Location
9 jobs
Job results
Staff Software Engineer, GPU Inference
Cerebras SystemsToronto, Canada +1
Build and optimize Cerebras’s production GPU inference stack across APIs, vLLM, PyTorch, ROCm, distributed systems, and AMD infrastructure. The role requires 8+ years of software engineering experience, strong C++ and Python skills, and deep expertise in GPU performance, reliability, and model serving.
Salary not listedHybrid8+ YOEML Engineering
CoDesign & NextGen Performance Engineer
Cerebras SystemsSunnyvale, CA
Characterize, analyze, and optimize performance of state-of-the-art AI models on Cerebras' wafer-scale hardware. Build performance models, optimize kernels and compilers, debug runtime behavior, and develop visualization tools to influence next-gen AI architecture.
Salary not listedOn-site3+ YOEML Engineering
AI Engineer, Model Quality and Performance
Cerebras SystemsSunnyvale, CA
Own model quality and performance for Cerebras inference by building AI agent-driven eval suites, automating benchmarking, and creating customer-specific tooling. Requires strong AI agent experience and tooling intuition.
Salary not listedOn-siteML Engineering
Senior Performance Engineer, Inference
Cerebras SystemsSunnyvale, CA
Senior Performance Engineer benchmarks Cerebras inference performance against competitors for real workloads, measuring metrics like tokens/second and TCO, while modeling competitor pricing to inform sales strategies. Requires 5+ years in ML systems and deep expertise in open-source inference stacks and transformer optimizations.
Salary not listedOn-site5+ YOEML Engineering
Principal ML Investigator
Cerebras SystemsSunnyvale, CA
Leads new ML research team focusing on post-training/RL, dataset optimization, LLM pretraining, sparsity, and domain-specific agents. Adapts algorithms for Cerebras hardware, builds teams, and collaborates on hardware/software design. Requires PhD and ML leadership experience.
Salary not listedOn-siteML Engineering
Advanced Technology: R&D Engineer - AI/ML, HPC
Cerebras SystemsSunnyvale, CA
Designs and implements AI/ML and scientific computing workloads on wafer-scale hardware to set performance benchmarks. Leads algorithm-hardware co-design, builds performance models, and contributes to technology roadmap with publications in top venues. PhD preferred in CS, Engineering, or related field.
Salary not listedOn-siteML Engineering
Applied Machine Learning Research Scientist
Cerebras SystemsUnited States
Build and optimize scalable machine learning systems for LLM pretraining, fine-tuning, alignment, and evaluation. The role requires 4+ years of ML systems experience, strong Python and PyTorch skills, and familiarity with transformers; experience with LLMs, reinforcement learning, and distributed training is preferred.
Salary not listedOn-site4+ YOEML Engineering
ML Software Tool Development Engineer
Cerebras SystemsUnited States
Build system-level debugging, validation, observability, and anomaly-analysis tooling for Cerebras’s AI hardware and software stack. The role requires strong C++ and Python skills, experience debugging complex hardware/software systems, and familiarity with compilers, runtimes, or high-performance computing.
Salary not listedOn-siteML Engineering
Get new-job notifications on iOS
Hotfix on iOS
Get a push summary when new jobs match your saved alerts.
Software Engineer, GPU Inference
Cerebras SystemsUnited States
Build and optimize Cerebras’s production GPU prefill and inference stack across APIs, serving runtimes, ROCm, distributed systems, and hardware. The role requires 5+ years of software engineering experience, strong C++ and Python skills, and hands-on experience operating high-performance model-serving systems.