Owns the server platform architecture for Cerebras AI clusters, translating runtime workloads into CPU, memory, I/O, PCIe, networking, and firmware requirements. The role requires deep Linux and x86 expertise, performance modeling, benchmarking, vendor leadership, and 8–10+ years of relevant industry experience depending on degree.
Salary not listed
On-site10+ YOEHardware Engineering
About the role
Responsibilities
Own the architecture for all server roles in Cerebras clusters, including server types, configurations, and lifecycle strategy.
Define and maintain server formulas, including counts and ratios per CS-3 count, cluster size, and workload type, with capacity planning and headroom policy.
Specify platform configurations covering CPU SKU and core strategy, AMD, Intel, and ARM vendor roadmaps, memory topology, PCIe topology and lane budgeting, NIC selection and placement, and local NVMe policy.
Translate software and runtime flows into measurable hardware requirements, including CPU utilization, memory bandwidth and latency, bursty I/O patterns, queueing, and concurrency limits.
Develop performance and scaling models; validate them with microbenchmarks and workload-level experiments; identify bottlenecks and drive cross-stack fixes.
Define the OS, BIOS, firmware, and driver baseline for each server type.
Evaluate emerging server technologies, including new CPU generations, memory technologies, CXL, NVMe, and SmartNIC/DPU capabilities, through proof-of-concept evaluations.
Lead technical vendor engagements with OEMs, ODMs, and component vendors to influence roadmaps, request platform controls, and resolve performance or reliability issues.
Define qualification and acceptance criteria for performance, stability, and operability, and partner with the Infrastructure Hardware TPM on qualification plans and production adoption.
Support bring-up and deployment debugging in lab and staging environments; drive root-cause analysis for regressions spanning firmware, drivers, the OS, and runtime behavior.
Requirements
PhD in Computer Science or Electrical/Computer Engineering and 8+ years of industry experience, or a master's or bachelor's degree in Computer Science or Electrical Engineering and 10+ years of industry experience.
5+ years of experience in server platform architecture, systems performance engineering, or large-scale infrastructure design for AI/ML, HPC, or performance-sensitive distributed systems.
Deep understanding of x86 server architecture, including CPU microarchitecture, cache hierarchies, NUMA, memory controllers and channels, and memory bandwidth versus latency tradeoffs.
Strong Linux systems knowledge, including profiling and performance analysis, scheduling and syscall overheads, memory management behavior, and practical tuning methodology.
Experience with high-performance I/O paths, including NIC behavior, RDMA/RoCE concepts, and NVMe performance characteristics.
Ability to create capacity and performance models and validate them empirically with rigorous benchmarking plans.
Experience working directly with vendors and partners to evaluate platforms, resolve issues, and influence roadmaps.
Strong cross-functional communication skills and ability to drive technical decisions through tradeoff documents and reviews.
Familiarity with application and systems software, including C, C++, and Python.
Benefits
Opportunity to build a breakthrough AI platform beyond the constraints of GPUs.
Ability to publish and open source cutting-edge AI research.
Work on one of the fastest AI supercomputers in the world.
Startup vitality with job stability.
Non-corporate work culture that respects individual beliefs.
Continuous learning, growth, and support.
Skills
x86 server architectureLinuxnumacpu microarchitecturememory bandwidthpcierdmarocenvmecxlsmartnicdpubiosfirmwarePython
Leads the design, development, and qualification of complex precision mechanical systems for a modular micromanufacturing platform. Requires 10+ years of multiphysics mechanical engineering experience, multiple product launches, cross-functional technical leadership, and a bachelor's degree in mechanical engineering or equivalent experience.
180k – 230k/yrOn-site10+ YOEHardware Engineering
Staff Design Verification Engineer
Cerebras SystemsSunnyvale, CA
Leads design verification for Cerebras AI chip architecture, developing scalable SystemVerilog/UVM environments, tests, assertions, and coverage plans while debugging across simulation, emulation, and silicon bring-up. Requires 10+ years of verification experience and strong programming and hardware-debugging skills.
250k – 300k/yrOn-site10+ YOEHardware Engineering
Senior Staff Quality Control Engineer
CrusoeAbilene, TX
Provides senior technical leadership for quality across mission-critical data center construction, focusing on MEP installation, commissioning readiness, defect prevention, and turnover integrity. Requires 10+ years of construction QA/QC or related experience and strong expertise in electrical, mechanical, controls, and commissioning processes.
Salary not listedOn-site10+ YOEHardware Engineering
Mechanical Engineer - Datacenter
xAIMemphis, TN +1
Designs, analyzes, tests, and supports manufacturing of fluid-handling and HVAC systems for advanced data center facilities. The role requires an engineering bachelor’s degree, hands-on fluids experience, and collaboration across engineering, construction, procurement, and manufacturing teams.
Salary not listedOn-site10+ YOEHardware Engineering
Senior/Staff Engineer : Post Silicon- Bring Up
Cerebras SystemsBengaluru, India +2
Leads post-silicon bring-up and optimization of Cerebras Wafer Scale Engines, developing validation infrastructure, diagnostics, testing automation, and performance improvements across hardware and software. Requires 7–10+ years of industry experience and strong ASIC, coding, debugging, and cross-functional collaboration skills.