# Compute Server Platform Architect

**Company:** [Cerebras Systems](https://hotfix.jobs/companies/cerebras-systems)
**Location:** Unspecified
**Role:** Hardware Engineering
**Experience:** 10+ years
**Skills:** x86 server architecture, Linux, numa, cpu microarchitecture, memory bandwidth, pcie, rdma, roce, nvme, cxl, smartnic, dpu, bios, firmware, Python
**Posted:** 2026-02-18

> Owns the server platform architecture for Cerebras AI clusters, translating runtime workloads into CPU, memory, I/O, PCIe, networking, and firmware requirements. The role requires deep Linux and x86 expertise, performance modeling, benchmarking, vendor leadership, and 8–10+ years of relevant industry experience depending on degree.

## Job Description

## Responsibilities
- Own the architecture for all server roles in Cerebras clusters, including server types, configurations, and lifecycle strategy.
- Define and maintain server formulas, including counts and ratios per CS-3 count, cluster size, and workload type, with capacity planning and headroom policy.
- Specify platform configurations covering CPU SKU and core strategy, AMD, Intel, and ARM vendor roadmaps, memory topology, PCIe topology and lane budgeting, NIC selection and placement, and local NVMe policy.
- Translate software and runtime flows into measurable hardware requirements, including CPU utilization, memory bandwidth and latency, bursty I/O patterns, queueing, and concurrency limits.
- Develop performance and scaling models; validate them with microbenchmarks and workload-level experiments; identify bottlenecks and drive cross-stack fixes.
- Define the OS, BIOS, firmware, and driver baseline for each server type.
- Evaluate emerging server technologies, including new CPU generations, memory technologies, CXL, NVMe, and SmartNIC/DPU capabilities, through proof-of-concept evaluations.
- Lead technical vendor engagements with OEMs, ODMs, and component vendors to influence roadmaps, request platform controls, and resolve performance or reliability issues.
- Define qualification and acceptance criteria for performance, stability, and operability, and partner with the Infrastructure Hardware TPM on qualification plans and production adoption.
- Support bring-up and deployment debugging in lab and staging environments; drive root-cause analysis for regressions spanning firmware, drivers, the OS, and runtime behavior.

## Requirements
- PhD in Computer Science or Electrical/Computer Engineering and 8+ years of industry experience, or a master's or bachelor's degree in Computer Science or Electrical Engineering and 10+ years of industry experience.
- 5+ years of experience in server platform architecture, systems performance engineering, or large-scale infrastructure design for AI/ML, HPC, or performance-sensitive distributed systems.
- Deep understanding of x86 server architecture, including CPU microarchitecture, cache hierarchies, NUMA, memory controllers and channels, and memory bandwidth versus latency tradeoffs.
- Strong Linux systems knowledge, including profiling and performance analysis, scheduling and syscall overheads, memory management behavior, and practical tuning methodology.
- Experience with high-performance I/O paths, including NIC behavior, RDMA/RoCE concepts, and NVMe performance characteristics.
- Ability to create capacity and performance models and validate them empirically with rigorous benchmarking plans.
- Experience working directly with vendors and partners to evaluate platforms, resolve issues, and influence roadmaps.
- Strong cross-functional communication skills and ability to drive technical decisions through tradeoff documents and reviews.
- Familiarity with application and systems software, including C, C++, and Python.

## Benefits
- Opportunity to build a breakthrough AI platform beyond the constraints of GPUs.
- Ability to publish and open source cutting-edge AI research.
- Work on one of the fastest AI supercomputers in the world.
- Startup vitality with job stability.
- Non-corporate work culture that respects individual beliefs.
- Continuous learning, growth, and support.

## Similar roles

- [Staff Mechanical Engineer, Matter Compiler](https://hotfix.jobs/jobs/f2770410-ea27-4107-8614-5d3cae6bfd71) - Atomicmachines - Emeryville, CA - $180k – $230k/yr
- [Staff Design Verification Engineer](https://hotfix.jobs/jobs/19fb92b1-d28c-416e-a506-ec8740823813) - Cerebras Systems - Sunnyvale, CA - $250k – $300k/yr
- [Senior Staff Quality Control Engineer](https://hotfix.jobs/jobs/df954d94-948c-4bd9-878a-c244b95b1589) - Crusoe - Abilene, TX
- [Mechanical Engineer - Datacenter](https://hotfix.jobs/jobs/b46e0b44-fc6b-4664-9e45-8056b75239dc) - xAI - Memphis, TN
- [Senior/Staff Engineer : Post Silicon- Bring Up](https://hotfix.jobs/jobs/587653bc-f218-490f-9c45-a297af1de188) - Cerebras Systems - Bengaluru, India - $175k – $275k/yr

**Apply:** https://hotfix.jobs/jobs/9f8eec11-dd0f-4ba3-b97c-3ff7e68be4a5
**Canonical:** https://hotfix.jobs/jobs/9f8eec11-dd0f-4ba3-b97c-3ff7e68be4a5