ML Systems Performance Engineer
Optimizes end-to-end machine learning inference performance across kernels, compilers, systems, and clusters. The role requires strong computer architecture knowledge, experience with performance modeling and profiling, and proficiency in C++ and Python.
Salary not listed
On-site3+ YOEML Engineering
About the job
Responsibilities
- Build performance models at the kernel and end-to-end levels to estimate performance for state-of-the-art and customer machine learning models.
- Optimize and debug kernel microcode and compiler algorithms to improve machine learning model inference speed, throughput, and compute utilization on the Cerebras Wafer Scale Engine.
- Debug and understand runtime performance on the system and cluster.
- Develop tools and infrastructure to visualize performance data collected from the Wafer Scale Engine and compute cluster.
Requirements
- Bachelor's, master's, or PhD in Electrical Engineering or Computer Science.
- Strong background in computer architecture.
- Understanding of low-level deep learning and large language model mathematics.
- Strong analytical and problem-solving skills.
- 3+ years of experience in a relevant domain, such as computer architecture, CPU/GPU performance, kernel optimization, or high-performance computing.
- Experience working with CPU/GPU simulators.
- Experience with performance profiling and debugging on a system pipeline.
- Proficiency with C++ and Python.
Benefits
- Work on a breakthrough AI platform beyond the constraints of GPUs.
- Publish and open-source cutting-edge AI research.
- Work on a high-performance AI supercomputer.
- Startup vitality with job stability.
- A non-corporate work culture that respects individual beliefs.
Skills
C++PythonComputer ArchitectureDeep LearningLLMsKernel OptimizationCompiler AlgorithmsCpu/Gpu PerformanceHigh-Performance ComputingCpu/Gpu SimulatorsPerformance ProfilingPerformance ModelingWafer Scale Engine