Skip to content
Cerebras SystemsCerebras SystemsBengaluru, India

ML Systems Performance Engineer

Optimizes end-to-end machine learning inference performance across kernels, compilers, systems, and clusters. The role requires strong computer architecture knowledge, experience with performance modeling and profiling, and proficiency in C++ and Python.

Salary not listed
On-site3+ YOEML Engineering

About the job

Responsibilities

  • Build performance models at the kernel and end-to-end levels to estimate performance for state-of-the-art and customer machine learning models.
  • Optimize and debug kernel microcode and compiler algorithms to improve machine learning model inference speed, throughput, and compute utilization on the Cerebras Wafer Scale Engine.
  • Debug and understand runtime performance on the system and cluster.
  • Develop tools and infrastructure to visualize performance data collected from the Wafer Scale Engine and compute cluster.

Requirements

  • Bachelor's, master's, or PhD in Electrical Engineering or Computer Science.
  • Strong background in computer architecture.
  • Understanding of low-level deep learning and large language model mathematics.
  • Strong analytical and problem-solving skills.
  • 3+ years of experience in a relevant domain, such as computer architecture, CPU/GPU performance, kernel optimization, or high-performance computing.
  • Experience working with CPU/GPU simulators.
  • Experience with performance profiling and debugging on a system pipeline.
  • Proficiency with C++ and Python.

Benefits

  • Work on a breakthrough AI platform beyond the constraints of GPUs.
  • Publish and open-source cutting-edge AI research.
  • Work on a high-performance AI supercomputer.
  • Startup vitality with job stability.
  • A non-corporate work culture that respects individual beliefs.

Skills

C++PythonComputer ArchitectureDeep LearningLLMsKernel OptimizationCompiler AlgorithmsCpu/Gpu PerformanceHigh-Performance ComputingCpu/Gpu SimulatorsPerformance ProfilingPerformance ModelingWafer Scale Engine