# Senior Runtime Engineer

**Company:** [Cerebras Systems](https://hotfix.jobs/companies/cerebras-systems)
**Location:** Unspecified
**Role:** Backend Engineering
**Experience:** 5+ years
**Skills:** C++, C, Python, PyTorch, Distributed Systems, Networking, Inter-Process Communication, Multithreading, Memory Management, Performance Optimization, Data Structures, Concurrency, Profiling, Compiler Internals, Hpc
**Posted:** 2025-10-28

> Design and optimize distributed runtime software for large-scale AI training and inference across heterogeneous clusters. The role requires 3+ years of high-performance or distributed systems experience, strong C/C++ skills, and expertise in concurrency, memory management, and performance optimization.

## Job Description

## Responsibilities
- Design and implement distributed runtime components to efficiently manage large-scale execution workloads.
- Develop and optimize high-performance data and communication pipelines that fully utilize CPU, memory, storage, and network resources.
- Enable scalable execution across multiple compute nodes, ensuring high concurrency and minimal bottlenecks.
- Collaborate closely with ML and compiler teams to integrate new model architectures, training regimes, and hardware-specific optimizations.
- Diagnose and resolve complex performance issues across the software stack using profiling and instrumentation tools.
- Contribute to system design, architecture reviews, and roadmap planning for large-scale AI workloads.

## Requirements
- 3+ years of experience developing high-performance or distributed systems software.
- Strong programming skills in C/C++, with expertise in multithreading, memory management, and performance optimization.
- Experience with distributed systems, networking, or inter-process communication.
- Solid understanding of data structures, concurrency, and system-level resource management, including CPU, I/O, and memory.
- Proven ability to debug, profile, and optimize code across scales, from threads to clusters.
- Bachelor's, master's, or equivalent experience in Computer Science, Electrical Engineering, or a related field.

## Nice-to-Haves
- Familiarity with machine learning training or inference pipelines, especially distributed training and large-model scaling.
- Exposure to Python and PyTorch, particularly for model training or performance tuning.
- Experience with compiler internals, custom hardware interfaces, or low-level protocol design.
- Prior work on high-performance clusters, HPC systems, or custom hardware/software co-design.
- Deep curiosity about unlocking new levels of performance for large-scale AI workloads.

## Benefits
- Work on a breakthrough AI platform beyond the constraints of GPUs.
- Opportunities to publish and open-source cutting-edge AI research.
- Work on one of the fastest AI supercomputers in the world.
- Job stability with startup vitality.
- A non-corporate work culture that respects individual beliefs.
- Continuous learning, growth, and support.

## Similar jobs

- [Senior Software Engineer](https://hotfix.jobs/jobs/493ff3e3-27c5-4773-84f0-525cc87f1091) - ZoomInfo - Remote - $140k – $220k/yr
- [Senior Software Engineer](https://hotfix.jobs/jobs/c42e6986-1722-4cdc-8d1d-b4c56b919eaa) - Pindrop - Remote - $130k – $170k/yr
- [Senior Backend Engineer](https://hotfix.jobs/jobs/c0691142-d546-4d35-bda4-798fd0dae1b1) - Mintlify - San Francisco, CA - $190k – $265k/yr
- [Lead Distributed Systems Engineer](https://hotfix.jobs/jobs/4f810a88-8f54-43b0-9138-f407165fadaa) - Cape - New York, NY - $250k – $325k/yr
- [Senior Software Engineer, Device Identity](https://hotfix.jobs/jobs/bd779c69-a026-4775-949d-bff228853342) - Okta - Toronto, Canada - CA$136k – CA$187k/yr

**Apply:** https://hotfix.jobs/jobs/76759b6e-8ef0-4b4b-9692-eb9946304d16
**Canonical:** https://hotfix.jobs/jobs/76759b6e-8ef0-4b4b-9692-eb9946304d16