# Senior Software Development Engineer in Test - AI Cluster

**Company:** [Cerebras Systems](https://hotfix.jobs/companies/cerebras-systems)
**Location:** Toronto, Canada
**Role:** QA Engineering
**Experience:** 5+ years
**Skills:** Python, Go, C/C++, Kubernetes, AWS, Docker, Prometheus, Grafana, Pdb, Gdb, Strace, Pcie, Networking, Storage, Ml Inference
**Posted:** 2026-07-15

> The Senior SDET will design and automate reliability, performance, security, and failure testing for large-scale AI clusters spanning distributed software and specialized hardware. The role requires 5+ years of testing experience, strong programming and debugging skills, and familiarity with datacenter infrastructure and cloud technologies.

## Job Description

## Responsibilities

- Design and execute tests for cutting-edge AI infrastructure.
- Define optimized test strategies and methodologies for large-scale distributed systems.
- Break distributed-system challenges into smaller components that can be unit tested.
- Use an automation-first approach to test cluster features, including high availability, failure scenarios, performance, stress, and security.
- Champion cluster security, reliability, 99.9999% uptime, and ease of use through observability.
- Test AI cluster components, including Kubernetes, Prometheus, Grafana, ML wafer-scale accelerators, CPU runtime nodes, SwarmX interconnects, and MemoryX interconnects.

## Requirements

- Bachelor's or master's degree in computer science, electrical engineering, AI, data science, or a related field.
- **5+ years of experience** testing enterprise software, distributed systems, datacenter hardware, or datacenter software.
- Strong coding skills in **Python**, **Golang**, or **C/C++**.
- Strong debugging skills for large distributed systems, hardware, and software.
- Experience with debugging tools such as **pdb**, **gdb**, **strace**, and network monitors.
- Understanding of operating-system internals, including memory management, file systems, security, and performance.
- Understanding of datacenter layouts and device performance characteristics, including servers, memory, BIOS, PCIe, networking, and storage.
- Experience with cloud technologies such as **AWS**, **Kubernetes**, and Docker.

## Nice to Have

- Experience with monitoring tools such as **Grafana** and **Prometheus**.
- Understanding of and experience with ML model training and inference.
- Understanding of ML hardware accelerators, including GPUs and custom accelerator ASICs.

## Benefits

- Opportunity to build a breakthrough AI platform beyond the constraints of GPUs.
- Ability to publish and open-source cutting-edge AI research.
- Work on one of the fastest AI supercomputers in the world.
- Job stability with startup vitality.
- A non-corporate work culture that respects individual beliefs.

## Similar jobs

- [Senior Software Development Engineer in Test](https://hotfix.jobs/jobs/9dd83fa4-8f54-447b-a9c3-a931af3991e9) - Dialpad - Kitchener, Canada - CA$149k – CA$174k/yr
- [Senior Quality Engineer](https://hotfix.jobs/jobs/7af0e7ee-6cfd-4b65-9ef1-1bf0e4bc2c7b) - Circle - Remote - $120k – $130k/yr
- [Senior Software Development Engineer in Test](https://hotfix.jobs/jobs/86218182-10dd-447f-a264-9fb9ea0f417d) - Dialpad - Vancouver, Canada - CA$151k – CA$175k/yr
- [Ingénieur de Test en Fabrication, Solutions Urbains](https://hotfix.jobs/jobs/6ce144b9-ade5-418f-b12c-bd65871ba381) - Lyft - Longueuil, Canada
- [Software Engineer - Testing Frameworks](https://hotfix.jobs/jobs/e2c14dc9-5b7e-41e9-819d-bd621388ce52) - Baseten - San Francisco, CA - $165k – $330k/yr

**Apply:** https://hotfix.jobs/jobs/1e0b08fd-3943-45fe-869f-c34e2a53c45c
**Canonical:** https://hotfix.jobs/jobs/1e0b08fd-3943-45fe-869f-c34e2a53c45c