Senior Software Development Engineer in Test - AI Cluster
The Senior SDET will design and automate reliability, performance, security, and failure testing for large-scale AI clusters spanning distributed software and specialized hardware. The role requires 5+ years of testing experience, strong programming and debugging skills, and familiarity with datacenter infrastructure and cloud technologies.
About the job
Responsibilities
- Design and execute tests for cutting-edge AI infrastructure.
- Define optimized test strategies and methodologies for large-scale distributed systems.
- Break distributed-system challenges into smaller components that can be unit tested.
- Use an automation-first approach to test cluster features, including high availability, failure scenarios, performance, stress, and security.
- Champion cluster security, reliability, 99.9999% uptime, and ease of use through observability.
- Test AI cluster components, including Kubernetes, Prometheus, Grafana, ML wafer-scale accelerators, CPU runtime nodes, SwarmX interconnects, and MemoryX interconnects.
Requirements
- Bachelor's or master's degree in computer science, electrical engineering, AI, data science, or a related field.
- 5+ years of experience testing enterprise software, distributed systems, datacenter hardware, or datacenter software.
- Strong coding skills in Python, Golang, or C/C++.
- Strong debugging skills for large distributed systems, hardware, and software.
- Experience with debugging tools such as pdb, gdb, strace, and network monitors.
- Understanding of operating-system internals, including memory management, file systems, security, and performance.
- Understanding of datacenter layouts and device performance characteristics, including servers, memory, BIOS, PCIe, networking, and storage.
- Experience with cloud technologies such as AWS, Kubernetes, and Docker.
Nice to Have
- Experience with monitoring tools such as Grafana and Prometheus.
- Understanding of and experience with ML model training and inference.
- Understanding of ML hardware accelerators, including GPUs and custom accelerator ASICs.
Benefits
- Opportunity to build a breakthrough AI platform beyond the constraints of GPUs.
- Ability to publish and open-source cutting-edge AI research.
- Work on one of the fastest AI supercomputers in the world.
- Job stability with startup vitality.
- A non-corporate work culture that respects individual beliefs.
Skills
Python, Go, C/C++, Kubernetes, AWS, Docker, Prometheus, Grafana, Pdb, Gdb, Strace, Pcie, Networking, Storage, Ml Inference
Similar jobs
QA Engineering jobsLeads quality engineering for AI Contact Center systems by building scalable automation frameworks and API, UI, capacity, security, and performance tests. Requires 6+ years of software development experience, strong Python or Java skills, and expertise in cloud services, REST testing, and CI integration.
Senior Quality Engineer partnering across product teams to assess risk, build web and API test automation, investigate technical failures, and improve quality practices. Requires strong Playwright and TypeScript experience, exploratory testing skills, CI/CD knowledge, and the ability to debug application code.
Owns automated quality engineering for AI Voice Agent services, spanning UI, backend, APIs, and audio/text interactions. The role requires 6+ years of software engineering or SDET experience, strong programming skills, cloud-native testing expertise, and hands-on experience with LLM systems and evaluation.
Conçoit et maintient des systèmes de test logiciels, électriques et matériels pour valider des vélos et trottinettes en production. Le rôle exige une formation technique, de l’expérience en automatisation, systèmes embarqués, firmware et dépannage directement sur le plancher de fabrication.
Build and own the testing frameworks, integration environments, performance tooling, and resilience capabilities that enable reliable AI inference infrastructure. The role requires strong Go or Python skills, distributed-systems testing experience, Kubernetes and Docker expertise, and the ability to improve test reliability at scale.