Performance Engineer
Performance Engineer optimizes throughput and robustness of large-scale ML distributed systems by solving novel performance issues. Requires significant software engineering experience at supercomputing scale and interest in ML.
About the job
You may be a good fit if you:
- Have significant software engineering or machine learning experience, particularly at supercomputing scale
- Are results-oriented, with a bias towards flexibility and impact
- Pick up slack, even if it goes outside your job description
- Enjoy pair programming (we love to pair!)
- Want to learn more about machine learning research
- Care about the societal impacts of your work
Strong candidates may also have experience with:
- High performance, large-scale ML systems
- GPU/Accelerator programming
- ML framework internals
- OS internals
- Language modeling with transformers
Representative projects:
- Implement low-latency high-throughput sampling for large language models
- Implement GPU kernels to adapt our models to low-precision inference
- Write a custom load-balancing algorithm to optimize serving efficiency
- Build quantitative models of system performance
- Design and implement a fault-tolerant distributed system running with a complex network topology
- Debug kernel-level network latency spikes in a containerized environment
Skills
Machine Learning, Gpu Programming, Distributed Systems, Ml Frameworks, Os Internals, Transformers, Kubernetes, Load Balancing, Performance Optimization, High-Throughput Systems
Similar jobs
DevOps / SRE jobsBuild and operate scalable build systems, CI pipelines, and developer infrastructure for consumer-device software. The role requires 5+ years of engineering experience, expertise with Bazel or comparable build systems, and experience improving CI reliability and performance at scale.
Designs, operates, and improves secure enterprise networks spanning offices, campuses, cloud environments, and connectivity services. The role combines architecture, production operations, troubleshooting, observability, security, and infrastructure automation.
Build and operate the Kubernetes-based cloud and on-premises infrastructure powering large-scale crawling, search, and ML workloads. The role requires 5+ years in DevOps, platform engineering, or cloud infrastructure, with strong Kubernetes, cloud, Docker, Terraform, and distributed-systems experience.
Build and operate an AI-first CI/CD and agent-operations platform for Salesforce and custom GTM applications. The role focuses on governed releases, approval workflows, observability, rollback, sandboxing, and SOX-compliant auditability.
Owns reliability standards, incident management, observability, failure testing, and automation for a high-throughput AI infrastructure platform. The role requires deep Linux, networking, software, cloud-native, and distributed-systems experience, along with the ability to influence teams across the organization.