Engineering Manager, ML Acceleration
Engineering Manager leading teams to optimize ML inference and training systems, remove bottlenecks, and scale compute resources efficiently. Requires 1+ years management experience in performance/distributed systems and ML/AI background.
About the job
Responsibilities
- Provide front-line leadership of engineering efforts to improve model performance and scale our inference and training systems
- Become familiar with the team’s technical stack enough to make targeted contributions as an individual contributor
- Manage day-to-day execution of the team's work
- Prioritize the team’s work and manage projects in a highly dynamic, fast paced environment
- Coach and support your reports in understanding, and pursuing, their professional growth
- Maintain a deep understanding of the team's technical work and its implications for AI safety
You may be a good fit if you
- Have 1+ years of management experience in a technical environment, particularly performance or distributed systems
- Have a background in machine learning, AI, or a similar related technical field
- Are deeply interested in the potential transformative effects of advanced AI systems and are committed to ensuring their safe development
- Excel at building strong relationships with stakeholders at all levels
- Are a quick learner, capable of understanding and contributing to discussions on complex technical topics
- Have experience managing teams through periods of rapid growth and change
- Are a quick study: this team sits at the intersection of a large number of different complex technical systems that you’ll need to understand (at a high level of abstraction) to be effective
Strong candidates may also have experience with
- High performance, large-scale ML systems
- GPU/Accelerator programming
- ML framework internals
- OS internals
- Language modeling with transformers
Annual Salary: $500,000 — $850,000 USD
Skills
Machine Learning, Distributed Systems, Performance Optimization, Gpu Programming, Ml Frameworks, Transformers, Os Internals, Large-Scale Ml Systems, Inference Systems, Training Systems
Similar jobs
Engineering Management jobsLeads teams building and operating the control plane for Anthropic’s large-scale inference fleet, improving routing, capacity, performance, reliability, and cost. Requires deep production-systems expertise, engineering management experience, and strong cross-functional leadership.
Leads the engineering team responsible for ChatGPT Search Infrastructure, setting technical direction for scalable, low-latency search platforms and integrations. Requires experience managing senior engineers and deep expertise in distributed systems, online serving, experimentation, and AI-powered products.
Leads the client platform engineering organization responsible for secure, reliable endpoint services across major operating systems. The role combines people management, platform strategy, architecture oversight, operational excellence, and cross-functional partnership.
Leads and grows the engineering team building AI-native artifacts such as documents, spreadsheets, slides, and dashboards. The role sets technical direction across product, infrastructure, rendering, storage, reliability, and model integration while partnering closely with research and product teams.
Leads the Data Platform Engineering team, building scalable batch and streaming infrastructure for complex healthcare data while guiding architecture, operations, and team development. Requires 5+ years of engineering experience and 2+ years managing engineering teams.