Performance Modeling Lead
Leads a team building performance modeling frameworks to evaluate AI infrastructure tradeoffs in compute, memory, networking, and topology. Guides architectural decisions, influences vendor designs, and partners with ML/systems teams; requires deep AI workload and systems expertise.
About the job
Key Responsibilities
- Build and own a performance modeling framework/toolchain to evaluate AI systems across multiple levels of abstraction.
- Analyze and quantify architectural tradeoffs across compute, memory, networking, storage, and system topology.
- Develop performance models to guide decisions on: scale-up vs. scale-out architectures, interconnect and network design, memory hierarchy and system balance.
- Translate modeling outputs into clear recommendations for internal teams and external hardware vendors.
- Influence reference designs and vendor roadmaps through data-driven insights.
- Partner closely with machine learning, systems, and hardware teams to understand workload characteristics and requirements.
- Lead and grow a small team (2–3 engineers), setting technical direction and maintaining high standards for modeling rigor.
- Continuously improve modeling fidelity by validating against real system behavior and measurements.
Qualifications
- Experience owning or building performance modeling frameworks used to drive real system design decisions.
- Deep knowledge of AI/ML workloads, including training and/or inference at scale.
- Understand system-level tradeoffs across compute, memory, and networking in large-scale distributed systems.
- Comfortable working across abstraction layers—from workload behavior to hardware implementation.
- Experience using modeling (analytical or simulation) to inform architectural decisions.
- Can operate in ambiguous problem spaces and turn open-ended questions into structured analysis.
- Communicate clearly and influence both internal teams and external partners.
Preferred Skills
- Experience working with hardware vendors (ODM/JDM, silicon, networking).
- Background in data center infrastructure or hyperscale systems.
- Familiarity with accelerators (GPUs/ASICs) and interconnects (e.g., NVLink, InfiniBand, Ethernet).
- Experience influencing hardware roadmaps or reference architectures.
- Prior experience leading or mentoring engineers.
Skills
Performance Modeling, Ai/Ml Workloads, Gpus, Asics, Nvlink, InfiniBand, Ethernet, Distributed Systems, Scale-Up Architectures, Scale-Out Architectures
Similar jobs
Engineering Management jobsLeads and develops an applied AI engineering team delivering reliable, customer-facing agents, integrations, and scientific infrastructure for pharma and biotech organizations. The role combines hands-on engineering, partner ownership, cross-functional product influence, and responsible AI deployment in life sciences.
Leads an engineering team building AI-native internal productivity tools, partnering with business stakeholders and owning reliable cloud services. Requires 6+ years of engineering experience, 2+ years managing engineering teams, and experience with web applications, APIs, and cloud infrastructure.
Leads engineering teams building Garner’s core healthcare product, including data ingestion systems and new product experiences with potential AI enhancements. Requires 6+ years of engineering experience, 2+ years of engineering management, and familiarity with web applications, APIs, and AWS.
Leads a player-coach engineering team building reliable, scalable multimodal data products across image, video, audio, and document workflows. The role combines hands-on backend and distributed-systems architecture with team management, recruiting, and product delivery.
Leads the team and platform responsible for Gusto’s agentic coding environments, AI code review, and supporting CI and security infrastructure. The role requires substantial engineering leadership experience, hands-on use of agentic coding tools, vendor-contract ownership, and strong judgment on AI autonomy, risk, and infrastructure spending.