Inference Technical Lead, Sora
Leads GPU inference engineering for Sora, optimizing model serving efficiency, kernel-level performance, and scalability. Collaborates with research and product teams to build reliable infrastructure for multimodal AI models.
About the job
Responsibilities
- Perform engineering efforts focused on improving model serving, inference performance, and system efficiency
- Drive optimizations from a kernel and data movement perspective to improve system throughput and reliability
- Partner closely with research and product teams to ensure our models perform effectively at scale
- Design, build, and improve critical serving infrastructure to support Sora’s growth and reliability needs
Requirements
- Deep expertise in model performance optimization, particularly at the inference layer
- Strong background in kernel-level systems, data movement, and low-level performance tuning
- Excited about scaling high-performing AI systems that serve real-world, multimodal workloads
- Can navigate ambiguity, set technical direction, and drive complex initiatives to completion
Skills
Gpu Inference, Model Serving, Kernel Optimization, Performance Tuning, Data Movement, Inference Optimization, Ai Systems, Multimodal Workloads, CUDA, PyTorch
Similar jobs
ML Engineering jobsLeads and builds a team of Applied Scientists developing production algorithmic systems for healthcare optimization, LLM applications, and member engagement. Requires 6+ years of relevant industry experience, strong technical judgment, and hands-on expertise across machine learning and optimization.
Build production machine learning systems for model customization, post-training, evaluation, and AWS-native API integration. The role requires 7+ years of relevant engineering experience and expertise in deep learning, transformers, LLM fine-tuning, and production ML infrastructure.
Leads a hands-on AI engineering team developing, evaluating, and deploying large-scale multimodal and video models. The role combines post-training, inference optimization, product experimentation, technical roadmap ownership, and people management.
Leads Discord’s Safety ML team, setting technical direction and overseeing production machine learning systems for content understanding, account integrity, and platform abuse. Requires substantial machine learning and engineering management experience, hands-on technical depth, and experience delivering ML systems at scale.
Senior AI Engineer responsible for production LLM agents that enrich business identity data through web discovery, verification, classification, and risk scoring. The role requires strong asynchronous Python, agent and evaluation expertise, browser automation, and experience operating AI systems in production.