Software Engineer, Inference - Performance Optimization
Models inference performance across application, model, and fleet layers using microbenchmarks to build cost-to-serve estimates. Analyzes workloads end-to-end, enhances bottleneck detection tools, and collaborates on optimizations for latency, throughput, and cost.
About the job
Responsibilities
- Build and refine performance models that translate microbenchmark results into cost-to-serve estimates.
- Analyze inference workloads end to end across applications, models, and fleet infrastructure.
- Enhance tooling to identify bottlenecks across layers for latency and throughput.
- Partner with other teams to turn performance insights into concrete improvements and project how future changes affect inference.
Requirements
- Enjoy reasoning from first principles about distributed systems, model inference, and hardware efficiency.
- Comfortable working across abstraction layers, from application behavior to kernels, accelerators, networking, and fleet scheduling.
- Deep expertise with performance profiling, benchmarking, analysis, and optimization.
- Enjoy collaborating with engineering and research teams to improve real production systems.
Skills
Performance Modeling, Profiling, Benchmarking, Distributed Systems, Model Inference, Hardware Optimization, Kernels, Accelerators, Networking, Fleet Scheduling
Similar jobs
ML Engineering jobsBuild and optimize OpenAI’s inference stack for AWS Trainium across high-performance kernels, compilers, runtimes, and model execution. The role requires systems programming and accelerator experience, with opportunities to solve end-to-end performance problems for frontier-scale AI models.
Build and deploy LLM-powered tools, agents, and ecosystem infrastructure with life sciences research institutions. The role requires deep scientific or biomedical research experience, production software development expertise, and the ability to translate partner workflows into scalable AI systems.
Build research infrastructure and tooling that enables AI models to design silicon, including reinforcement learning environments, EDA integrations, evaluations, and experiment workflows. The role requires strong software engineering fundamentals and comfort working across research, tooling, and chip-design systems.
Build and optimize the production LLM inference runtime for frontier models on OpenAI’s custom silicon. The role spans scheduling, distributed execution, memory and KV-cache management, performance tooling, and hardware-software co-design.
Build and operate machine learning models for sales roleplay, scoring, and coaching products, owning the lifecycle from fine-tuning and evaluation through production and on-device deployment. The role emphasizes open-source models, latency and privacy optimization, and rigorous model testing.