Software Engineer, Inference - Multi Modal
Build and optimize high-performance inference infrastructure for OpenAI's multimodal models handling image, audio, and other non-text inputs at scale. Collaborate with research and product teams on low-latency production systems using GPU workloads and inference tooling.
About the job
Responsibilities
- Design and implement inference infrastructure for large-scale multimodal models.
- Optimize systems for high-throughput, low-latency delivery of image and audio inputs and outputs.
- Enable experimental research workflows to transition into reliable production services.
- Collaborate closely with researchers, infra teams, and product engineers to deploy state-of-the-art capabilities.
- Contribute to system-level improvements including GPU utilization, tensor parallelism, and hardware abstraction layers.
Requirements
- Experience building and scaling inference systems for LLMs or multimodal models.
- Worked with GPU-based ML workloads and understand the performance dynamics of large models, especially with complex data like images or audio.
- Enjoy experimental, fast-evolving work and collaborating closely with research.
- Comfortable dealing with systems that span networking, distributed compute, and high-throughput data handling.
- Familiarity with inference tooling like vLLM, TensorRT-LLM, or custom model parallel systems.
- Own problems end-to-end and excited to operate in ambiguous, fast-moving spaces.
Nice to Have
- Experience working with image generation or audio synthesis models in production.
- Exposure to distributed ML training or system-efficient model design.
Skills
Inference Systems, LLMs, Multimodal Models, Gpu Workloads, vLLM, Tensorrt-Llm, Tensor Parallelism, Distributed Compute, Networking, High-Throughput Data Handling
Similar jobs
ML Engineering jobsBuild and optimize OpenAI’s inference stack for AWS Trainium across high-performance kernels, compilers, runtimes, and model execution. The role requires systems programming and accelerator experience, with opportunities to solve end-to-end performance problems for frontier-scale AI models.
Build and deploy LLM-powered tools, agents, and ecosystem infrastructure with life sciences research institutions. The role requires deep scientific or biomedical research experience, production software development expertise, and the ability to translate partner workflows into scalable AI systems.
Build research infrastructure and tooling that enables AI models to design silicon, including reinforcement learning environments, EDA integrations, evaluations, and experiment workflows. The role requires strong software engineering fundamentals and comfort working across research, tooling, and chip-design systems.
Build and optimize the production LLM inference runtime for frontier models on OpenAI’s custom silicon. The role spans scheduling, distributed execution, memory and KV-cache management, performance tooling, and hardware-software co-design.
Build and operate machine learning models for sales roleplay, scoring, and coaching products, owning the lifecycle from fine-tuning and evaluation through production and on-device deployment. The role emphasizes open-source models, latency and privacy optimization, and rigorous model testing.