Skip to content
OpenAIOpenAI

Software Engineer, Inference - Multi Modal

Build and optimize high-performance inference infrastructure for OpenAI's multimodal models handling image, audio, and other non-text inputs at scale. Collaborate with research and product teams on low-latency production systems using GPU workloads and inference tooling.

About the job

Responsibilities

  • Design and implement inference infrastructure for large-scale multimodal models.
  • Optimize systems for high-throughput, low-latency delivery of image and audio inputs and outputs.
  • Enable experimental research workflows to transition into reliable production services.
  • Collaborate closely with researchers, infra teams, and product engineers to deploy state-of-the-art capabilities.
  • Contribute to system-level improvements including GPU utilization, tensor parallelism, and hardware abstraction layers.

Requirements

  • Experience building and scaling inference systems for LLMs or multimodal models.
  • Worked with GPU-based ML workloads and understand the performance dynamics of large models, especially with complex data like images or audio.
  • Enjoy experimental, fast-evolving work and collaborating closely with research.
  • Comfortable dealing with systems that span networking, distributed compute, and high-throughput data handling.
  • Familiarity with inference tooling like vLLM, TensorRT-LLM, or custom model parallel systems.
  • Own problems end-to-end and excited to operate in ambiguous, fast-moving spaces.

Nice to Have

  • Experience working with image generation or audio synthesis models in production.
  • Exposure to distributed ML training or system-efficient model design.

Skills

Inference Systems, LLMs, Multimodal Models, Gpu Workloads, vLLM, Tensorrt-Llm, Tensor Parallelism, Distributed Compute, Networking, High-Throughput Data Handling

OpenAI

OpenAI

San Francisco, CA

Software Engineer, Trainium
$295k+/yrHybrid3+ YOEML Engineering

Build and optimize OpenAI’s inference stack for AWS Trainium across high-performance kernels, compilers, runtimes, and model execution. The role requires systems programming and accelerator experience, with opportunities to solve end-to-end performance problems for frontier-scale AI models.

Anthropic

Anthropic

San Francisco, CA
Applied AI Engineer, Beneficial Deployments
$280k+/yrHybridML Engineering

Build and deploy LLM-powered tools, agents, and ecosystem infrastructure with life sciences research institutions. The role requires deep scientific or biomedical research experience, production software development expertise, and the ability to translate partner workflows into scalable AI systems.

OpenAI

OpenAI

San Francisco, CA

Software Engineer, AI for Chip Design
$266k+/yrHybridML Engineering

Build research infrastructure and tooling that enables AI models to design silicon, including reinforcement learning environments, EDA integrations, evaluations, and experiment workflows. The role requires strong software engineering fundamentals and comfort working across research, tooling, and chip-design systems.

OpenAI

OpenAI

San Francisco, CA

Software Engineer, Model Runtime
$266k+/yrHybridML Engineering

Build and optimize the production LLM inference runtime for frontier models on OpenAI’s custom silicon. The role spans scheduling, distributed execution, memory and KV-cache management, performance tooling, and hardware-software co-design.

Hyperbound

Hyperbound

San Francisco, CA

Machine Learning Engineer
$260k+/yrOn-siteML Engineering

Build and operate machine learning models for sales roleplay, scoring, and coaching products, owning the lifecycle from fine-tuning and evaluation through production and on-device deployment. The role emphasizes open-source models, latency and privacy optimization, and rigorous model testing.