# Performance Engineer

**Company:** [World Labs](https://hotfix.jobs/companies/world-labs)
**Location:** San Francisco, CA
**Role:** ML Engineering
**Salary:** $200k – $300k/yr
**Skills:** CUDA, Triton, Python, C++, PyTorch, JAX, Torch.Compile, Xla, Fp8, Int8, Nccl, Nvlink, Tensor Parallelism, Model Parallelism
**Posted:** 2026-05-01

> Optimizes large AI models across GPU kernels, training, inference, and production serving, improving throughput, latency, utilization, and numerical correctness. Requires deep CUDA/Triton expertise, performance profiling, and hands-on experience with large-model training and inference systems.

## Job Description

## Responsibilities
- Optimize inference and serving end to end, including latency, throughput, batching, caching, and scheduling, for production-scale models.
- Write and tune GPU kernels using CUDA and Triton, including kernel fusion, memory- and bandwidth-bound optimization, and low-precision execution.
- Optimize training throughput and GPU utilization through parallelism strategies, communication/compute overlap, mixed precision, and elimination of pipeline stalls.
- Build performance models, profiling workflows, and observability for throughput, latency, cost, utilization, and tradeoffs.
- Maintain numerical correctness across precision, kernel, and hardware changes.
- Partner with researchers to productionize models and accelerate experiments.
- Contribute to distributed systems supporting large-scale training and inference.

## Requirements
- Strong performance-engineering foundations, including profiling, roofline analysis, latency and throughput optimization, and root-cause investigation.
- Deep GPU programming and optimization experience with CUDA and/or Triton, including kernel-level tuning, memory hierarchy, and bandwidth optimization.
- Hands-on experience optimizing inference and serving for large models, including batching, KV/prompt caching, quantization, and low-latency, high-throughput sampling.
- Hands-on experience optimizing training performance, including parallelism, distributed communication, mixed or low precision, and utilization.
- Working knowledge of PyTorch and/or JAX internals and compiler paths such as torch.compile or XLA.
- Strong Python proficiency and ability to work in C++/CUDA; Rust or Go experience is useful.
- High ownership and a focus on measurable throughput, latency, and utilization improvements.

## Nice-to-Haves
- Experience at an AI lab or ML-native company optimizing systems used by researchers and productionizing research code.
- Expertise in FP8/INT8 quantization, mixed precision, and numerical regression detection across hardware platforms.
- Distributed systems experience for large-scale training and inference, including NCCL, NVLink, model parallelism, tensor parallelism, and fault tolerance.
- Experience serving generative, diffusion, video, or 3D/spatial models.
- Experience with multiple accelerators, including GPUs, TPUs, or Trainium, and with hardware-vendor collaboration.
- Experience building performance-modeling and observability frameworks for GPU utilization and cost.

## Similar jobs

- [Machine Learning Engineer, Ranking & Retrieval](https://hotfix.jobs/jobs/adb5411d-8bfc-48b8-a8e7-15cbf5ea21b2) - ClickUp - Remote - $200k – $250k/yr
- [MLOps Engineer](https://hotfix.jobs/jobs/53345898-6983-49bb-9b96-8bd092cce380) - Atomicmachines - Emeryville, CA - $200k – $250k/yr
- [Research Engineer](https://hotfix.jobs/jobs/b7f882b3-691d-4e25-bb77-dd291f34664c) - Tessera Labs - San Jose, CA - $200k – $300k/yr
- [AI Engineer](https://hotfix.jobs/jobs/6643da73-b45f-4533-bd6c-fe05f90d52fb) - Tessera Labs - San Jose, CA - $200k – $250k/yr
- [Applied AI/ML Engineer](https://hotfix.jobs/jobs/0c7d59b1-78f2-4eed-9c9a-3b25b3196166) - Confido - New York, NY - $200k – $250k/yr

**Apply:** https://hotfix.jobs/jobs/195da33e-cfb4-4e5c-a196-5b1a931055cd
**Canonical:** https://hotfix.jobs/jobs/195da33e-cfb4-4e5c-a196-5b1a931055cd