Skip to content
World LabsWorld Labs

Performance Engineer

Optimizes large AI models across GPU kernels, training, inference, and production serving, improving throughput, latency, utilization, and numerical correctness. Requires deep CUDA/Triton expertise, performance profiling, and hands-on experience with large-model training and inference systems.

About the job

Responsibilities

  • Optimize inference and serving end to end, including latency, throughput, batching, caching, and scheduling, for production-scale models.
  • Write and tune GPU kernels using CUDA and Triton, including kernel fusion, memory- and bandwidth-bound optimization, and low-precision execution.
  • Optimize training throughput and GPU utilization through parallelism strategies, communication/compute overlap, mixed precision, and elimination of pipeline stalls.
  • Build performance models, profiling workflows, and observability for throughput, latency, cost, utilization, and tradeoffs.
  • Maintain numerical correctness across precision, kernel, and hardware changes.
  • Partner with researchers to productionize models and accelerate experiments.
  • Contribute to distributed systems supporting large-scale training and inference.

Requirements

  • Strong performance-engineering foundations, including profiling, roofline analysis, latency and throughput optimization, and root-cause investigation.
  • Deep GPU programming and optimization experience with CUDA and/or Triton, including kernel-level tuning, memory hierarchy, and bandwidth optimization.
  • Hands-on experience optimizing inference and serving for large models, including batching, KV/prompt caching, quantization, and low-latency, high-throughput sampling.
  • Hands-on experience optimizing training performance, including parallelism, distributed communication, mixed or low precision, and utilization.
  • Working knowledge of PyTorch and/or JAX internals and compiler paths such as torch.compile or XLA.
  • Strong Python proficiency and ability to work in C++/CUDA; Rust or Go experience is useful.
  • High ownership and a focus on measurable throughput, latency, and utilization improvements.

Nice-to-Haves

  • Experience at an AI lab or ML-native company optimizing systems used by researchers and productionizing research code.
  • Expertise in FP8/INT8 quantization, mixed precision, and numerical regression detection across hardware platforms.
  • Distributed systems experience for large-scale training and inference, including NCCL, NVLink, model parallelism, tensor parallelism, and fault tolerance.
  • Experience serving generative, diffusion, video, or 3D/spatial models.
  • Experience with multiple accelerators, including GPUs, TPUs, or Trainium, and with hardware-vendor collaboration.
  • Experience building performance-modeling and observability frameworks for GPU utilization and cost.

Skills

CUDA, Triton, Python, C++, PyTorch, JAX, Torch.Compile, Xla, Fp8, Int8, Nccl, Nvlink, Tensor Parallelism, Model Parallelism

ClickUp

ClickUp

United States

Machine Learning Engineer, Ranking & Retrieval
$200k+/yrRemote5+ YOEML Engineering

Build and operate large-scale ranking and retrieval systems that power search relevance, including hybrid lexical/vector search, embeddings, query understanding, and permission-aware retrieval. Requires a bachelor's degree and 5+ years of ML engineering experience in ranking or information retrieval.

Atomicmachines

Atomicmachines

Emeryville, CA

MLOps Engineer
$200k+/yrOn-site5+ YOEML Engineering

Build and operate production ML infrastructure spanning training, deployment, serving, monitoring, data pipelines, and feedback-driven retraining. The role requires strong MLOps and DevOps experience, Python and SQL proficiency, and ownership of reliable cloud-based systems.

Tessera Labs

Tessera Labs

San Jose, CA
Research Engineer
$200k+/yrOn-siteML Engineering

Build and scale post-training, reinforcement-learning, evaluation, and inference systems for long-horizon agents operating over complex enterprise software. The role requires strong Python and PyTorch or JAX skills, distributed GPU experience, empirical rigor, and the ability to take research results into production.

Tessera Labs

Tessera Labs

San Jose, CA
AI Engineer
$200k+/yrHybrid3+ YOEML Engineering

Build and operate production AI agents that transform enterprise processes, data, and code. The role focuses on tool layers, retrieval, context management, evaluations, monitoring, auditability, and guardrails, requiring strong Python and TypeScript plus experience with production LLM systems and traditional machine learning.

Confido

Confido

New York, NY

Applied AI/ML Engineer
$200k+/yrOn-site5+ YOEML Engineering

Build and productionize applied AI/ML systems for document understanding, agentic workflows, and demand forecasting using rich, messy enterprise data. The role requires 3+ years of production AI/ML experience, strong evaluation and monitoring practices, and a STEM master’s degree.