Skip to content

AI Inference Engineer

Develops APIs and optimizes large-scale ML model inference using Python, Rust, C++, PyTorch, and CUDA. Benchmarks performance, improves reliability, and implements LLM optimizations on GPU architectures.

About the job

Responsibilities

  • Develop APIs for AI inference that will be used by both internal and external customers
  • Benchmark and address bottlenecks throughout our inference stack
  • Improve the reliability and observability of our systems and respond to system outages
  • Explore novel research and implement LLM inference optimizations

Qualifications

  • Experience with ML systems and deep learning frameworks (e.g. PyTorch, TensorFlow, ONNX)
  • Familiarity with common LLM architectures and inference optimization techniques (e.g. continuous batching, quantization, etc.)
  • Understanding of GPU architectures or experience with GPU kernel programming using CUDA

Skills

Python, Rust, C++, PyTorch, Triton, CUDA, Kubernetes, TensorFlow, Onnx

Cinder

Cinder

New York, NY

AI/ML Engineer
$220k+/yrHybrid5+ YOEML Engineering

Build and operate production machine-learning systems for content safety, from messy customer data through classification, evaluation, and inference. The role requires 5+ years of ML engineering experience, strong Python and MLOps skills, and sound judgment across classical models and LLMs.

Perplexity

Perplexity

San Francisco, CA

Member of Technical Staff
$220k+/yrOn-siteML Engineering

Build AI agent harnesses, models, and product capabilities that enable agents to perform complex work across digital environments. The role combines applied AI research and software engineering, requiring Python proficiency, strong product judgment, and experience with agent tooling, reinforcement learning, or browser technologies.

Mercor

Mercor

San Francisco, CA

Member of Technical Staff, Enterprise Evals Platform
$220k+/yrOn-siteML Engineering

Builds the platform, verifiers, environments, and grading infrastructure used to evaluate enterprise AI agents at scale. The role combines strong software engineering with expertise in agent runtimes, evaluation design, benchmarks, and production failure analysis.

Applied Intuition

Applied Intuition

Sunnyvale, CA

Machine Learning Performance Engineer - Offboard Training & Inference
$215k+/yrOn-siteML Engineering

Optimizes distributed machine learning training and high-throughput offline inference across large accelerator clusters. The role focuses on profiling, scaling efficiency, cluster goodput, GPU performance, and cost-effective processing of autonomy data.

Firecrawl

Firecrawl

San Francisco, CA

Machine Learning Engineer
$210k+/yrHybrid3+ YOEML Engineering

Build and operate production machine-learning systems for search ranking, relevance, extraction quality, and LLM-driven features. The role requires production ML ownership, ranking or relevance expertise, large-scale data experience, Python, and rigorous experimentation skills.