Skip to content

Machine Learning Engineer - Inference

Optimizes and builds production inference systems for large language models at scale using PyTorch and high-performance tooling. Requires 3+ years experience in production code, OS concepts, and AI inference systems.

About the job

Responsibilities

  • Design and build the production systems that power the Together AI inference engine, enabling reliability and performance at scale.
  • Develop and optimize runtime inference services for large-scale AI applications.
  • Collaborate with researchers, engineers, product managers, and designers to bring new features and research capabilities to the world.
  • Conduct design and code reviews to ensure high standards of quality.
  • Create services, tools, and developer documentation to support the inference engine.
  • Implement robust and fault-tolerant systems for data ingestion and processing.

Requirements

  • 3+ years of experience writing high-performance, well-tested, production-quality code.
  • Proficiency with Python and PyTorch.
  • Demonstrated experience in building high performance libraries and tooling.
  • Excellent understanding of low-level operating systems concepts including multi-threading, memory management, networking, storage, performance, and scale.
  • Preferred: Knowledge of existing AI inference systems such as TGI, vLLM, TensorRT-LLM, Optimum.
  • Preferred: Knowledge of AI inference techniques such as speculative decoding.
  • Preferred: Knowledge of CUDA/Triton programming.
  • Nice to have: Knowledge of Rust, Cython and compilers.

Compensation

US base salary range: $160,000 - $230,000 + equity + benefits. Salary determined by location, level, experience, skills, and job-related knowledge.

Skills

PyTorch, Python, CUDA, Triton, vLLM, Tensorrt-Llm, Tgi, Optimum, Rust, Cython

Pindrop

Pindrop

United States

Research Scientist II
$160k+/yrRemote3+ YOEML Engineering

Research Scientist II building and improving fraud risk models and scam detection systems using audio, behavioral, and metadata signals. Requires an advanced degree and 3+ years of applied ML experience with Python and modern ML frameworks.

Office Hours

Office Hours

San Francisco, CA
Software Engineer - Benchmarking
$160k+/yrRemote4+ YOEML Engineering

Build reproducible systems for AI model benchmarking, including datasets, evaluation pipelines, containerized environments, scoreboards, and analysis tools. The role requires 4+ years of professional engineering experience, strong Python, dataset rigor, and Docker expertise.

Applied Intuition

Applied Intuition

Sunnyvale, CA

Software Engineer - Prediction and Planning ML
$151k+/yrOn-site3+ YOEML Engineering

Develop and deploy ML-first behavior prediction and planning systems for autonomous vehicles, forecasting the motion and interactions of road users. Requires a bachelor's degree, deep learning lifecycle expertise, and at least three years of production software experience with C++ or Python.

OnePay

OnePay

United States

Forward Deployed Engineer (AI and Automation)
$150k+/yrRemote5+ YOEML Engineering

Build and operate production AI agents, automation workflows, and integrations that improve complex business processes. The role requires 5+ years of software engineering experience, modern LLM and agent-framework expertise, systems integration skills, and strong cross-functional collaboration.

LangChain

LangChain

New York, NY
AI Engineer, Enablement
$150k+/yrOn-site3+ YOEML Engineering

Build and teach reliable AI agent systems through customer workshops, technical content, guidance, and reference implementations. The role requires strong Python and agent-development experience plus a background delivering customer-facing technical training.