Machine Learning Engineer - Inference
Optimizes and builds production inference systems for large language models at scale using PyTorch and high-performance tooling. Requires 3+ years experience in production code, OS concepts, and AI inference systems.
About the job
Responsibilities
- Design and build the production systems that power the Together AI inference engine, enabling reliability and performance at scale.
- Develop and optimize runtime inference services for large-scale AI applications.
- Collaborate with researchers, engineers, product managers, and designers to bring new features and research capabilities to the world.
- Conduct design and code reviews to ensure high standards of quality.
- Create services, tools, and developer documentation to support the inference engine.
- Implement robust and fault-tolerant systems for data ingestion and processing.
Requirements
- 3+ years of experience writing high-performance, well-tested, production-quality code.
- Proficiency with Python and PyTorch.
- Demonstrated experience in building high performance libraries and tooling.
- Excellent understanding of low-level operating systems concepts including multi-threading, memory management, networking, storage, performance, and scale.
- Preferred: Knowledge of existing AI inference systems such as TGI, vLLM, TensorRT-LLM, Optimum.
- Preferred: Knowledge of AI inference techniques such as speculative decoding.
- Preferred: Knowledge of CUDA/Triton programming.
- Nice to have: Knowledge of Rust, Cython and compilers.
Compensation
US base salary range: $160,000 - $230,000 + equity + benefits. Salary determined by location, level, experience, skills, and job-related knowledge.
Skills
PyTorch, Python, CUDA, Triton, vLLM, Tensorrt-Llm, Tgi, Optimum, Rust, Cython
Similar jobs
ML Engineering jobsResearch Scientist II building and improving fraud risk models and scam detection systems using audio, behavioral, and metadata signals. Requires an advanced degree and 3+ years of applied ML experience with Python and modern ML frameworks.
Build reproducible systems for AI model benchmarking, including datasets, evaluation pipelines, containerized environments, scoreboards, and analysis tools. The role requires 4+ years of professional engineering experience, strong Python, dataset rigor, and Docker expertise.
Develop and deploy ML-first behavior prediction and planning systems for autonomous vehicles, forecasting the motion and interactions of road users. Requires a bachelor's degree, deep learning lifecycle expertise, and at least three years of production software experience with C++ or Python.
Build and operate production AI agents, automation workflows, and integrations that improve complex business processes. The role requires 5+ years of software engineering experience, modern LLM and agent-framework expertise, systems integration skills, and strong cross-functional collaboration.
Build and teach reliable AI agent systems through customer workshops, technical content, guidance, and reference implementations. The role requires strong Python and agent-development experience plus a background delivering customer-facing technical training.