Member of Technical Staff
Engineer optimizing RL inference stack for workloads from ablations to production training. Requires experience with large-scale distributed systems, LLM inference, and proficiency in Python/C++/Rust with PyTorch/JAX/CUDA.
About the job
Responsibilities
- Design and optimize our inference stack for all shapes of RL workloads at xAI, from small scale ablations to production training runs.
- Analyze, profile and address performance bottlenecks in large scale RL systems.
- Work closely with the modelling team to efficiently implement novel RL techniques and algorithms.
Basic Qualifications
- Experience in building, debugging, and optimizing efficiency of large-scale distributed systems.
- Experience in LLM inference.
- Proficiency in programming languages such as Python, C++ and/or Rust; frameworks such as PyTorch, Jax, CUDA.
- Willingness to dive deep and solve hardcore problems at all levels of the stack.
Preferred Skills and Experience
- Strong knowledge in quantization and numerics in LLM inference and training.
- Experience in developing inference engines, e.g. SGLang, vLLM.
Skills
Python, C++, Rust, PyTorch, JAX, CUDA, Llm Inference, Quantization, Distributed Systems, Sglang, vLLM
Similar jobs
ML Engineering jobsBuild and operate the engineering systems that support post-training research, including reinforcement learning infrastructure, sandboxed execution, data pipelines, and agent scaffolding. The role requires strong Python and systems engineering skills, project ownership, and a relevant bachelor’s degree or equivalent experience.
Build research infrastructure and tooling that enables AI models to design silicon, including reinforcement learning environments, EDA integrations, evaluations, and experiment workflows. The role requires strong software engineering fundamentals and comfort working across research, tooling, and chip-design systems.
Build production AI capabilities for automated slide and document generation, working across LLM applications, data analysis, and content generation. The role requires 3+ years in machine learning and NLP, advanced Python, and experience with LLM frameworks and production systems.
Build and operate large-scale ranking and retrieval systems that power search relevance, including hybrid lexical/vector search, embeddings, query understanding, and permission-aware retrieval. Requires a bachelor's degree and 5+ years of ML engineering experience in ranking or information retrieval.
Develop and deploy machine learning models for biomedical research and AI products, collaborating with scientific, engineering, and product teams. Requires an advanced quantitative degree, substantial ML experience, Python proficiency, and experience bringing models into production or research applications.