Member of Technical Staff - Imagine Model
Develops multimodal AI models focused on image, video, and audio for high-fidelity generation, understanding, and agentic systems. Drives data curation, training, evaluation, and production integration of cutting-edge models.
About the job
Responsibilities
- Create and drive engineering agendas to advance multimodal capabilities, with emphasis on image and video generation, editing, understanding, controllable/long-horizon synthesis, agentic planning, RL training, and world simulation (including audio integration for richer video experiences).
- Improve data quality through annotation, filtering, augmentation, synthetic generation, captioning, and in-depth data studies, particularly for visual and audio data.
- Design evaluation frameworks, metrics, benchmarks, evals, and reward models tailored to image/video/audio quality and coherence.
- Implement efficient algorithms for state-of-the-art model performance, including real-time inference, distillation, and scalable serving for visual content.
- Develop scalable data collection and processing pipelines for multimodal (primarily image/video-focused) datasets.
- Collaborate cross-functionally to integrate AI solutions into production and rapidly iterate based on user feedback.
Basic Qualifications
- Track record in leading studies that significantly improve neural network capabilities and performance through better data or modeling.
- Experience in data-driven experiment designs, systematic analysis, and iterative model debugging.
- Experience developing or working with large-scale distributed machine learning systems.
- Ability to deliver optimal end-to-end user experiences.
- Hands-on contributor with initiative, excellence, strong work ethic, prioritization skills, and excellent communication.
Preferred Skills and Experience
- Experience in SFT, RL, evals, human/synthetic data collection, or agentic systems.
- Proficiency in Python, JAX/XLA, PyTorch, Rust/C++, Spark, Ray, and related large-scale frameworks.
- Domain expertise in multimodal applications such as graphics engines, rendering techniques, image/video understanding and generation, world models, real-time simulation, or controllable/long-horizon visual content creation (audio/speech processing or music/audio generation experience is a plus where it supports video).
- Experience with agentic RL training, controllable/long-horizon generation, or multimodal agents that reason and act across modalities (especially in visual domains).
Compensation and Benefits
$180,000 - $440,000 USD
Base salary is just one part of our total rewards package at xAI, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks.
Skills
PyTorch, Transformers, Computer Vision, Multimodal Ai, Reinforcement Learning, Image Generation, Video Generation, Data Pipelines, Inference Optimization, RLHF
Similar jobs
ML Engineering jobsBuild and operate the engineering systems that support post-training research, including reinforcement learning infrastructure, sandboxed execution, data pipelines, and agent scaffolding. The role requires strong Python and systems engineering skills, project ownership, and a relevant bachelor’s degree or equivalent experience.
Build research infrastructure and tooling that enables AI models to design silicon, including reinforcement learning environments, EDA integrations, evaluations, and experiment workflows. The role requires strong software engineering fundamentals and comfort working across research, tooling, and chip-design systems.
Build production AI capabilities for automated slide and document generation, working across LLM applications, data analysis, and content generation. The role requires 3+ years in machine learning and NLP, advanced Python, and experience with LLM frameworks and production systems.
Build and operate large-scale ranking and retrieval systems that power search relevance, including hybrid lexical/vector search, embeddings, query understanding, and permission-aware retrieval. Requires a bachelor's degree and 5+ years of ML engineering experience in ranking or information retrieval.
Develop and deploy machine learning models for biomedical research and AI products, collaborating with scientific, engineering, and product teams. Requires an advanced quantitative degree, substantial ML experience, Python proficiency, and experience bringing models into production or research applications.