Member of Technical Staff - Post-Training and RL
Develops advanced post-training and reinforcement learning techniques like RLHF/DPO and reward modeling to enhance AI model reasoning, truthfulness, and real-world capabilities at xAI. Seeks passionate AI enthusiasts obsessed with truth-seeking models; prior experience preferred but not required.
About the job
Responsibilities
- Work on critical post-training and reinforcement learning challenges, including reward modeling, preference optimization (RLHF/DPO), and RL for improving reasoning, truthfulness, and real-world capabilities.
Basic Qualifications
- Believe truth-seeking AI is the most important and challenging problem.
- Obsessed about building incredibly useful models through post-training and RL techniques.
- Power user of AI models and eager to push boundaries with reinforcement learning and alignment methods.
- Previous work on post-training, RLHF, or models used by millions is a big plus (relevant experience not required).
- Take pride in work and thrive in meritocratic environments.
Compensation and Benefits
- $180,000 - $600,000 USD
- Equity, comprehensive medical, vision, and dental coverage
- Access to 401(k) retirement plan
- Short & long-term disability insurance
- Life insurance
- Various other discounts and perks
Skills
Reinforcement Learning, RLHF, Dpo, Reward Modeling, Ai Alignment, Post-Training, PyTorch, JAX, Transformers, Machine Learning
Similar jobs
ML Engineering jobsBuild and deploy agentic systems that power AI-driven creative video workflows. The role requires 5+ years of experience, production ML or agentic pipeline development, context engineering, and expertise in evaluation and agent infrastructure.
Build and advance agentic machine-learning systems for multimodal creative tasks, with a focus on video understanding, reasoning, control, and tool use. The role requires strong production ML or agent-pipeline experience and deep knowledge of modern LLM techniques.
Build evaluation methods, RL environments, agent tooling, and scalable infrastructure that make subjective qualities such as design and taste measurable for frontier AI models. The role requires experience with evaluations, RL environments, ML or post-training, plus strong backend engineering skills.
Build and scale generative video and multimodal models, optimizing training and inference for efficiency, throughput, and ultra-low latency. The role requires deep learning systems expertise, strong PyTorch/CUDA experience, and the ability to move research models into production.
Build production-grade AI agents, evaluation infrastructure, and developer tooling that make AI-assisted engineering faster, safer, and reusable across teams. The role requires software engineering experience, platform or internal developer-product experience, and hands-on expertise with LLM integration and orchestration.