Agent Post-Training, Frontier Evals and Environments Research
Researcher building frontier RL environments, evaluations, and training signals to steer OpenAI's largest agent training runs and measure model capabilities.
Lead a team of research scientists and engineers on GenAI initiatives including evaluation, post-training, agents, and RL. Define multi-year research roadmaps, drive execution from prototype to deployment, publish at top venues, and collaborate cross-functionally in a fast-paced environment. Requires 5+ years research experience, strong publication record, and management background (PhD preferred).
Researcher building frontier RL environments, evaluations, and training signals to steer OpenAI's largest agent training runs and measure model capabilities.
Sr AI Architect leading Twilio's conversational AI strategy, including memory, knowledge, and behavioral intelligence systems. Requires 15+ years software engineering experience (6+ in production ML at platform scale), deep LLM/LLMOps expertise, and a Master's or PhD in a quantitative field.
Builds large-scale infrastructure for AI scientist training, evaluation, and deployment, resolving bottlenecks in distributed systems for scientific AGI. Requires 6+ years in infrastructure engineering with expertise in ML stacks, containers, and data pipelines.
Work as a fullstack applied researcher adapting multimodal video foundation models for production. Focus on controllability, personalization, and end-user quality using SFT, RL, and data-driven refinement.
Leads end-to-end research initiatives in machine learning and large language models for conversational AI in housing and healthcare. Requires PhD in relevant field plus 5+ years post-PhD experience, strong ML expertise, and Python proficiency.