Build production machine learning systems for model customization, post-training, evaluation, and AWS-native API integration. The role requires 7+ years of relevant engineering experience and expertise in deep learning, transformers, LLM fine-tuning, and production ML infrastructure.
295k – 445k/yr
On-site7+ YOEML Engineering
About the role
Responsibilities
Partner with strategic customers and internal teams to define target model behaviors, diagnose failure modes, and translate real-world needs into training, evaluation, and system requirements.
Build and scale production ML systems for model customization, post-training, and fine-tuning-as-a-service workflows.
Investigate whether training and customization workflows produce intended outcomes, and identify changes to data, evaluation, training, or infrastructure that improve performance.
Partner with backend and infrastructure engineers to integrate ML capabilities into AWS-native API environments.
Improve post-training systems, tooling, APIs, and developer workflows based on partner deployment learnings.
Bring model improvements, training workflows, and evaluation best practices into production.
Design systems that enable strategic partners and enterprise customers to safely customize OpenAI models.
Debug and improve systems spanning model behavior, training data, APIs, distributed infrastructure, and customer-facing product surfaces.
Operate with high ownership in an ambiguous 0→1 environment where reliability matters.
Requirements
Master’s or PhD in Computer Science, Machine Learning, or a related field, or equivalent practical experience.
7+ years of professional engineering experience in relevant ML, infrastructure, or product-driven engineering roles.
Strong ML engineering experience building, training, fine-tuning, evaluating, or deploying production AI systems.
Hands-on experience with deep learning, transformer models, and frameworks such as PyTorch or TensorFlow.
Familiarity with training and fine-tuning large language models, including supervised fine-tuning, distillation, preference optimization, reinforcement learning, or other post-training techniques.
Strong software engineering fundamentals, including data structures, algorithms, systems design, and high-quality production code in Python, Rust, or similar languages.
Experience with model customization, evaluation systems, data pipelines, distributed systems, cloud infrastructure, or production ML platform tradeoffs.
Ability to collaborate across model behavior, APIs, infrastructure, Research, Safety, product engineering, and external technical partners.
Comfort moving quickly through ambiguity, owning problems end-to-end, and learning as needed.
Research Engineer/Scientist shaping personalities and behaviors of personalized AI models like ChatGPT using RL, reward modeling, synthetic data, and post-training methods. Requires strong ML engineering and research experience with large models.
295k – 555k/yrHybrid7+ YOEML Engineering
Agent Post-Training, Artifacts Research
OpenAISan Francisco, CA
Train frontier models to generate polished artifacts (docs, spreadsheets, slides) by owning post-training improvements across RL, data, evals, and alignment. Requires strong ML fundamentals and hands-on LLM/RL experience.
295k – 445k/yrOn-site7+ YOEML Engineering
Agent Post-Training, Computer Use Research
OpenAISan Francisco, CA
Train frontier models to operate computers, browsers, and desktops. Design experiments, build evals, own post-training pipelines (RL, data, graders), and ship improvements into OpenAI agents.
295k – 445k/yrOn-site7+ YOEML Engineering
Agent Post-Training, Connectors Research
OpenAISan Francisco, CA
Train frontier agents to interface with professional software via code, APIs, and structured integrations. Design experiments, own post-training improvements (RL, evals, data), and ship capabilities into major model runs.
295k – 445k/yrOn-site7+ YOEML Engineering
Context Researcher
OpenAISan Francisco, CA
Context Researcher on the Agent Post-Training team scaling compute on context for frontier agent models. Designs experiments, owns post-training improvements, builds evals, and ships capabilities into Codex and ChatGPT.