Model Quality Software Engineer, Claude Code
Software Engineer on Claude Code team builds evaluation systems, tooling, and infrastructure to enhance AI coding capabilities. Collaborates with researchers in fast-paced environment; requires 5+ years experience building complex systems.
About the job
Responsibilities
- Design and build eval systems that measure model capabilities across diverse coding tasks
- Build tooling and infrastructure that enables researchers to run experiments at scale
- Develop pipelines for data collection, processing, and analysis
- Create internal tools that improve researcher productivity and accelerate iteration cycles
- Serve as a bridge between product and research—bring strong product intuition to inform which capabilities matter most
- Work closely with researchers to translate research questions into engineering solutions
- Own systems end-to-end—from design through production reliability
You may be a good fit if you:
- Have built and owned complex systems—pipelines, infrastructure, or software that orchestrates many components and handles significant state and logic
- Thrive in high-intensity environments with fast iteration cycles
- Take full ownership of problems and drive them to completion independently
- Are a power user of agentic coding tools and have strong intuition about model capabilities and limitations
- Are comfortable diving into unfamiliar technical domains and figuring things out quickly
- Care deeply about correctness and reliability in the systems you build
- Are excited to work at the boundary between engineering and AI research
- Have at least 5 years of work experience
Strong candidates may also have experience with:
- Writing or maintaining eval/evaluation frameworks
- Reinforcement learning systems
- Working in high-performance, demanding environments—trading firms, quant funds, competitive research labs, or fast-moving startups where intensity is the norm
- Have research computing or scientific infrastructure background
- Have a strong quantitative foundation (math, physics, or related fields)
- Python and TypeScript
Logistics
Education requirements: We require at least a Bachelor's degree in a related field or equivalent experience.
Location-based hybrid policy: Currently, we expect all staff to be in one of our offices at least 25% of the time. However, some roles may require more time in our offices.
Skills
Python, TypeScript, Evaluation Frameworks, Reinforcement Learning, Data Pipelines, Research Infrastructure, Scientific Computing
Similar jobs
ML Engineering jobsDevelop multimodal perception and authentication systems combining visual, audio, and other sensor signals for real-world AI products. The role requires machine learning expertise, practical research-to-system experience, and proficiency in Python and PyTorch with comfort in C++.
Build and optimize OpenAI’s inference stack for AWS Trainium across high-performance kernels, compilers, runtimes, and model execution. The role requires systems programming and accelerator experience, with opportunities to solve end-to-end performance problems for frontier-scale AI models.
Build and operate the engineering systems that support post-training research, including reinforcement learning infrastructure, sandboxed execution, data pipelines, and agent scaffolding. The role requires strong Python and systems engineering skills, project ownership, and a relevant bachelor’s degree or equivalent experience.
Operates and improves the infrastructure powering large-scale post-training and reinforcement learning runs, partnering with researchers to debug failures, improve reliability, and automate recovery. Requires 4+ years operating distributed production systems and strong Python, Go, or C++ skills.
Research-focused engineer advancing agentic model capabilities across synthetic data, task environments, evaluations, training, and usability improvements. Requires strong Python engineering, deep learning framework experience, scalable distributed training skills, and scientific experimentation ability.