Software Engineer, Evaluation Platform / Infra
Build an end-to-end evaluation platform for frontier AI models, spanning backend systems, data pipelines, APIs, and user-facing applications. The role requires software engineering experience, model evaluation expertise, and the ability to collaborate closely with researchers.
About the job
Responsibilities
- Design, build, and maintain a platform for authoring, running, tracking, and analyzing model evaluations used in daily research and model releases.
- Work across evaluation libraries, distributed backend systems, data pipelines, APIs, and user-facing applications.
- Build flexible abstractions for evaluation tasks, environments, graders, datasets, and model outputs.
- Ensure evaluation results are reproducible and trustworthy through versioning, provenance, observability, failure recovery, and quality controls.
- Partner with researchers to identify bottlenecks and turn bespoke workflows into self-serve systems.
- Collaborate with pre-training, post-training, applied, and research tooling teams.
Requirements
- Bachelor's degree or equivalent practical experience in computer science, engineering, machine learning, or a related field.
- 2+ years of post-graduate software engineering or machine learning engineering experience, excluding internships.
- Hands-on experience building or maintaining evaluations, benchmarks, graders, or model-quality systems for large language or multimodal models.
- Strong software engineering fundamentals and experience building reliable, maintainable systems.
- Proficiency in at least one backend programming language; Python and Rust are primarily used.
- Experience with databases, data pipelines, distributed systems, or other data-intensive infrastructure.
- Ability to work across the stack and own projects from discovery through deployment and operation.
- Experience collaborating with cross-functional partners and subject-matter experts.
Nice-to-haves
- Experience building frameworks, SDKs, or developer tools with thoughtful abstractions and strong user experience.
- Experience with distributed job execution, workflow orchestration, sandboxed environments, or large-scale data processing.
- Experience building interfaces for inspecting complex data, comparing experiments, or debugging model behavior.
- Familiarity with large language or multimodal model evaluation, including model-based grading, human evaluation, or synthetic data.
- Experience working closely with researchers to turn rapidly evolving needs into durable systems.
- Experience at a startup or on a small team building technically complex products end to end.
Compensation and Benefits
- Expected annual salary: $300,000–$475,000 USD, depending on background, skills, and experience.
- Health, dental, and vision benefits.
- Unlimited paid time off.
- Paid parental leave.
- Relocation support as needed.
- Visa sponsorship.
Skills
Python, Rust, React, TypeScript, Distributed Systems, Data Pipelines, APIs, Databases, Workflow Orchestration, LLMs, Multimodal Models, Observability
Similar jobs
Fullstack Engineering jobsBuild full-stack research infrastructure used to manage training runs, evaluations, experiments, and model visualizations. The role partners closely with researchers and requires strong software engineering fundamentals, cross-stack ownership, and at least two years of post-graduate engineering experience.
Build and maintain AI-enabled software systems for a healthcare technology company, collaborating across disciplines and protecting sensitive user data. The role requires at least two years of software engineering experience, particularly with distributed event-driven architectures.
Builds and deploys software and machine learning components for behavior prediction, decision-making, and environmental interaction systems. Requires a bachelor’s degree and at least one year of relevant experience with Python, Go or C++, scalable data pipelines, deep learning, and cloud infrastructure.
Software engineering intern designing high-performance systems for real-time trading, distributed data processing, and machine-learning research infrastructure. The role requires strong computer science fundamentals, significant coding experience, and familiarity with Linux, Git, databases, and cloud infrastructure.
Build and deploy full-stack features for an AI code-review product, tackling LLM memory, codebase indexing, and semantic search challenges. The role requires a computer science or equivalent degree, at least one year of software or DevOps engineering experience, and JavaScript/TypeScript expertise.