Quality Systems Lead
Own Encord’s data-quality standards, automated evaluation systems, and audit operations for human-generated AI data. The role combines hands-on Python and SQL development with statistical quality measurement and leadership of a distributed audit team.
About the job
Responsibilities
- Build automated dataset quality evaluation and root-cause detection using model-assisted and LLM-as-judge screening, agreement analysis, anomaly detection, and drift detection.
- Hire, train, calibrate, and manage an audit team of approximately ten specialists based in India.
- Convert audit output into labeled ground truth for training and validating automated evaluation systems.
- Manage the balance between automated and manual quality coverage over time.
- Own quality standards for each delivered data type, including rubrics, worked edge cases, golden sets, and customer-agreed acceptance criteria.
- Build scoring systems for annotator and reviewer performance to inform routing, staffing, and offboarding decisions.
- Set certification pass thresholds for project queues.
- Report quality KPIs including accuracy against client specifications, inter-annotator agreement, rework rate, cost of rework, and coverage.
- Partner with Project Management on remediation while maintaining independent quality standards.
- Own quality unit economics, including cost per audited unit and coverage per pound spent.
- Partner with Product and Engineering to make quality measurement a native Encord platform capability.
Requirements
- 4+ years owning technical and operational outcomes in a data, AI, or service-delivery environment where quality was measured.
- Hands-on Python and SQL development experience.
- Practical experience applying models to quality or evaluation problems, such as LLM-as-judge, model-assisted QA, automated evaluation, anomaly detection, or classifier-based screening.
- Practical experience with sampling methodology and agreement statistics, including Cohen's kappa, Fleiss' kappa, and F1 against ground truth.
- Experience hiring, training, and managing a team, ideally an audit, review, or QA team and preferably a distributed team.
- Track record of building a quality framework or function with reporting used by leadership and customers.
- Ability to write clear rubrics for distributed teams, maintain calibrated standards, and diagnose recurring quality failures.
Nice-to-haves
- Experience with annotation, evaluation, or model-training workflows.
- Familiarity with the quality evidence accepted by frontier AI labs.
- STEM degree or background in data science or research engineering.
- Multilingual delivery and linguistic quality assessment experience.
Compensation and Benefits
- Competitive salary, commission, and equity.
- 25 days of annual leave plus UK public holidays.
- Annual learning and development budget.
- Travel opportunities for customer visits, events, and conferences across the UK and Europe.
- Company lunches twice a week.
- Monthly socials and twice-yearly team offsites.
Skills
Python, SQL, Llm-As-Judge, Automated Evaluation, Anomaly Detection, Sampling Methodology, Cohen'S Kappa, Fleiss' Kappa, F1 Score, Quality Assurance, Model-Assisted Qa, Drift Detection
Similar jobs
QA Engineering jobsSenior Quality Engineer partnering across product teams to assess risk, build web and API test automation, investigate technical failures, and improve quality practices. Requires strong Playwright and TypeScript experience, exploratory testing skills, CI/CD knowledge, and the ability to debug application code.
Evaluates and defines quality standards for Spotify’s AI voice models across products, languages, and use cases. The role designs qualitative and quantitative evaluation programs, benchmarks competing technologies, and translates findings into product and research recommendations.
Leads quality strategy and hands-on engineering for a shared MLOps platform, building test frameworks, deployment safety tooling, and observability systems. The role requires strong platform engineering expertise, technical leadership, and experience guiding teams across complex ML systems.
The Software Quality Engineer will define QA strategies, build automated test frameworks, and validate AI agents and LLM-powered applications. The role requires 5+ years of quality engineering experience, strong Python or TypeScript skills, GitHub Actions, Playwright, and familiarity with AI/ML testing.