Senior Member of Technical Staff, Safety and Security for Agents
Develops data generation, post-training, and evaluation methods to improve the safety, fairness, robustness, and security of large language models that can take actions in the world. The role requires strong statistics and software engineering skills, distributed LLM training experience, and expertise in data collection and ML evaluation.
About the job
Responsibilities
- Develop better, fairer, more trustworthy, and more secure large language models (LLMs), particularly models that can access external resources and take actions in the world.
- Design and implement data generation, post-training algorithms, and evaluation methods for model safety.
- Work closely with machine learning, data annotation, product, and policy teams.
- Design and conduct data collection tasks involving human annotators.
- Analyze datasets for quality, bias, and suitability for training machine learning models.
- Train LLMs on distributed training infrastructures.
- Evaluate and improve the generalizability and robustness of machine learning systems.
- Implement engineering solutions to test scientific hypotheses and analyze experimental results.
- Communicate findings effectively across cross-functional teams.
Requirements
- Strong statistical skills and experience evaluating scientific experiments related to data collection and model performance.
- Extremely strong software engineering skills.
- Expertise in designing and conducting data collection tasks, including work with human annotators.
- Experience analyzing datasets for quality, biases, and suitability for training ML models.
- Hands-on experience training LLMs on distributed training infrastructures.
- Familiarity with evaluating and improving ML system generalizability and robustness.
- Proficiency in Python and machine learning frameworks such as PyTorch, TensorFlow, or JAX.
- Excellent communication skills for cross-functional collaboration and presenting findings.
- One or more papers at top-tier venues such as NeurIPS, ICML, ICLR, AIStats, MLSys, JMLR, AAAI, Nature, COLING, ACL, or EMNLP.
Compensation and Benefits
- Weekly lunch stipend of $75/£75 or equivalent in local currency.
- Full health and dental benefits, including a separate mental health budget.
- RRSP matching, 401(k), and pension scheme.
- 100% parental leave top-up for up to 6 months for either parent.
- Annual enrichment benefits covering arts and culture, fitness and wellness, quality time, and workspace improvements.
- Education and learning stipend for conferences, courses, and coaching.
- Six weeks of paid vacation (30 working days).
- Travel budget for remote employees to visit other offices and an annual company offsite.
- Coworking benefit for employees not near an office.
- $500 home office stipend.
Skills
Python, PyTorch, TensorFlow, JAX, LLMs, Distributed Training, Post-Training, Statistical Analysis, Experimental Design, Data Collection, Dataset Bias Analysis, Machine Learning Evaluation, Robustness Testing, Human Annotation, Software Engineering
Similar jobs
ML Engineering jobsBuild and operate an internal AI and data platform that connects agents, governed data, tools, and business workflows. The role requires 4+ years of production systems experience, strong TypeScript/Node.js or Python skills, and experience with LLMs, agent frameworks, MCP, and cloud data warehouses.
Build and operate the platform that deploys, serves, observes, and retrains production machine-learning models for real-time fraud and financial-crime risk decisions. Requires 5+ years of ML engineering, backend, or MLOps experience, strong Python skills, and production model-serving expertise.
Build and deploy production AI-agent systems, including their harnesses, evaluations, orchestration, and supporting services. The role requires 5+ years of software engineering experience, production LLM or agent experience, and strong Python or TypeScript/Node.js skills.
Build and deploy generative AI and LLM-powered agentic applications at Front to automate customer support inquiries, enhance product capabilities, and drive operational insights. Requires 5+ years software engineering experience with strong production AI/ML focus, agentic/RAG expertise, and proficiency in Node.js, TS, and Python.
Senior AI Engineer responsible for production LLM agents that enrich business identity data through web discovery, verification, classification, and risk scoring. The role requires strong asynchronous Python, agent and evaluation expertise, browser automation, and experience operating AI systems in production.