Senior Research Engineer, Safety
Research and build safety models, evaluations, and runtime safeguards for conversational AI agents, addressing prompt injection, unsafe tool use, privacy, and policy risks. Requires 4+ years in AI/ML engineering, research, or safety plus experience deploying and evaluating language models or agentic systems.
About the job
Responsibilities
- Research and build safeguards against prompt injection, unsafe tool use, sensitive-data disclosure, policy violations, and hallucinated commitments.
- Build adversarial evaluations, simulations, red-team datasets, and regression suites informed by production failures.
- Develop and deploy classifiers, judges, reward signals, post-training methods, and runtime safeguards for safer agent behavior.
- Analyze production traces and incidents to identify root causes, test mitigations, and measure their impact.
- Partner with Security, Product, Infrastructure, Legal, and customer-facing teams to turn enterprise requirements into scalable safeguards and rollout practices.
Requirements
- 4+ years of experience in AI/ML engineering, research, or AI safety.
- Hands-on experience evaluating, post-training, or deploying language models or agentic systems.
- Experience with modern post-training techniques, such as reinforcement learning, preference optimization, distillation, model routing, and synthetic-data generation.
- Experience with adversarial testing, model red teaming, prompt injection, policy enforcement, privacy, or safe tool use.
- Fluency in Python and modern machine learning tooling, with strong experimental judgment and engineering depth to ship production systems.
- Ability to own ambiguous, high-stakes technical problems and make clear risk and product tradeoffs.
Nice to Have
- Experience building safeguards for high-stakes or regulated enterprise workflows.
- Familiarity with human-in-the-loop review, incident response, or responsible rollout frameworks for machine learning systems.
Compensation and Benefits
- Base salary: $200,000–$400,000, plus equity.
- Medical, dental, and vision benefits for employees and families.
- Life insurance and disability benefits.
- Retirement plan.
- Parental leave.
- Fertility and family-building benefits.
- Monthly wellness and lifestyle stipend.
- Daily office lunches and snacks.
- Flexible vacation policy.
Skills
Python, Machine Learning, Language Models, Agentic Systems, Reinforcement Learning, Preference Optimization, Distillation, Model Routing, Synthetic Data, Red Teaming, Prompt Injection, Privacy, Safe Tool Use, Incident Response
Similar jobs
AI Research jobsBuild and expand customer-facing agentic AI products, MCP integrations, and automated reconciliation workflows for private fund management. The role requires senior-level software engineering, strong systems thinking, product judgment, and hands-on experience building and evaluating AI systems.
Build Vanta’s organizational intelligence layer by shipping prototypes, internal tools, and AI agent workflows that make cross-source data useful to EPD, GTM, and other teams. The role requires recent hands-on LLM product work, independent problem scoping, and strong judgment around AI quality, reliability, cost, and latency.
Conduct research and build open foundation models and training systems aimed at accelerating scientific discovery. The role requires a PhD-level background and substantial experience training foundation models, with expertise in agentic training or multimodal data preferred.
Own applied multimodal ML research from clinical problem definition through production, developing and rigorously evaluating computer vision, NLP, and deep learning systems for radiology. Requires strong Python and PyTorch expertise, 4+ years of relevant experience, and an MS, PhD, or equivalent practical experience.
Owns reusable patterns, standards, and tooling for production agentic service workflows, guiding platform priorities, automation measurement, and quality governance. Requires 8+ years in operations, product, or AI, hands-on agentic workflow experience, and strong LLM, metrics, and cross-functional influence skills.