What You’ll Do
- Review flagged content, safety escalations, and account-level abuse signals daily. Triage cases, apply policy judgment, and take action (content flags, account review, bans/recovery) on an ongoing basis.
- Use patterns from casework to design and refine safety policy across the product stack, working with engineering, legal, safety research, and security stakeholders.
- Build and maintain tooling and automation that make casework faster and more consistent: triage agents, ban/recovery workflows, abuse detection frameworks and templates.
- Partner with product teams to embed safety into the product experience: model refusals, content flagging, account review, and safety protections, informed by queue observations.
- Improve observability and detection for safety-relevant events (model safety trends, abuse patterns, malicious behavior in production).
Skills and Qualifications
Minimum qualifications:
- 2+ years in an operational trust & safety, content moderation, or fraud/abuse ops role with direct, recurring responsibility for a case queue.
- Experience owning policy definition, operationalization, and enforcement end to end, evidenced by specific policies or enforcement programs built or run.
- Direct case experience with at least one of: cybersecurity abuse, CBRN-relevant risk, youth safety, or prompt injection, in a production environment.
- Working familiarity with model safety and abuse risk categories (jailbreaks, prompt injection, scaled abuse) and how to identify and mitigate them in a live product, evidenced by specific cases handled.
- Practical experience using AI tools (Claude, Codex, or similar) to build or accelerate operational workflows.
Preferred qualifications:
- Experience with safety and integrity operations specifically on AI-powered products or LLM APIs, and their unique abuse patterns.
- Track record of turning recurring case patterns into reusable tooling, workflows, or process improvements, while still owning the underlying queue.
- Experience training, onboarding, or setting the quality bar for other moderators or reviewers.
You’ll Thrive in This Role if
- You have thoughtful opinions about safe, trustworthy, frictionless user experiences and test those opinions against real cases in the queue.
- You can translate safety and technical constraints into clear product trade-offs and feature requirements, growing out of daily casework.
- You bias toward speed and learning, measured by how much faster or better the queue runs.
Logistics
Compensation: Depending on background, skills and experience, the expected annual salary range for this position is $190,000 - $300,000.
Benefits: Generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.