# Safeguards Enforcement Analyst, User Well-being

**Company:** [Anthropic](https://hotfix.jobs/companies/anthropic)
**Location:** San Francisco, CA, New York, NY, Washington, DC
**Role:** Data Analytics
**Salary:** $245k – $285k/yr
**Skills:** SQL, Data Analysis, Content Moderation, Trust And Safety, Generative AI, Prompt Engineering, Llm Classification, Experiment Design, Evaluation Datasets, Quality Assurance, Workflow Management, Policy Enforcement, Detection Models, Precision And Recall, Claude Code
**Posted:** 2026-08-03

> Analyzes and improves mental-health safeguards for generative AI by evaluating interventions, tuning detection systems, reviewing flagged content, and identifying policy gaps. Requires trust-and-safety or related well-being experience, experimentation and measurement skills, SQL or comparable data analysis, and sound judgment in high-consequence cases.

## Job Description

## Responsibilities

- Support the design and execution of interventions, define key metrics, and curate evaluation datasets.
- Partner with Engineering and Data Science teams to build, tune, and validate detection models for automated intervention systems, including threshold-setting and precision/recall tradeoffs.
- Monitor intervention and detection-system performance over time.
- Review flagged content to drive enforcement and policy improvements.
- Support in-product features that connect users to crisis resources, working with Product, Legal, and external partners on referral pathways and user-facing content.
- Provide detailed feedback to the Safeguards Policy Design team on policy gaps based on real scenarios.
- Track emerging AI policy and external research on AI’s relationship to mental health and use findings to inform decisions and workflows.

## Requirements

- Experience in trust and safety, product policy, content moderation, or a related field, with direct exposure to mental health, suicide and self-harm, or related well-being harms.
- Experience designing or running experiments, evaluations, or measurement studies to determine whether an intervention worked.
- Experience translating policy definitions into measurable rubrics, review guidelines, or classification criteria for human reviewers or automated systems.
- Experience managing or coordinating content-review operations, including quality assurance and workflow management.
- Proficiency in SQL and/or other data-analysis tools.
- Experience with generative AI products, including writing effective prompts for content review, classification, or evaluation.
- Experience turning open questions and data into concise, insightful analysis.
- Experience identifying emerging risks and communicating findings to cross-functional stakeholders.
- Understanding of challenges involved in implementing product policies at scale in content moderation.
- Sound judgment in ambiguous, high-consequence cases and comfort making decisions and escalating appropriately when signals are incomplete.

## Nice-to-haves

- Subject-matter expertise in mental health through academia, clinical practice, crisis intervention, trust and safety, or related settings.
- Experience building or evaluating LLM-based classification systems.
- Experience using agentic tools, such as Claude Code, to scale analysis or automate recurring work.
- Experience working within crisis support.

## Compensation and benefits

- Annual salary: **$245,000–$285,000 USD**.
- Competitive compensation and benefits.
- Optional equity donation matching.
- Generous vacation and parental leave.
- Flexible working hours.
- Office space for collaboration.
- Minimum education: bachelor’s degree or an equivalent combination of education, training, and/or experience.
- Required field of study: a field relevant to the role, as demonstrated through coursework, training, or professional experience.
- Location-based hybrid policy: staff are currently expected to be in an office at least 25% of the time; some roles may require more.
- Visa sponsorship may be available depending on the role and candidate.

## Similar jobs

- [Data Operations Manager, Human Data](https://hotfix.jobs/jobs/a7b9d0c5-fa0e-47b4-b15e-0aa9d254ca98) - Anthropic - San Francisco, CA - $270k – $365k/yr
- [Marketing Analytics Manager](https://hotfix.jobs/jobs/03cc4786-dadb-4872-be5e-e98c6f1c7485) - Baseten - San Francisco, CA - $200k – $240k/yr
- [Growth & Monetization Analyst](https://hotfix.jobs/jobs/69aca69e-af2f-4cba-a4e3-38be3618eafe) - Eve - San Francisco, CA - $160k – $200k/yr
- [Marketing Data & Systems Analyst](https://hotfix.jobs/jobs/a2062d00-2ee9-4889-9ea3-223a3e5dfb2a) - Idme - Mountain View, CA - $155k – $190k/yr
- [Product Analytics](https://hotfix.jobs/jobs/9b514970-625f-42c5-9c18-30066067534d) - OnePay - Remote - $150k – $180k/yr

**Apply:** https://hotfix.jobs/jobs/333dfa26-71b8-441a-a0ca-6dcb754981a9
**Canonical:** https://hotfix.jobs/jobs/333dfa26-71b8-441a-a0ca-6dcb754981a9