Safeguards Enforcement Analyst, Safety Evaluations
Conduct safety evaluations for AI models, monitor results, drive mitigations, and build scalable processes in collaboration with policy, engineering, and stakeholders. Requires trust & safety experience, process-building skills, and comfort in ambiguity.
About the job
Responsibilities
- Support model launch readiness by running evaluations, monitoring results, and surfacing issues to stakeholders
- Partner with policy and domain experts on evaluation lifecycle: risk identification, scoping, creation, and maintenance
- Manage evaluation outcomes with stakeholders, interpret results, and drive mitigations
- Strategically improve eval quality, processes, and paradigms as models advance
- Build processes for product-specific evaluations as products expand
- Design tooling improvements for evolving eval needs and self-serve capabilities
- Write and maintain documentation for evaluation creation, execution, and interpretation
You may be a good fit if you:
- Have experience in trust & safety, content operations, policy enforcement, or related operational roles at tech companies
- Thrive in ambiguous, fast-moving environments
- Have built processes or programs from scratch (zero-to-one work)
- Possess strong program management skills for complex, multi-stakeholder efforts
- Are eager to adopt technical tools and AI-assisted workflows
- Manage multiple workstreams with strong prioritization and context-switching
- Are a strong generalist comfortable with varied work
- Make judgment calls with incomplete info and escalate appropriately
- Communicate clearly and concisely
Strong candidates may also have:
- Experience with high-stakes timelines like launches or incident response
- Coordinating across engineering, policy, and product teams
- Building/maintaining SOPs, runbooks, and documentation
- Proficiency with data tools (SQL, dashboards, spreadsheets)
- Comfort with sensitive content
Skills
SQL, Dashboards, Spreadsheets, Claude Code, Evaluation Tooling, Ai Safety Evaluations, Policy Enforcement, Program Management, SOPs, Runbooks
Similar jobs
Build and run the global spares sourcing program for critical datacenter equipment (chillers, generators, switchgear) at multi-GW scale. Qualify/dual-source suppliers, set stocking levels tied to failure data and criticality, negotiate VMI/consignment terms, and partner with reliability teams to ensure zero downtime from parts shortages.
Own state and local policy agenda for AI data center infrastructure. Build relationships with officials and regulators while drafting testimony and navigating permitting and utility approvals.
Own public affairs and community relations for Fluidstack's data center and power infrastructure projects. Build local support through town halls and coalitions, manage media and opposition responses, and partner with internal teams to align external messaging with project realities.
Leads child safety investigations, enforcement decisions, mandatory-reporting workflows, and process improvements involving sensitive content and abuse signals. The role requires Trust & Safety investigation experience, sound judgment, strong documentation, and cross-functional collaboration.
Own and maintain Primavera P6 CPM schedules, cost tracking, and productivity reporting for greenfield hyperscale data center construction projects. Drive critical path compression, lead reviews with GCs/executives, and deliver defensible as-built records on $500M+ programs.