Skip to content
OpenAIOpenAI

Model Policy Manager

Defines and maintains policies for AI model behavior in high-risk domains like agentic systems and user safety. Collaborates with research, engineering, and product teams to operationalize policies into measurable safeguards using empirical data and red-teaming.

About the job

Responsibilities

  • Design and maintain model policies across safety-relevant domains, including dual-use, agentic, and emerging frontier-risk areas.
  • Translate risk and harm models into clear behavioral specifications, evaluation criteria, grading guidance, and system-level safeguards.
  • Define practical boundaries between beneficial uses of AI and assistance that could materially enable harm, exploitation, misuse, or unsafe outcomes.
  • Build policy artifacts that support model training, evaluation, and deployment.
  • Partner with safety researchers, engineers, product teams, and other stakeholders to operationalize policy into scalable model behavior and measurable safeguards.
  • Use red-teaming results, deployment data, model failures, over-refusals, under-refusals, and ambiguous edge cases to improve policy and evaluation quality over time.
  • Identify emerging capability areas where frontier AI systems could create new safety challenges or lower barriers to harm.
  • Study real-world deployments to identify where model behavior succeeds, fails, or drifts from the intended safety posture.
  • Combine longer-horizon safety research with hands-on launch and deployment work.
  • Contribute to system cards, safety reports, policy documentation, launch reviews, and external communications on OpenAI's approach to model safety and risk mitigation.
  • Design and run human data campaigns, including gold set construction, labeling guidance, calibration, adjudication, and eval coverage analysis, to ensure policies can be reliably measured and improved.

Requirements

  • Strong judgment about how advanced AI systems may affect real-world risk, especially in ambiguous, fast-moving, or high-impact areas.
  • Experience building or applying policies, taxonomies, harm models, threat models, or risk frameworks for complex technical, social, or adversarial systems.
  • Ability to move across domains without needing to be the deepest subject-matter expert in every area, while knowing when to seek expert input.
  • Can turn fuzzy questions into structured policy frameworks, evaluation criteria, operational guidance, and enforceable model behavior.
  • Comfortable using empirical evidence, including evaluations, red-teaming results, deployment observations, and model failure modes, to inform policy decisions.
  • Think in systems across policy, data, graders, classifiers, training, deployment safeguards, measurement, monitoring, and escalation workflows.
  • Technical judgment about what model behavior can realistically be trained, measured, evaluated, and enforced at scale.
  • Work well across research, engineering, product, policy, domain experts, and operational teams.
  • Write clearly about complex tradeoffs where safety, user value, and implementation constraints all matter.
  • Pragmatic approach to safety, focused on reducing real-world risk while preserving legitimate, beneficial, and socially valuable uses of AI.
  • Enjoy fast-paced, collaborative research environments where priorities shift as models, evidence, and risks change.
  • Stay grounded in implementation details, empirical results, and what can actually be trained or measured.

Skills

Ai Safety, Red-Teaming, Risk Frameworks, Threat Models, Policy Frameworks, Evaluation Criteria, Harm Models, System Safeguards, Model Training, Model Deployment

OpenAI

OpenAI

San Francisco, CA
Red Team Specialist - Cyber
$198k+/yrHybridSecurity Engineering

The Red Team Specialist evaluates AI models for cyber capabilities, safeguard failures, and agentic-system abuse risks. The role combines hands-on security testing, automated evaluation infrastructure, risk assessment, and cross-functional communication.

Vercel

Vercel

San Francisco, CA
Software Engineer, Trust & Safety
$196k+/yrHybrid5+ YOESecurity Engineering

Build and operate trust and safety systems that detect and mitigate abuse at internet scale. The role combines security engineering, large-scale data analysis, and applied LLM techniques, requiring 5+ years of relevant experience and strong Python and JavaScript/TypeScript skills.

Tailscale

Tailscale

United States
Security Infrastructure Engineer
CA$218k+/yrRemoteSecurity Engineering

This role builds and improves infrastructure security controls across cloud, operating system, Kubernetes, network, and CI/CD environments. It requires cloud security expertise, programming and Infrastructure as Code proficiency, threat-modeling experience, and the ability to lead infrastructure containment during security incidents.

Fluidstack

Fluidstack

New York, NY
Security Engineer, Threat Intelligence
$220k+/yrOn-siteSecurity Engineering

The Security Engineer will track advanced adversaries targeting frontier AI infrastructure, build intelligence pipelines, conduct threat hunts, and create production detections. The role requires hands-on malware and infrastructure analysis, production programming, and close collaboration with detection and incident response teams.

1Password

1Password

United States
Manager, Security Incident Response
$192k+/yrRemote5+ YOESecurity Engineering

Leads and develops a security incident response team while driving automation, AI-assisted workflows, operational maturity, and response strategy. The role requires 5+ years of incident response experience, people leadership, technical depth, and calm management of high-severity incidents.