# Researcher, Agent Safety, Oversight and System Mitigations

**Company:** [OpenAI](https://hotfix.jobs/companies/openai)
**Location:** San Francisco, CA
**Role:** AI Research
**Salary:** $380k – $500k/yr
**Skills:** Ai Safety, Ai Alignment, Security, Threat Modeling, Sandboxing, Process Isolation, Permission Boundaries, Red Teaming, Evaluation Design, Experimental Infrastructure, Agentic Systems, Data Exfiltration, Codex
**Posted:** 2026-09-03

> Researcher or engineer focused on designing, evaluating, and productionizing oversight systems and safety mitigations for autonomous AI agents. The role requires strong systems or security reasoning, threat-modeling ability, and experience building practical evaluations and controls.

## Job Description

## Responsibilities
- Design, build, and evaluate system-level controls for agent actions, including agent-based review.
- Plan how controls integrate with sandboxing, process isolation, and permission boundaries.
- Collaborate with Codex harness engineering to productionize AI controls.
- Red-team end-to-end agentic systems to assess prevention of data exfiltration, unsafe tool use, and other harmful outcomes.
- Measure and improve the safety–productivity tradeoff by reducing missed harmful actions, unnecessary blocks, approval burden, and latency.

## Requirements
- Strong systems or security instincts, with the ability to reason about isolation boundaries, permissions, attack surfaces, and failure modes.
- Ability to turn ambiguous safety questions into threat models, reproducible experiments, and practical mitigations.
- Experience building robust experimental infrastructure and designing evaluations that distinguish effective mitigations from brittle ones.
- Deep interest in frontier AI alignment, safety, and control.

## Nice-to-haves
- Background in AI control or security.

## Compensation
- Annual compensation range: $380,000–$500,000.

## Similar jobs

- [Researcher, Agent Safety, Training and Evaluations](https://hotfix.jobs/jobs/d80336da-e453-4999-9f26-85a125b679d9) - OpenAI - San Francisco, CA - $380k – $500k/yr
- [Research Engineer, Takeoff Intel](https://hotfix.jobs/jobs/398824a2-65cc-4e28-aaeb-26c3b6610876) - Anthropic - San Francisco, CA - $350k – $850k/yr
- [Research, Coding Agents](https://hotfix.jobs/jobs/53bac391-9bcb-48e0-84f9-2dbe7cd426b8) - Thinking Machines Lab - San Francisco, CA - $350k – $475k/yr
- [Research, Safety](https://hotfix.jobs/jobs/420d2360-0401-41e6-8665-9a5112542a0b) - Thinking Machines Lab - San Francisco, CA - $350k – $475k/yr
- [Research Scientist, Life Sciences](https://hotfix.jobs/jobs/c9d5ef0c-39c0-4e12-b2b6-d7f30615fb7e) - Anthropic - San Francisco, CA - $300k – $320k/yr

**Apply:** https://hotfix.jobs/jobs/4544e3bb-bb96-43d2-a96c-cd364b641660
**Canonical:** https://hotfix.jobs/jobs/4544e3bb-bb96-43d2-a96c-cd364b641660