Researcher, Alignment CoT Monitorability at OpenAI
- Company: OpenAI
- Location: San Francisco
- Employment type: Full-time
- Salary: $295K – $500K • Offers Equity
- Posted: 2026-08-04
- Technologies: OpenAI
About the role
ABOUT THE TEAM The CoT Monitorability team at OpenAI studies whether and when the chain-of-thought of frontier reasoning models is monitorable enough to support scalable oversight. We study how to measure monitorability https://openai.com/index/evaluating-chain-of-thought-monitorability/, which training mechanisms affect monitorability, and speculative methods to improve monitorability. While we mostly focus on CoT monitorability at the moment, we care more generally about any form of monitorability, auditing methods, and improving alignment. We were the first to show https://openai.com/index/chain-of-thought-monitoring/ that chain-of-thought monitoring can be a practical additional safety mechanism, and today our monitoring systems are actively used on OpenAI’s largest RL training runs to detect misbehavior.…
Apply on OpenAI's official careers page: https://jobs.ashbyhq.com/openai/82492010-ea96-449d-9949-b726b1a22616