Jacob Steinhardt - Oversight Foundation Models
September 1, 2026
Abstract
As AI models become more capable, their behaviors become too complex for us to reliably tell whether they are doing what we intended. One promising response is to enlist AI itself: training "oversight assistants" whose job is to understand other AI models and report back to us in terms we can act on. We'd ideally like to ask these assistants questions like:
• What are important situations where the model sandbags?
• Does the model have an objective it wouldn't admit to if asked directly?
• Does the model treat a user differently once it infers something about their identity, and along what axis?
• Is the model's chain of thought load-bearing, or is it a post-hoc rationalization of an answer that was already settled on?
• Is it reward hacking on this input, or actually trying to solve the task?
The assistant should provide answers to these questions backed by reliable empirical evidence.
To get such an assistant, we lay out a vision for building a foundation model for oversight: an AI system mid-trained (or pre-trained) on a large, diverse corpus of experiments on a given "subject model", RLVR'd on a large number of verified oversight tasks, and then fine-tuned to answer natural-language questions about the subject model.
In this talk we'll describe the general vision of oversight foundation models, then lay out our initial results, including training oversight assistants of up to 1.1T parameters and computing scaling laws for hard tasks such as behavior elicitation.