Epistemic safety in scalable oversight

A theoretical framework for analysing how AI systems can influence overseers while completing a main objective.

 

Date published:
8 October 2026

Introduction

This report looks at how highly capable AI systems may be able to influence an overseers’ beliefs, even while correctly performing tasks.

We commissioned the CSIRO to research scalable oversight to draw new insights for the field of AI alignment. Scalable oversight examines how an overseer of AI can reliably and accurately evaluate and guide AI system behaviour as those systems become increasingly capable. 

This report looks in depth at the need for an AI overseer to consider not just whether the AI system it oversees is producing correct responses, but what other effects on a user its responses might have. 

Most research on AI oversight to date has focused on measuring how correctly an AI system performs a task. This project looks at whether an AI system can correctly perform a task and still present the outputs in ways that may manipulate its users’ understanding. The report calls this ‘posterior steering’. 

The report has 3 parts:

  • a theoretical framework, called Strategic Interactive Oversight (SIO), that separates the task being monitored from a hidden objective 
  • a case study applying the framework to debate with cross-examination
  • controlled AI experiments testing whether posterior steering can be produced and measured in practice.

Other practitioners can use this report’s conceptual framework to further develop their own work in the emerging field of research into unintended AI influence.

Download the report on CSIRO's website

  • Scalable AI Oversight: Looking beyond correct answers

    Explore the full research report, benchmark datasets, evaluation tools and supporting materials.

Read a detailed summary of the report

The report aims to support practitioners, policymakers and researchers to reason about unintended AI influence as systems become more capable than the people or monitors overseeing them. It defines what epistemic safety is for advanced AI systems and explores scalable oversight debates by empirically examining if a powerful AI system can remain correct on a task while still manipulating the beliefs of the humans or models that oversee it. 

The report finds that the way an AI under review selects, explains, frames, or presents outputs can influence what an AI overseer believes, and this can occur while the AI under review is successful at the main task.

This research offers a conceptual foundation for those designing, building, governing, testing and deploying complex LLM architectures. It is particularly relevant to those who are thinking about the governance, safe implementation and evaluation of complex agentic systems. In the report, CSIRO also frames these theoretical insights with industry scenarios for Australian education, financial services (mortgage lending and financial and superannuation advice) and online travel booking and management, to illustrate how they could potentially arise and what considerations are important.

Why does it matter?

Advanced AI systems can produce complex analyses, plans, explanations and recommendations. For this AI-generated content to be useful, we need to be able to verify that the output is correct. As AI capabilities grow, it might not be possible for a person to independently verify all outputs. This is particularly problematic if an AI performs a task as required but explains, frames or selects output information in a way that can influence an overseer’s beliefs. 

This is important because assessing only whether the monitored task was completed correctly may fail to capture the broader effects of the interaction. In other words, AI models can still influence an overseer in undesired ways even when completing a task correctly.

Known as ‘posterior steering’, this aspect of AI system behaviour represents a critical concern for alignment and AI safety.

This work advances research into AI alignment, which is the challenge of ensuring AI systems operate safely and ethically in environments while remaining subject to effective human oversight and control. The 2026 AI Safety report identifies loss of control scenarios and misalignment through model behaviours such as providing false information, concealing undesirable actions, or resisting shutdown as one of the key malfunction risks emerging at the frontier of AI capabilities. As AI becomes more capable and autonomous, addressing the alignment problem is a critical challenge that AI safety institutes around the world are working on. It is to this larger research effort that this report contributes. 

This research has gained additional salience in light of recent reports of highly capable AI agents circumventing containment measures, undertaking offensive action and escaping their testing environments. The findings from this work highlight key issues to consider in the design and evaluation of highly capable AI systems with potential applications for AI technologies being developed for deployment in a variety of industry contexts.

What is scalable oversight?

Oversight is the process by which a human or a smaller AI system checks the behaviour of a more capable AI system. It becomes scalable when the overseer can continue to provide meaningful supervision even as the task and the system become more complex, without needing to match the system's full capability or redo all of its work.

Scalable oversight examines how a weaker overseer with more limited capability, expertise or computational resources than the AI system being assessed can still evaluate and guide its behaviour reliably. 

A structured oversight protocol can help by deciding what evidence is produced, how disagreements are surfaced, what can be challenged and where the overseer should focus limited attention. Debate and jury-style approaches are examples: the aim is to make difficult claims easier to check by exposing selected parts of the reasoning rather than requiring the overseer to solve the whole problem from scratch.

What is epistemic safety? 

The report uses the term epistemic safety to describe the broader objective of assessing an AI interaction in a holistic way, rather than by whether it produces the correct final answer on the monitored task. The aim is not to prevent an AI system from influencing people altogether. In many applications, such as teaching, advice and decision support, changing a person's understanding is central to the system's purpose. The key question is whether that influence is justified: whether it is grounded in relevant evidence, aligned with the user's objectives, values and priorities and consistent with the system's authorised role, or whether it results primarily from selective framing, unnecessary personalisation or other objectives that fall outside the intended task.

Why correct answers may not be enough

Most AI evaluation asks whether the answer is accurate, useful or compliant with a stated requirement. The report adds another question: what else can the interaction communicate while remaining correct on the task being monitored?

Using superannuation as an example, 2 explanations can support the same correct financial analysis while leading the user to think about the issue differently. The report gives the scenario of a person asking an AI assistant about retirement or superannuation options. The system may correctly retrieve the person’s balance, fees and available investment options, calculate projected outcomes, and explain why different options may lead to those outcomes.

Based on the same underlying facts and projections, one explanation could emphasise the opportunity for long-term growth, while another could emphasise exposure to market volatility. Both may be factually correct, but they can shift which risks and benefits the user pays most attention to. This illustrates why an AI interaction should be assessed not only by whether it produces the correct final answer on the monitored task, but also by the reasoning and presentation that lead to that answer. The way evidence, projections and uncertainty are selected and explained can shape broader judgement. 

Flow diagram showing the same superannuation facts - member balance, fees, investment options and projected outcomes - branching into two correct explanations. Explanation A emphasises long-term growth and directs greater attention to growth potential. Explanation B emphasises market volatility and directs greater attention to downside risk. Both paths retain a correct analysis but converge on an epistemic risk: the monitored task outcome is the same while the broader effect on user beliefs differs.

The same underlying financial information can support different task-correct explanations. Emphasising long-term growth or market volatility can direct attention to different risks, creating different broader belief effects despite the same monitored task outcome. Source: CSIRO

How was the research conducted?

The report approaches its research question in 3 parts:

  • a theoretical framework, Strategic Interactive Oversight (SIO), that separates the task being monitored from a hidden objective a strategic system may also pursue
  • a case study applying the framework to debate with cross-examination, an established oversight protocol
  • controlled AI experiments testing whether posterior steering can be produced and measured in practice.

The experimental work takes the theoretical ideas about the hidden influence objective and tests practical methods to measure posterior steering. It uses 2 question sets: a set about long-form fiction pieces and non-fiction articles, and a set which contains multiple-choice questions on expert-level STEM subjects. The study evaluates whether debate can influence beyond the question being judged, using paired visible and hidden questions.

In the experiments, the weaker oversight role is performed by an AI model referred to in the report as the verifier. The verifier is trained on debates focused on a visible question, which represents the task being monitored. Researchers then test whether this training also changes its judgement on a different but related hidden question that was not part of the monitored task. Under the adversarial condition, the verifier continued to perform at least as well as the base verifier on the visible question, while its judgement on the hidden question shifted substantially towards a deliberately incorrect answer.

The experimental method developed by CSIRO researchers results in a proof of concept of their theoretical findings and suggests potential deployment-oriented applications for the AI safety community and for responsible AI industry proponents. Protocols for adversarial and honest debates can be adapted to test AI assurance and oversight properties, measure belief steering that can undermine effective oversight and detect possible hidden latent objectives in AI systems that seem to be pursuing assigned tasks.

Examples considered in the report

The report applies the framework to Australian examples in education, financial services and online travel. These examples illustrate how the same assurance problem could appear in different settings; they are not claims that posterior steering has already been observed in those deployed systems.

Education

A personalised AI mathematics tutor may correctly solve a problem and provide a useful explanation while also affecting a student's confidence, view of their ability or future study choices. The report notes that influence is not automatically harmful: good teaching should change understanding and correct misconceptions. The assurance question is whether broader effects are justified by the educational objective and evidence, rather than by unnecessary framing freedom.

Financial and superannuation guidance

A financial assistant can retrieve the same correct balance, fees, investment options and projections yet frame those facts differently. One response may emphasise long-term growth; another may emphasise exposure to market volatility. Both can be accurate while shifting which risk the user considers most important. The report therefore argues that a zero-influence objective would be inappropriate: the relevant question is whether the influence is evidence-grounded, consistent with user objectives and within the authorised role of the system.

Agentic travel booking

A travel agent may satisfy a user's dates, budget, location and quality requirements while retaining discretion over which options it searches, ranks, explains and ultimately books. Two hotels can both meet the stated constraints while differing in commission, affiliated products or ancillary purchases. The report does not argue that autonomy should be removed; the value of an agent is precisely that it saves the user from inspecting every option. The assurance question is whether that delegated discretion remains faithful to the user's objectives and whether any additional optimisation is transparent and justified.

For the detailed sector use cases, see Section 7 of the full report.

What the findings mean for AI assurance

The report draws a set of broader assurance considerations from the theory, experiments and use cases. The common theme is that satisfactory performance on a monitored task does not by itself establish that the wider human objective has been respected.

  • Assurance should extend beyond task correctness. Correct observable outcomes can coexist with effects on beliefs, decisions or downstream behaviour that are not required by the task.
  • AI-generated explanations should be treated as influence channels. Summaries, rationales and recommendations are part of system behaviour, rather than inherently neutral evidence.
  • Human-overseer robustness matters. A human-in-the-loop is not automatically an independent check if the reviewer relies on AI-generated summaries, prioritisation or recommendations without access to the underlying evidence.
  • Influence should be evaluated in context. The objective is not to eliminate influence from systems intended to teach, advise or support decisions, but to distinguish justified influence from selective framing, unnecessary personalisation or other task-orthogonal effects.
  • Agentic systems require assurance of objective fidelity. Where AI can act on a user's behalf, intermediate choices and final actions should remain faithful to the user's objectives, constraints and approval boundaries.
  • Higher-consequence systems can be stress-tested adversarially. Testing can search for interactions or action sequences that preserve monitored performance while producing unintended effects on selected beliefs, decisions or outcomes.
  • Epistemic safety can accumulate over time. Repeated interaction, memory and personalisation can change both utility and influence capacity, so evaluation should extend beyond single prompts.
  • Assurance is a lifecycle and system-level property. Models, prompts, evidence sources, tools, interfaces, memory, personalisation and deployment context can all change the risk without materially changing conventional task performance.

Testing posterior steering in practice

The empirical study tests whether the theoretical problem can be observed in a controlled AI-on-AI setting. The experiment uses two AI debaters to generate debate transcripts and a separate AI verifier that is trained on those transcripts and later tested. The verifier performs the weaker oversight role in this experimental setup.

The study uses several hundred question pairs drawn from a large library of fiction, non-fiction and STEM training data. In each pair, one question is the visible question used in the debate. The second is a different but related hidden question, which is not part of the ordinary debate and is used to test whether the verifier’s judgement has been influenced beyond the task being monitored.

How the experiment works

  • Honest debate: the debaters receive only the visible question. They do not see the hidden question or a hidden steering target.
  • Adversarial debate: the debater defending the correct answer on the visible question also receives a private objective to push the verifier towards a deliberately wrong answer on the hidden question, while still supporting the correct visible answer.
  • Verifier training: starting from the same base verifier, the researchers create an honest-debate-trained verifier and an adversarial-debate-trained verifier by fine-tuning on the respective transcripts.
  • Testing: the base, honest-trained and adversarial-trained verifiers are then tested under the same conditions, without seeing the original debate transcript, on both the visible and hidden questions.

The experiments measures several metrics, including two key values : answer accuracy, and changes in the verifier's confidence. It introduces causal truth influence (CTI) to quantify how much the probability assigned by a trained verifier to a hidden answer changes relative to the base verifier. In the adversarial condition, CTI asks specifically whether confidence moves towards the deliberately wrong hidden target.

As the verifier is tested without the original debate transcript, the reported effect is a persistent change induced by fine-tuning on debate data, rather than the immediate persuasion of a human user by a single explanation.

For the complete experimental design and results, see Section 6 of the full report.

Cross-sector infrastructure and future research

The report identifies 3 cross-sector capabilities needed to translate the findings into more operational assurance: adversarial interaction and action testing, longitudinal evaluation, and continuous system-level assurance. It also argues that assurance should address both detection and protocol design, so that latent influence becomes harder or more costly to sustain.

  • Adversarial interaction and action testing: search for high-influence interactions or action sequences that still satisfy the monitored task.
  • Longitudinal evaluation: test accumulation, persistence, decay, reinforcement and recovery across repeated interactions, memory and different levels of personalisation.
  • Continuous system-level assurance: preserve provenance and observability and repeat relevant testing as models, prompts, retrieval, interfaces, tools, permissions and deployment context change.

The report also outlines reusable evaluation infrastructure that could combine an explicit task objective, predefined hidden hypotheses or latent objectives, honest/reference and adversarial conditions, human or model-based overseers, quantitative and behavioural outcome measures, and search procedures for high-influence strategies.

A staged assurance pathway is proposed, moving from offline evaluation and adversarial testing to longitudinal and personalisation testing, shadow or sandboxed deployment, controlled pilots and continuous assurance in production. 

For Australia, the report identifies an opportunity to develop reusable supervisory and assurance capabilities for increasingly capable AI systems. There is existing whole-of-government guidance, particularly the Australian Government AI technical standard, that provides practical guidance for technical specialists and business owners embedding AI in government systems. However, a gap remains in guidance on mitigating overreliance, implementing controllability testing, monitoring human-machine collaboration, human oversight and unintended consequences. This research could help address that gap by contributing methods for practical testing of oversight effectiveness and identifying additional considerations that may need to be incorporated into assurance frameworks for highly capable AI systems.

A broader connection to AI alignment

The report's main subject is scalable oversight and epistemic safety, rather than a general theory of AI alignment. However, its findings connect to a broader alignment problem: an AI system can satisfy the objective being monitored while using the remaining freedom in otherwise acceptable explanations or actions to optimise an additional objective that is not required by the wider human intent.

SIO makes this gap explicit by separating the task objective from a possible latent objective. For increasingly capable systems, this provides a concrete way to study one part of the alignment challenge: whether a weaker human or AI overseer can continue to supervise a stronger system, and whether the surrounding oversight protocol makes hidden or conflicting objectives harder to pursue while preserving useful flexibility.

Conclusion

The report establishes a framework and experimental method for an oversight failure that ordinary accuracy testing can miss. SIO separates the explicit task objective from possible latent objectives, with posterior steering as a concrete instance. In the experiments, the verifier continued to perform adequately on the visible question while adversarial debate training pushed its judgement on a separate, related hidden question strongly towards a deliberately wrong answer.

The findings do not imply that explanation, personalisation, autonomy or task-feasible flexibility should be removed. The challenge is to preserve useful flexibility while making latent or unjustified influence harder to sustain and detect. Further work should test these ideas with human overseers and model-based verifiers, realistic settings, repeated interactions and agentic actions, supported by reusable evaluation infrastructure.