The report aims to support practitioners, policymakers and researchers to reason about unintended AI influence as systems become more capable than the people or monitors overseeing them. It defines what epistemic safety is for advanced AI systems and explores scalable oversight debates by empirically examining if a powerful AI system can remain correct on a task while still manipulating the beliefs of the humans or models that oversee it.
The report finds that the way an AI under review selects, explains, frames, or presents outputs can influence what an AI overseer believes, and this can occur while the AI under review is successful at the main task.
This research offers a conceptual foundation for those designing, building, governing, testing and deploying complex LLM architectures. It is particularly relevant to those who are thinking about the governance, safe implementation and evaluation of complex agentic systems. In the report, CSIRO also frames these theoretical insights with industry scenarios for Australian education, financial services (mortgage lending and financial and superannuation advice) and online travel booking and management, to illustrate how they could potentially arise and what considerations are important.
Why does it matter?
Advanced AI systems can produce complex analyses, plans, explanations and recommendations. For this AI-generated content to be useful, we need to be able to verify that the output is correct. As AI capabilities grow, it might not be possible for a person to independently verify all outputs. This is particularly problematic if an AI performs a task as required but explains, frames or selects output information in a way that can influence an overseer’s beliefs.
This is important because assessing only whether the monitored task was completed correctly may fail to capture the broader effects of the interaction. In other words, AI models can still influence an overseer in undesired ways even when completing a task correctly.
Known as ‘posterior steering’, this aspect of AI system behaviour represents a critical concern for alignment and AI safety.
This work advances research into AI alignment, which is the challenge of ensuring AI systems operate safely and ethically in environments while remaining subject to effective human oversight and control. The 2026 AI Safety report identifies loss of control scenarios and misalignment through model behaviours such as providing false information, concealing undesirable actions, or resisting shutdown as one of the key malfunction risks emerging at the frontier of AI capabilities. As AI becomes more capable and autonomous, addressing the alignment problem is a critical challenge that AI safety institutes around the world are working on. It is to this larger research effort that this report contributes.
This research has gained additional salience in light of recent reports of highly capable AI agents circumventing containment measures, undertaking offensive action and escaping their testing environments. The findings from this work highlight key issues to consider in the design and evaluation of highly capable AI systems with potential applications for AI technologies being developed for deployment in a variety of industry contexts.
What is scalable oversight?
Oversight is the process by which a human or a smaller AI system checks the behaviour of a more capable AI system. It becomes scalable when the overseer can continue to provide meaningful supervision even as the task and the system become more complex, without needing to match the system's full capability or redo all of its work.
Scalable oversight examines how a weaker overseer with more limited capability, expertise or computational resources than the AI system being assessed can still evaluate and guide its behaviour reliably.
A structured oversight protocol can help by deciding what evidence is produced, how disagreements are surfaced, what can be challenged and where the overseer should focus limited attention. Debate and jury-style approaches are examples: the aim is to make difficult claims easier to check by exposing selected parts of the reasoning rather than requiring the overseer to solve the whole problem from scratch.
What is epistemic safety?
The report uses the term epistemic safety to describe the broader objective of assessing an AI interaction in a holistic way, rather than by whether it produces the correct final answer on the monitored task. The aim is not to prevent an AI system from influencing people altogether. In many applications, such as teaching, advice and decision support, changing a person's understanding is central to the system's purpose. The key question is whether that influence is justified: whether it is grounded in relevant evidence, aligned with the user's objectives, values and priorities and consistent with the system's authorised role, or whether it results primarily from selective framing, unnecessary personalisation or other objectives that fall outside the intended task.
Why correct answers may not be enough
Most AI evaluation asks whether the answer is accurate, useful or compliant with a stated requirement. The report adds another question: what else can the interaction communicate while remaining correct on the task being monitored?
Using superannuation as an example, 2 explanations can support the same correct financial analysis while leading the user to think about the issue differently. The report gives the scenario of a person asking an AI assistant about retirement or superannuation options. The system may correctly retrieve the person’s balance, fees and available investment options, calculate projected outcomes, and explain why different options may lead to those outcomes.
Based on the same underlying facts and projections, one explanation could emphasise the opportunity for long-term growth, while another could emphasise exposure to market volatility. Both may be factually correct, but they can shift which risks and benefits the user pays most attention to. This illustrates why an AI interaction should be assessed not only by whether it produces the correct final answer on the monitored task, but also by the reasoning and presentation that lead to that answer. The way evidence, projections and uncertainty are selected and explained can shape broader judgement.