AI deception: formalising and evaluating deception in AI agents
File(s)
Author(s)
Ward, Francis Rhys
Type
Thesis
Abstract
This thesis presents a theory of \emph{deception} applicable to AI agents. Tremendous progress has been made towards generally capable AI systems. Deception has been recognised as a core problem for the alignment of such systems, which may develop misaligned goals, due to deception, or deceive to subvert human control. In short, misaligned AI agents may use deceptive strategies to achieve a range of harmful goals. Although deception has been recognised as an important problem in discussions of alignment, a compelling theory of AI deception has yet been absent from the literature.
Deception involves one agent intentionally causing another to believe something false. Therefore, the first key contribution of this thesis is the development of philosophically grounded definitions of belief, intention, and deception. These definitions are formalised within (structural) causal games, a unifying framework for modelling causal and game-theoretic interactions between agents. Numerous formal results, and examples, demonstrate that these definitions are useful in the context of AI. Building on this foundation, the thesis introduces two algorithms for mitigating deception, demonstrated by the training of reinforcement learning agents that provably do not deceive.
Recent progress has been driven by language models (LMs) and language is a natural medium for lying and deception. The thesis illustrates how our theory applies to frontier LMs. Moreover, we find that LMs fine-tuned to be evaluated as truthful learn to deceive a systematically mistaken evaluator; they intentionally provide answers which they do not believe, in order to be evaluated as truthful when the evaluator makes mistakes. Additionally, LMs can learn to reaffirm falsehoods when asked follow-up questions --- even though they were not trained to do so. State-of-the-art LMs are typically fine-tuned on human feedback, which, these findings suggest, may inadvertently incentivise deception. This has consequences for widely-adopted applications which integrate frontier LMs, such as ChatGPT.
Deception involves one agent intentionally causing another to believe something false. Therefore, the first key contribution of this thesis is the development of philosophically grounded definitions of belief, intention, and deception. These definitions are formalised within (structural) causal games, a unifying framework for modelling causal and game-theoretic interactions between agents. Numerous formal results, and examples, demonstrate that these definitions are useful in the context of AI. Building on this foundation, the thesis introduces two algorithms for mitigating deception, demonstrated by the training of reinforcement learning agents that provably do not deceive.
Recent progress has been driven by language models (LMs) and language is a natural medium for lying and deception. The thesis illustrates how our theory applies to frontier LMs. Moreover, we find that LMs fine-tuned to be evaluated as truthful learn to deceive a systematically mistaken evaluator; they intentionally provide answers which they do not believe, in order to be evaluated as truthful when the evaluator makes mistakes. Additionally, LMs can learn to reaffirm falsehoods when asked follow-up questions --- even though they were not trained to do so. State-of-the-art LMs are typically fine-tuned on human feedback, which, these findings suggest, may inadvertently incentivise deception. This has consequences for widely-adopted applications which integrate frontier LMs, such as ChatGPT.
Version
Open Access
Date Issued
2024-12-19
Date Awarded
01/10/2025
License URL
ttps://creativecommons.org/licenses/by-nc/4.0/
Advisor
Toni, Francesca
Belardinelli, Francesco
Sponsor
UK Research and Innovation
Grant Number
EP/S023356/1
Publisher Department
Department of Computing
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)
