publications | John Burden

2026

Predictable artificial intelligence

Lexin Zhou, Pablo A.M. Casares, Fernando Martínez-Plumed, and 12 more authors

Artificial Intelligence, 2026

DOI
Pressure Reveals Character: Behavioural Alignment Evaluation at Depth

Nora Petrova, and John Burden

2026

2025

Formalising Human-in-the-Loop: Computational Reductions, Failure Modes, and Legal-Moral Responsibility

Maurice Chiodo, Dennis Müller, Paul Siewert, and 3 more authors

arXiv preprint arXiv:2505.10426, 2025
Framing the Game: How Context Shapes LLM Decision-Making

Isaac Robinson, and John Burden

arXiv preprint arXiv:2503.04840, 2025
arXiv
General Scales Unlock AI Evaluation with Explanatory and Predictive Power

Lexin Zhou, Lorenzo Pacchiardi, Fernando Martínez-Plumed, and 23 more authors

arXiv preprint arXiv:2503.06378, 2025

Abs Bib HTML

Ensuring safe and effective use of AI requires understanding and anticipating its performance on novel tasks, from advanced scientific challenges to transformed workplace activities. So far, benchmarking has guided progress in AI, but it has offered limited explanatory and predictive power for general-purpose AI systems, given the low transferability across diverse tasks. In this paper, we introduce general scales for AI evaluation that can explain what common AI benchmarks really measure, extract ability profiles of AI systems, and predict their performance for new task instances, in- and out-of-distribution. Our fully-automated methodology builds on 18 newly-crafted rubrics that place instance demands on general scales that do not saturate. Illustrated for 15 large language models and 63 tasks, high explanatory power is unleashed from inspecting the demand and ability profiles, bringing insights on the sensitivity and specificity exhibited by different benchmarks, and how knowledge, metacognition and reasoning are affected by model size, chain-of-thought and distillation. Surprisingly, high predictive power at the instance level becomes possible using these demand levels, providing superior estimates over black-box baseline predictors based on embeddings or finetuning, especially in out-of-distribution settings (new tasks and new benchmarks). The scales, rubrics, battery, techniques and results presented here represent a major step for AI evaluation, underpinning the reliable deployment of AI in the years ahead.
@article{zhou2025generalscalesunlockai, title = {{General Scales Unlock AI Evaluation with Explanatory and Predictive Power}}, author = {Zhou, Lexin and Pacchiardi, Lorenzo and Martínez-Plumed, Fernando and Collins, Katherine M. and Moros-Daval, Yael and Zhang, Seraphina and Zhao, Qinlin and Huang, Yitian and Sun, Luning and Prunty, Jonathan E. and Li, Zongqian and Sánchez-García, Pablo and Chen, Kexin Jiang and Casares, Pablo A. M. and Zu, Jiyun and Burden, John and Mehrbakhsh, Behzad and Stillwell, David and Cebrian, Manuel and Wang, Jindong and Henderson, Peter and Wu, Sherry Tongshuang and Kyllonen, Patrick C. and Cheke, Lucy and Xie, Xing and Hernández-Orallo, José}, year = {2025}, eprint = {2503.06378}, archiveprefix = {arXiv}, primaryclass = {cs.AI}, journal = {arXiv preprint arXiv:2503.06378}, url = {https://kinds-of-intelligence-cfi.github.io/ADELE/}, }
A Framework to Categorise Modified General-Purpose AI Models as New Models Based on Behavioural Changes

Lorenzo Pacchiardi, John Burden, Fernando Martinez-Plumed, and 1 more author

In Collection of External Scientific Studies on General-Purpose AI Models under the EU AI Act, 2025

JRC143257

DOI
A Framework for General-Purpose AI Model Categorisation

John Burden, Lorenzo Pacchiardi, Fernando Martínez Plumed, and 1 more author

In Collection of External Scientific Studies on General-Purpose AI Models under the EU AI Act, 2025

JRC143256

DOI
Conversational complexity for assessing risk in large language models

John Burden, Manuel Cebrian, and Jose Hernandez-Orallo

EPJ Data Science, 2025

DOI
I Spy With My Model’s Eye: Visual Search as a Behavioural Test for MLLMs

John Burden, Jonathan Prunty, Ben Slater, and 3 more authors

2025
Paradigms of AI Evaluation: Mapping Goals, Methodologies and Culture

John Burden, Marko Tešić, Lorenzo Pacchiardi, and 1 more author

Jun 2025

arXiv:2502.15620 [cs]

Abs DOI

Research in AI evaluation has grown increasingly complex and multidisciplinary, attracting researchers with diverse backgrounds and objectives. As a result, divergent evaluation paradigms have emerged, often developing in isolation, adopting conflicting terminologies, and overlooking each other’s contributions. This fragmentation has led to insular research trajectories and communication barriers both among different paradigms and with the general public, contributing to unmet expectations for deployed AI systems. To help bridge this insularity, in this paper we survey recent work in the AI evaluation landscape and identify six main paradigms. We characterise major recent contributions within each paradigm across key dimensions related to their goals, methodologies and research cultures. By clarifying the unique combination of questions and approaches associated with each paradigm, we aim to increase awareness of the breadth of current evaluation approaches and foster cross-pollination between different paradigms. We also identify potential gaps in the field to inspire future research directions.

2024

Evaluating AI Evaluation: Perils and Prospects

John Burden

Jun 2024

_eprint: 2407.09221
Conversational Complexity for Assessing Risk in Large Language Models

John Burden, Manuel Cebrian, and Jose Hernandez-Orallo

Sep 2024

arXiv:2409.01247 [cs, math]

Abs

Large Language Models (LLMs) present a dual-use dilemma: they enable beneficial applications while harboring potential for harm, particularly through conversational interactions. Despite various safeguards, advanced LLMs remain vulnerable. A watershed case was Kevin Roose’s notable conversation with Bing, which elicited harmful outputs after extended interaction. This contrasts with simpler early jailbreaks that produced similar content more easily, raising the question: How much conversational effort is needed to elicit harmful information from LLMs? We propose two measures: Conversational Length (CL), which quantifies the conversation length used to obtain a specific response, and Conversational Complexity (CC), defined as the Kolmogorov complexity of the user’s instruction sequence leading to the response. To address the incomputability of Kolmogorov complexity, we approximate CC using a reference LLM to estimate the compressibility of user instructions. Applying this approach to a large red-teaming dataset, we perform a quantitative analysis examining the statistical distribution of harmful and harmless conversational lengths and complexities. Our empirical findings suggest that this distributional analysis and the minimisation of CC serve as valuable tools for understanding AI safety, offering insights into the accessibility of harmful information. This work establishes a foundation for a new perspective on LLM safety, centered around the algorithmic complexity of pathways to harm.
The Animal-AI Environment: A Virtual Laboratory For Comparative Cognition and Artificial Intelligence Research

Konstantinos Voudouris, Ibrahim Alhas, Wout Schellaert, and 11 more authors

Oct 2024

arXiv:2312.11414 [cs]

Abs

The Animal-AI Environment is a unique game-based research platform designed to facilitate collaboration between the artificial intelligence and comparative cognition research communities. In this paper, we present the latest version of the Animal-AI Environment, outlining several major new features that make the game more engaging for humans and more complex for AI systems. New features include interactive buttons, reward dispensers, and player notifications, as well as an overhaul of the environment’s graphics and processing for significant improvements in agent training time and quality of the human player experience. We provide detailed guidance on how to build computational and behavioural experiments with the Animal-AI Environment. We present results from a series of agents, including the state-of-the-art Deep Reinforcement Learning agent, Dreamer-v3, on newly designed tests and the Animal-AI Testbed of 900 tasks inspired by research in the field of comparative cognition. The Animal-AI Environment offers a new approach for modelling cognition in humans and non-human animals, and for building biologically-inspired artificial intelligence.

2023

Animal-AI 3: What’s New & Why You Should Care

Konstantinos Voudouris, Ibrahim Alhas, Wout Schellaert, and 9 more authors

Dec 2023

arXiv:2312.11414 [cs]

Abs

The Animal-AI Environment is a unique game-based research platform designed to serve both the artificial intelligence and cognitive science research communities. In this paper, we present Animal-AI 3, the latest version of the environment, outlining several major new features that make the game more engaging for humans and more complex for AI systems. New features include interactive buttons, reward dispensers, and player notifications, as well as an overhaul of the environment’s graphics and processing for significant increases in agent training time and quality of the human player experience. We provide detailed guidance on how to build computational and behavioural experiments with Animal-AI 3. We present results from a series of agents, including the state-of-the-art Deep Reinforcement Learning agent (dreamer-v3), on newly designed tests and the Animal-AI Testbed of 900 tasks inspired by research in comparative psychology. Animal-AI 3 is designed to facilitate collaboration between the cognitive sciences and artificial intelligence. This paper serves as a stand-alone document that motivates, describes, and demonstrates Animal-AI 3 for the end user.
9. From Turing’s Speculations to an Academic Discipline: A History of AI Existential Safety

John Burden, Sam Clarke, and Jess Whittlestone

In The Era of Global Risk, Aug 2023

Abs DOI

Through a short history of research into artificial intelligence, this chapter emphasises the need to ensure the safety of such technology. From longstanding fears of hubris to contemporary technical challenges, concerns around this modern technology are discussed, and the need to make progress on AI existential safety to avoid dangers to future generations is scrutinised.
Inferring Capabilities from Task Performance with Bayesian Triangulation

John Burden, Konstantinos Voudouris, Ryan Burnell, and 3 more authors

Sep 2023

arXiv:2309.11975 [cs]

Abs

As machine learning models become more general, we need to characterise them in richer, more meaningful ways. We describe a method to infer the cognitive profile of a system from diverse experimental data. To do so, we introduce measurement layouts that model how task-instance features interact with system capabilities to affect performance. These features must be triangulated in complex ways to be able to infer capabilities from non-populational data – a challenge for traditional psychometric and inferential tools. Using the Bayesian probabilistic programming library PyMC, we infer different cognitive profiles for agents in two scenarios: 68 actual contestants in the AnimalAI Olympics and 30 synthetic agents for O-PIAAGETS, an object permanence battery. We showcase the potential for capability-oriented evaluation.
Rethink reporting of evaluation results in AI

Ryan Burnell, Wout Schellaert, John Burden, and 13 more authors

Science, Sep 2023

Abs DOI

Aggregate metrics and lack of access to results limit understanding Artificial intelligence (AI) systems have begun to be deployed in high-stakes contexts, including autonomous driving and medical diagnosis. In contexts such as these, the consequences of system failures can be devastating. It is therefore vital that researchers and policy-makers have a full understanding of the capabilities and weaknesses of AI systems so that they can make informed decisions about where these systems are safe to use and how they might be improved. Unfortunately, current approaches to AI evaluation make it exceedingly difficult to build such an understanding, for two key reasons. First, aggregate metrics make it hard to predict how a system will perform in a particular situation. Second, the instance-by-instance evaluation results that could be used to unpack these aggregate metrics are rarely made available (1). Here, we propose a path forward in which results are presented in more nuanced ways and instance-by-instance evaluation results are made publicly available.
Your Prompt is My Command: On Assessing the Human-Centred Generality of Multimodal Models

Wout Schellaert, Fernando Martínez-Plumed, Karina Vold, and 6 more authors

Journal of Artificial Intelligence Research, Sep 2023
Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, and 447 more authors

Transactions on Machine Learning Research, Sep 2023
An International Consortium for Evaluations of Societal-Scale Risks from Advanced AI

Ross Gruetzemacher, Alan Chan, Kevin Frazier, and 11 more authors

Nov 2023

arXiv:2310.14455 [cs]

Abs

Given rapid progress toward advanced AI and risks from frontier AI systems (advanced AI systems pushing the boundaries of the AI capabilities frontier), the creation and implementation of AI governance and regulatory schemes deserves prioritization and substantial investment. However, the status quo is untenable and, frankly, dangerous. A regulatory gap has permitted AI labs to conduct research, development, and deployment activities with minimal oversight. In response, frontier AI system evaluations have been proposed as a way of assessing risks from the development and deployment of frontier AI systems. Yet, the budding AI risk evaluation ecosystem faces significant coordination challenges, such as a limited diversity of evaluators, suboptimal allocation of effort, and perverse incentives. This paper proposes a solution in the form of an international consortium for AI risk evaluations, comprising both AI developers and third-party AI risk evaluators. Such a consortium could play a critical role in international efforts to mitigate societal-scale risks from advanced AI, including in managing responsible scaling policies and coordinated evaluation-based risk response. In this paper, we discuss the current evaluation ecosystem and its shortcomings, propose an international consortium for advanced AI risk evaluations, discuss issues regarding its implementation, discuss lessons that can be learnt from previous international institutions and existing proposals for international AI governance institutions, and, finally, we recommend concrete steps to advance the establishment of the proposed consortium: (i) solicit feedback from stakeholders, (ii) conduct additional research, (iii) conduct a workshop(s) for stakeholders, (iv) analyze feedback and create final proposal, (v) solicit funding, and (vi) create a consortium.
Harms from Increasingly Agentic Algorithmic Systems

Alan Chan, Rebecca Salganik, Alva Markelius, and 19 more authors

In Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency, Nov 2023

event-place: Chicago, IL, USA

Abs DOI

Research in Fairness, Accountability, Transparency, and Ethics (FATE)1 has established many sources and forms of algorithmic harm, in domains as diverse as health care, finance, policing, and recommendations. Much work remains to be done to mitigate the serious harms of these systems, particularly those disproportionately affecting marginalized communities. Despite these ongoing harms, new systems are being developed and deployed, typically without strong regulatory barriers, threatening the perpetuation of the same harms and the creation of novel ones. In response, the FATE community has emphasized the importance of anticipating harms, rather than just responding to them. Anticipation of harms is especially important given the rapid pace of developments in machine learning (ML). Our work focuses on the anticipation of harms from increasingly agentic systems. Rather than providing a definition of agency as a binary property, we identify 4 key characteristics which, particularly in combination, tend to increase the agency of a given algorithmic system: underspecification, directness of impact, goal-directedness, and long-term planning. We also discuss important harms which arise from increasing agency – notably, these include systemic and/or long-range impacts, often on marginalized or unconsidered stakeholders. We emphasize that recognizing agency of algorithmic systems does not absolve or shift the human responsibility for algorithmic harms. Rather, we use the term agency to highlight the increasingly evident fact that ML systems are not fully under human control. Our work explores increasingly agentic algorithmic systems in three parts. First, we explain the notion of an increase in agency for algorithmic systems in the context of diverse perspectives on agency across disciplines. Second, we argue for the need to anticipate harms from increasingly agentic systems. Third, we discuss important harms from increasingly agentic systems and ways forward for addressing them. We conclude by reflecting on implications of our work for anticipating algorithmic harms from emerging systems.

2022

How Sure to Be Safe? Difficulty, Confidence and Negative Side Effects

John Burden, Jose Hernandez-Orallo, and Sean Heigeartaigh

In NeurIPS ML Safety Workshop, Nov 2022
Not a Number: Identifying Instance Features for Capability-Oriented Evaluation

Ryan Burnell, John Burden, Danaja Rutar, and 3 more authors

In Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence, IJCAI-22, Jul 2022

DOI
Evaluating object permanence in embodied agents using the animal-AI environment

Konstantinos Voudouris, Niall Donnelly, Danaja Rutar, and 4 more authors

EBeM’22: Workshop on AI Evaluation Beyond Metrics, July 25, 2022, Vienna, Austria, Jul 2022

Publisher: CEUR Workshop Proceedings
How general-purpose is a language model? usefulness and safety with human prompters in the wild

Pablo Antonio Moreno Casares, Bao Sheng Loe, John Burden, and 1 more author

In , Jul 2022

Issue: 5

2021

Latent Property State Abstraction For Reinforcement Learning

John Burden, Sajjad Kamali Siahroudi, and Daniel Kudenko

In Proceedings of the AAMAS Workshop on Adaptive Learning Agents (ALA), Jul 2021

2020

Automating abstraction for potential-based reward shaping

John Burden

Dec 2020

Publisher: University of York

Abs

Within the field of Reinforcement Learning (RL) the successful application of abstraction can play a huge role in decreasing the time required for agents to learn competent policies. Many examples of this speed-up have been observed throughout the literature. Reward Shaping is one such technique for utilising abstractions in this way. This thesis focuses on how an agent can learn its own abstractions from its own experiences to be used for Potential Based Reward Shaping. As the thesis progresses, the environments for which the abstraction construction is automated grow in complexity and scope — while also utilising less external knowledge of the domains. This culminates in the approaches }backslashtextit{Uniform Property State Abstraction} (UPSA) and }backslashtextit{Latent Property State Abstraction} (LPSA), which can both augment existing RL algorithms and allow them to construct abstractions from their own experience and then effectively make use of these abstractions to improve convergence time. Empirical results from this thesis demonstrate that this approach can outperform existing deep RL algorithms such as Deep Q-Networks over a range of domains.
Uniform State Abstraction for Reinforcement Learning

John Burden, and Daniel Kudenko

In 24th European Conference on Artificial Intelligence,, Dec 2020

DOI