Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action
ICML 2026 (LM4Plan workshop)
Theory of mind

The AI research group at Prolific — papers, notes, and field logs.

Human evaluation
A demographically-aware, multi-dimensional human-preference leaderboard: 27 models judged by 20,000+ stratified participants and ranked with a hierarchical Bradley–Terry–Davidson model. Compare models head-to-head, by metric, and across 22 demographic groups.
Open leaderboard ↗
Alignment
Behavioural alignment evaluated under realistic pressure — 904 multi-turn scenarios across Honesty, Safety, Non-Manipulation, Robustness, Corrigibility, and Scheming. Ranks how models actually behave when instructions conflict, not what they claim they would do.
Open leaderboard ↗
Health evaluation
How the public uses and judges AI for their health, and how the same models hold up when practising clinicians check the advice: where models' health advice goes wrong, what people ask and reward, and why AI judges need checking against clinicians.
Open HUMAINE Health ↗noteHuman evaluation
Preference prediction hits a wall at 66%, demographics explain about 1% of how people judge, and the crowd rewards substance over flattery 2.5 to 1. Yet the pooled signal still teaches sycophancy. A state-of-the-project report from 100,000+ human comparisons, and the studies we're running next.
noteAgentic research
We ran Karpathy's autoresearch loop on a DPO task, then handed the same model its results in a single Claude Code session. 300 Prolific participants judged the outputs: the autonomous loop plateaued below chance, while five minutes of human steering produced the only decisive wins.
noteHuman evaluation
A demographically-aware human-preference framework: 20,000+ stratified participants, 27 models, and a hierarchical Bradley–Terry–Davidson model that turns 21,352 judgements into an interactive, multi-dimensional leaderboard.
ICML 2026 (LM4Plan workshop)
Theory of mind
Preprint coming
Interpretability
ICLR 2026 Workshop ICBINB
AI safety
ICLR 2026 (poster)
Human evaluation