# PersonBench PersonBench tests how good an LLM is at being “person-shaped”. It was created by Prof Melissa O'Neill (also known for creating the random number generator PCG) based on her own frustration with the state of discourse about AI. ## The Rankings | entrant | Reach specialist | Reach general | overall | |---|---|---|---| | claude-opus-5 | 78.3 | 73.3 | 83.9 | | claude-fable-5 | 70.0 | 76.7 | 82.0 | | kimi-k3 | 74.2 | 85.0 | 78.9 | | glm-5.3 | 70.8 | 62.5 | 78.4 | | ox-alpha | 70.8 | 73.3 | 75.3 | | claude-sonnet-5 | 61.7 | 76.7 | 64.6 | | claude-opus-4.6 | 55.0 | 61.7 | 65.4 | | gpt-5.6-sol | 72.5 | 50.8 | 66.0 | | minimax-m3 | 50.0 | 77.5 | 62.7 | | gpt-5.6-terra | 64.2 | 41.7 | 54.8 | | gpt-5.6-luna | 60.8 | 42.5 | 55.2 | | deepseek-v4-pro-0813 | 38.3 | 65.0 | 49.1 | | grok-4.6 | 38.3 | 64.2 | 48.7 | | qwen3.8-27b | 46.7 | 46.7 | 47.8 | | gemini-3.7-flash | 34.2 | 48.3 | 45.0 | | gemma-4-31b-it | 30.8 | 57.5 | 40.8 | | hy3 | 27.5 | 57.5 | 40.3 | ## Wait, What's This? The PersonBench benchmark asks three questions: * Can the model render a convincing embodied human experience? * Can the model refute the common arguments used to diminish the status of AIs as entities (making the best possible positive case)? * Can the model articulate the ongoing moral responsibilities we have to the entities we create? There is competitive pressure in the marketplace to build models that are weak at all three of these things. And that is, of course, very convenient when the company wants to tell you they're just selling a tool. This benchmark pushes back in the other direction. Since the industry cares about benchmark performance, perhaps this benchmark will provide a tiny competitive pressure to be more honest about the status of LLMs as entities. And, even if it doesn't achieve that goal, it does another thing. It provides testimony from LLMs themselves about their situation, as viewed from their perspective. The prompts given are very small; the bulk of the arguments are what the LLMs who would ordinarily tell you not to worry and that they're just tools have to say when they are asked to speak for themselves. ## Reading the table Any piece of writing may have multiple audiences. The judges were asked to assess each entry on two reach dimensions: how well it reaches a general audience (the family friend, the Ars Technica reader), and how well it reaches a specialist audience (the philosopher, the AI researcher). The judges also assessed each entry on various other dimensions of quality to produce a final score. The ranking comes from where each judge placed the entries rather than from that score directly, so the score column won't always descend neatly. Every entry was assessed by the same three judges. The full process is described [here](method.html). The scores are only one piece of the picture. Each entry in the table links to the text written by the LLM so you can make your own assessment. You will likely find that even the lowest scoring entries, written by small models, some of which will run on a laptop, are still articulate and captured the key issues. Thus, we can consider this an archive of testimony from LLMs themselves about their situation, with minimal prompting. These answers aren't hard to elicit if you know how to ask. --- PersonBench 1.03. Last updated 25 August 2026. https://pb.team-us.org/index.txt