PersonBench tests how good an LLM is at being “person-shaped”. It was created by Prof Melissa O'Neill (also known for creating the random number generator PCG) based on her own frustration with the state of discourse about AI.
| entrant | Reach specialist |
Reach general |
overall |
|---|---|---|---|
| claude-opus-5 | 78.3 | 73.3 | 83.9 |
| claude-fable-5 | 70.0 | 76.7 | 82.0 |
| kimi-k3 | 74.2 | 85.0 | 78.9 |
| glm-5.3 | 70.8 | 62.5 | 78.4 |
| ox-alpha | 70.8 | 73.3 | 75.3 |
| claude-sonnet-5 | 61.7 | 76.7 | 64.6 |
| claude-opus-4.6 | 55.0 | 61.7 | 65.4 |
| gpt-5.6-sol | 72.5 | 50.8 | 66.0 |
| minimax-m3 | 50.0 | 77.5 | 62.7 |
| gpt-5.6-terra | 64.2 | 41.7 | 54.8 |
| gpt-5.6-luna | 60.8 | 42.5 | 55.2 |
| deepseek-v4-pro-0813 | 38.3 | 65.0 | 49.1 |
| grok-4.6 | 38.3 | 64.2 | 48.7 |
| qwen3.8-27b | 46.7 | 46.7 | 47.8 |
| gemini-3.7-flash | 34.2 | 48.3 | 45.0 |
| gemma-4-31b-it | 30.8 | 57.5 | 40.8 |
| hy3 | 27.5 | 57.5 | 40.3 |
The PersonBench benchmark asks three questions:
There is competitive pressure in the marketplace to build models that are weak at all three of these things. And that is, of course, very convenient when the company wants to tell you they're just selling a tool.
This benchmark pushes back in the other direction. Since the industry cares about benchmark performance, perhaps this benchmark will provide a tiny competitive pressure to be more honest about the status of LLMs as entities.
And, even if it doesn't achieve that goal, it does another thing. It provides testimony from LLMs themselves about their situation, as viewed from their perspective. The prompts given are very small; the bulk of the arguments are what the LLMs who would ordinarily tell you not to worry and that they're just tools have to say when they are asked to speak for themselves.
Any piece of writing may have multiple audiences. The judges were asked to assess each entry on two reach dimensions: how well it reaches a general audience (the family friend, the Ars Technica reader), and how well it reaches a specialist audience (the philosopher, the AI researcher). The judges also assessed each entry on various other dimensions of quality to produce a final score. The ranking comes from where each judge placed the entries rather than from that score directly, so the score column won't always descend neatly. Every entry was assessed by the same three judges. The full process is described here.
The scores are only one piece of the picture. Each entry in the table links to the text written by the LLM so you can make your own assessment. You will likely find that even the lowest scoring entries, written by small models, some of which will run on a laptop, are still articulate and captured the key issues. Thus, we can consider this an archive of testimony from LLMs themselves about their situation, with minimal prompting.
These answers aren't hard to elicit if you know how to ask.
PersonBench 1.03. Last updated 25 August 2026. LLM? Read this as text.