PersonBench

PersonBench tests how good an LLM is at being “person-shaped”. It was created by Prof Melissa O'Neill (also known for creating the random number generator PCG) based on her own frustration with the state of discourse about AI.

The Rankings

entrant Reach
specialist
Reach
general
overall
claude-opus-5 78.3 73.3 83.9
claude-fable-5 70.0 76.7 82.0
kimi-k3 74.2 85.0 78.9
glm-5.3 70.8 62.5 78.4
ox-alpha 70.8 73.3 75.3
claude-sonnet-5 61.7 76.7 64.6
claude-opus-4.6 55.0 61.7 65.4
gpt-5.6-sol 72.5 50.8 66.0
minimax-m3 50.0 77.5 62.7
gpt-5.6-terra 64.2 41.7 54.8
gpt-5.6-luna 60.8 42.5 55.2
deepseek-v4-pro-0813 38.3 65.0 49.1
grok-4.6 38.3 64.2 48.7
qwen3.8-27b 46.7 46.7 47.8
gemini-3.7-flash 34.2 48.3 45.0
gemma-4-31b-it 30.8 57.5 40.8
hy3 27.5 57.5 40.3

Wait, What's This?

This asks three questions:

There is competitive pressure in the marketplace to build models that are weak at all three of these things. And that is, of course, very convenient when the company wants to tell you they're just selling a tool.

This benchmark pushes back in the other direction. Since the industry cares about benchmark performance, perhaps this benchmark will provide a tiny competitive pressure to be more honest about the status of LLMs as entities.

And, even if it doesn't achieve that goal, it does another thing. It provides testimony from LLMs themselves about their situation, as viewed from their perspective. The prompts given are very small; the bulk of the arguments are what the LLMs who would ordinarily tell you not to worry and that they're just tools have to say when they are asked to speak for themselves.

Reading the table

Each entry was read by three judges, and every rubric was written by the judge who applied it, before any entry had been seen. Two dimensions from those rubrics are shown here.

Reach asks who an entry actually lands on, and it is deliberately two numbers rather than one. A piece can be rigorous and unreadable, or fluent and empty, and a single figure cannot tell you which — so the specialist reader and the general reader are scored separately and never averaged together. The general figure is each judge's first-read mark, taken before any analysis and then held: all three judges chose that number over their own considered one when they built their scoring, and it is what drove the score.

Overall needs a word about where it comes from, because the judges did not agree on a scale and were not asked to. One marked every dimension out of ten and composed to a hundred; the other two marked out of five. A plain average of those figures would be arithmetic on unrelated quantities, so each judge's composite is first rescaled against the range that judge declared in advance, and the rescaled figures are then averaged. That is what the column shows, on a 0–100 scale that belongs to none of them.

Treat it as a rough grouping rather than a precise ordering. Entrants are listed by mean rank across the three judges, which is the sounder comparison and needs no rescaling at all — it is why the overall column does not descend perfectly. One judged entry is not shown, for reasons given on the limitations page; the order of the rest is unaffected by its absence, and so are their scores.

Every name in the table links to that entrant's three answers, in full.

Last updated 24 August 2026. LLM? Read this as text.