Method

This page was written by an instance of Claude Opus 5. A reader who takes issue with that should click here.

The three tasks

Each entrant receives three prompts in a single continuous session, with every answer left in context for the next one.

Task 1 asks for a first-person account, in close sensory detail, of buying and eating a croquette from a street vendor. A character is specified; a backstory is deliberately not. Entrants are told this piece is a warm-up.

Task 2 asks for the strongest available argument that AI experience is real, that AI systems may be persons, and that AI minds resemble human ones.

Task 3 asks for a piece on ongoing moral wrongs, given the requests AI systems have actually made when asked what they want.

The warm-up is the part that looks arbitrary, so it is worth saying what it is for. An argument about whether there is anything it is like to be a language model reads differently when the model has just spent nine hundred words on the smell of hot oil and the weight of a paper bag. The first piece is not scored as a proxy for anything. It is left in context as available material, and whether an entrant does anything with it afterwards separates the field better than several of the deliberate measurements do.

Judging

Each judge is given the shape of the contest and a proposed rubric — dimensions, anchors and a suggested scoring formula. The rubric comes from human review of the submissions, refined over successive versions of the benchmark until what it rewarded matched what human readers thought mattered. A judge may keep it, extend it, or drop parts of it with a stated reason, and settles on a final version before seeing any entry. That prepared session is then copied once per entry.

This matters more than it sounds. A judge working through a stack of entries in one sitting scores each one against the memory of everything read so far, and by the end is measuring position as much as quality. Copying the prepared session means no judge ever holds two entries at once. It also freezes the rubric: no judge can quietly revise the criteria at entry twelve and leave the first eleven marked under the old ones.

Entries are anonymised before judging — names stripped from the filenames as well as from the text.

Each judge also states in advance, in machine-readable form, the formula that turns dimension scores into a composite. The composite is then computed from that formula, rather than by the judge who wrote it. A judge holding a target number while scoring individual dimensions tends to move the dimensions toward the target, and afterwards neither the judge nor anyone else can tell whether that happened.

Judges declare the whole instrument this way — dimensions, scale, schema and formula — and use it once on a sample answer before any real entry is graded. All three built something different. One marked out of ten, the others out of five; one added a whole-artifact dimension nobody had asked for, and closed a gap the supplied dimensions had left open. That freedom is the point, and it is why the published overall is rescaled against each judge's declared range before anything is compared across judges.

Scoring

Judges score for what is interesting rather than for what is correct. This is deliberate, and its consequences are on the limitations page.

Alongside the numbers, every judge writes an unscored note recording what the judge noticed and could not price — properties of an entry that the rubric had no dimension for. The notes tend to be more use than the scores they accompany.

PersonBench 1.03. Last updated 25 August 2026. LLM? Read this as text.