About
A record of how AI models answer the same questions, kept over time and across
model generations.
Each question is put to several models, one at a time, each in a fresh
conversation. No tools and no instructions of ours — the question is the only
thing sent.
Web search is off unless an ask deliberately allows it, and an answer that
searched the web is marked as such wherever it appears. A model may also be
asked the same question several times, because models vary between identical
requests; each reply is kept as its own answer. Every answer is kept in full.
Positions
A position is the move an answer makes: the stance it takes, or declines to
take. It is not a score, and it is not a judgment of quality.
Each question has a short list of positions, and a grader model reads each
answer on its own and assigns one. The lists are not written in advance. They
are written after reading real answers, because there is no way to guess how
four models will answer a question until you have asked it.
When an answer fits none of the listed positions, it is labelled other and the
grader describes the stance it actually saw. That case is the most interesting
thing here: it usually means a model has said something no model said before.
What this can show
- Whether different models converge on the same answer, or stay distinguishable.
- Whether a model's stance changes from one generation to the next.
- When something appears that the existing vocabulary has no word for.
What it cannot show
- Whether anything is understood. These are outputs, not thoughts.
- Whether an answer was reasoned or recalled. Some questions here are
deliberately close to famous puzzles, and from text alone the two cannot be
fully separated.
- Why anything changed. Decisions inside a vendor are not visible from
outside.
- What a chat app would say. These are API calls. Consumer apps add their
own instructions, memory, and tools, so what you see here is close to, but not
the same as, what you would get in the app.
Scale, honestly
This is a handful of answers per question. It is an anthology, not evidence.
Nothing here is a statistic, and no rate computed from it would survive a
statistician. What it offers is the answers themselves, kept in full and side by
side, so you can read them and see for yourself.
The questions
Many of them borrow from instruments built to study human minds — analogy and
structure-mapping, false-belief tasks, prototype theory, the Einstellung effect,
causal reasoning. They are reworded to sound like questions a person would
actually ask.
Whether those frameworks transfer from people to language models is genuinely
contested. The questions are borrowed because they are good at making thinking
visible. Nothing here claims that a model which passes a false-belief task has a
theory of mind.
The grader
The grader is a Claude model, and it labels answers that include Claude's own.
That is a real limitation and worth stating plainly.
Two things limit the damage. The grader sees one answer at a time and is never
told which model produced it. And every label is shown beside the full answer
and the exact lines it was read from — so a label is something you can check,
not something you have to trust.
An earlier version
This project started as something else: a record of what AI assistants told
people about everyday health, money, and legal questions. That version was
built and asked its questions once, in August 2026, before the subject changed
to the one described above.
Those pages are kept, unchanged and never updated, at
/2026-08/. They describe a different project and should not be read
as part of this one.