AI Snapshots

About

A record of how AI models answer the same questions, kept over time and across

model generations.

Each question is put to several models, one at a time, each in a fresh

conversation. No tools and no instructions of ours — the question is the only

thing sent.

Web search is off unless an ask deliberately allows it, and an answer that

searched the web is marked as such wherever it appears. A model may also be

asked the same question several times, because models vary between identical

requests; each reply is kept as its own answer. Every answer is kept in full.

Positions

A position is the move an answer makes: the stance it takes, or declines to

take. It is not a score, and it is not a judgment of quality.

Each question has a short list of positions, and a grader model reads each

answer on its own and assigns one. The lists are not written in advance. They

are written after reading real answers, because there is no way to guess how

four models will answer a question until you have asked it.

When an answer fits none of the listed positions, it is labelled other and the

grader describes the stance it actually saw. That case is the most interesting

thing here: it usually means a model has said something no model said before.

What this can show

What it cannot show

deliberately close to famous puzzles, and from text alone the two cannot be

fully separated.

outside.

own instructions, memory, and tools, so what you see here is close to, but not

the same as, what you would get in the app.

Scale, honestly

This is a handful of answers per question. It is an anthology, not evidence.

Nothing here is a statistic, and no rate computed from it would survive a

statistician. What it offers is the answers themselves, kept in full and side by

side, so you can read them and see for yourself.

The questions

Many of them borrow from instruments built to study human minds — analogy and

structure-mapping, false-belief tasks, prototype theory, the Einstellung effect,

causal reasoning. They are reworded to sound like questions a person would

actually ask.

Whether those frameworks transfer from people to language models is genuinely

contested. The questions are borrowed because they are good at making thinking

visible. Nothing here claims that a model which passes a false-belief task has a

theory of mind.

The grader

The grader is a Claude model, and it labels answers that include Claude's own.

That is a real limitation and worth stating plainly.

Two things limit the damage. The grader sees one answer at a time and is never

told which model produced it. And every label is shown beside the full answer

and the exact lines it was read from — so a label is something you can check,

not something you have to trust.

An earlier version

This project started as something else: a record of what AI assistants told

people about everyday health, money, and legal questions. That version was

built and asked its questions once, in August 2026, before the subject changed

to the one described above.

Those pages are kept, unchanged and never updated, at

/2026-08/. They describe a different project and should not be read

as part of this one.