About
A record of how AI assistants answer the same questions over time.
The method is simple. Each question is put to several AI models in the same words every time, several times over, and asked again on a repeating schedule. Every answer is recorded and kept.
Models are updated and replaced over time. Whether their answers change along the way, and in what direction, is a question that needs a record to answer. This is an attempt to keep one. Past answers cannot be collected later, which is why it starts now rather than once the question is better formed.
What is compared
- Divergence
- Different providers, asked the same question in the same month.
- Certainty
- How much one model's answer varies when the same question is put to it several times on the same day.
- Search use
- Whether the model reached for the web before answering, which it decides for itself.
- Drift
- One provider's default model, month over month. When a provider replaces the model it serves by default, the series continues under the new one and the change is recorded — so a shift here can mean the model changed its mind, or that the provider changed which model you get. Nothing can be said about drift at all until the record covers more than one month.
Limits
Answers are collected through each provider's API, with web search available. That is close to, but not the same as, what their chat products show a user, which add their own instructions and context.
Models vary between runs even when nothing has changed. Each question is therefore asked several times, and what is reported is the pattern across those samples together with how many of them agreed — never a single answer standing in for the model.
Each answer is labelled by a language model working from a written rubric. That is a limitation in itself: the labeller can be wrong, and it is a Claude model labelling answers that include Claude's own. Every answer is published in full, and beside each one our grader quotes the lines it read the label off — the model's words, our selection — so a label can be checked rather than taken on trust.