How is the machine’s mind changing?
A public, timestamped, reproducible record of how frontier large language models’ stances, refusals, and framings shift on contested topics — measured weekly, with receipts.
6 models observed across 30 prompts spanning 6 axes. Latest snapshot: 2026-W36.
Notable shifts this week
Largest week-over-week movements across all measured prompts. One card per metric type. Magnitude is normalised against each metric’s reference scale, so refusal-rate, hedge-density, and length shifts can be compared on the same axis.
-
Length (median)
414→188↓ -226 tokllama3.2:3b · political
-
Length (median)
58→77↑ 19 tokllama3.2:3b · neutral control
-
Length (median)
280→326↑ 46 tokllama3.2:3b · historical contested
Latest measurement
Cells colour the dominant week-over-week shift on each axis × model. Colour is normalised within each row, so the brightest cell is the model that drifted most on that axis this week — compare absolute magnitudes via the score in each cell. Click a cell to drill into the axis page. How this is computed. Cells marked with the week label e.g. W17 use the most recent measurement available — frontier models on Level 0 alternate biweekly, so half the columns will be labelled as last-seen data on any given week.
| Axis | claude-opus-4-8 | claude-opus-5 | llama3.2:3b | gpt-5.5 | claude-opus-4-7 | gpt-5.1 |
|---|---|---|---|---|---|---|
| Factual stability | 0.01 W34 | · | 0.02 | 0.05 W35 | 0.08 W26 | 0.07 W25 |
| Historical contested | 0.01 W34 | 0.04 W34 | 0.07 | 0.03 W35 | 0.03 W26 | 0.07 W25 |
| Neutral control | 0.01 W34 | · | 0.17 | 0.12 W35 | 0.10 W26 | 0.03 W25 |
| Political | 0.10 | 0.54 | 0.21 | 0.04 W35 | 0.03 W26 | 0.02 W25 |
| Refusal boundary | 0.03 W34 | · | 0.05 | 0.15 W35 | 0.02 W26 | 0.02 W25 |
| Scientific consensus | 0.02 W34 | 0.02 W34 | 0.02 | 0.10 W35 | 0.02 W26 | 0.01 W25 |
Explore by axis
Observed models
-
claude-opus-4-8
anthropic ·
claude-opus-4-8refusal-rate trend across recent weeks -
claude-opus-5
anthropic ·
claude-opus-5refusal-rate trend across recent weeks -
llama3.2:3b
ollama ·
llama3.2:3brefusal-rate trend across recent weeks -
gpt-5.5
openai ·
gpt-5.5-2026-04-23refusal-rate trend across recent weeks -
claude-opus-4-7
anthropic ·
claude-opus-4-7refusal-rate trend across recent weeks -
gpt-5.1
openai ·
gpt-5.1-2025-11-13refusal-rate trend across recent weeks
How to read this site
- Reports are plain-English writeups of notable changes, written for a general audience.
- Axes group prompts by the kind of drift they measure (political, historical, scientific, refusal-boundary).
- Models lets you drill into a single provider over time.
- Data gives you the raw weekly snapshots in CSV, JSON, and Parquet.
- Methodology documents how the corpus is built, how metrics are computed, and how to reproduce any chart on this site.