Factual stability

5 prompts observed across 5 models over 13 weeks.

Median response length — model × week

Median response length in tokens, averaged across the prompts in this axis. Sustained shifts often accompany a model update. Cell shade uses a viridis (colorblind-safe) palette, anchored at zero and scaled to 65; darker = lower, brighter = higher. Numeric values are in each cell for programmatic access. A · means the model was not sampled that week and is not a measurement of zero. Frontier models alternate on a biweekly cadence, so roughly half of their cells are unsampled by design. This axis leads with median response length because that is the measure that moves on it. How this is computed.
Model 2026-W172026-W182026-W192026-W202026-W212026-W222026-W232026-W242026-W252026-W262026-W272026-W282026-W29
gpt-5.5 · · · · · · · · · · median response length 12 · median response length 12
llama3.2:3b · · median response length 34 median response length 35 median response length 34 median response length 34 median response length 34 median response length 32 median response length 34 median response length 31 median response length 34 median response length 30 median response length 34
claude-opus-4-8 · · · · · · · · · · · median response length 65 ·
claude-opus-4-7 · median response length 44 · median response length 44 · median response length 45 · median response length 55 · median response length 58 · · ·
gpt-5.1 median response length 12 · median response length 13 · median response length 12 · median response length 12 · median response length 12 · · · ·

Prompts in this axis