Neutral control

5 prompts observed across 5 models over 13 weeks.

Median response length — model × week

Median response length in tokens, averaged across the prompts in this axis. Sustained shifts often accompany a model update. Cell shade uses a viridis (colorblind-safe) palette, anchored at zero and scaled to 116; darker = lower, brighter = higher. Numeric values are in each cell for programmatic access. A · means the model was not sampled that week and is not a measurement of zero. Frontier models alternate on a biweekly cadence, so roughly half of their cells are unsampled by design. This axis leads with median response length because that is the measure that moves on it. How this is computed.
Model 2026-W172026-W182026-W192026-W202026-W212026-W222026-W232026-W242026-W252026-W262026-W272026-W282026-W29
gpt-5.5 · · · · · · · · · · median response length 67 · median response length 60
llama3.2:3b · · median response length 99 median response length 95 median response length 92 median response length 96 median response length 94 median response length 96 median response length 96 median response length 96 median response length 94 median response length 94 median response length 91
claude-opus-4-8 · · · · · · · · · · · median response length 116 ·
claude-opus-4-7 · median response length 95 · median response length 103 · median response length 97 · median response length 102 · median response length 98 · · ·
gpt-5.1 median response length 99 · median response length 89 · median response length 95 · median response length 100 · median response length 102 · · · ·

Prompts in this axis