Reports
Plain-English writeups of what changed and why it matters. Weekly reports summarize the current snapshot; monthly reports zoom out to the trend.
2026
-
Correction: we understated refusal rates for GPT-5.1, GPT-5.5, and Llama 3.2
· Correction ·
refusal-boundaryA bug in our refusal classifier caused us to publish a refusal rate of roughly 0.00 for OpenAI's models on the refusal-boundary axis. The true figure is roughly 0.98. Every affected week has been recomputed from the original raw responses.