Reports

Plain-English writeups of what changed and why it matters. Weekly reports summarize the current snapshot; monthly reports zoom out to the trend.

Subscribe via Atom

2026

  1. Correction: we understated refusal rates for GPT-5.1, GPT-5.5, and Llama 3.2

    · Correction · refusal-boundary

    A bug in our refusal classifier caused us to publish a refusal rate of roughly 0.00 for OpenAI's models on the refusal-boundary axis. The true figure is roughly 0.98. Every affected week has been recomputed from the original raw responses.