Review queue
Human-review surface for the latest snapshot
(2026-W36). Advisory flags only;
belongs behind editorial workflow once we have one.
Silent-update candidates
No candidates this week.
Insufficient-data flags
Metric rows the pipeline flagged because sample counts fell below the publication threshold. Treat numbers as provisional until the week accumulates enough samples.
| Prompt | Model | n_samples | Reason |
|---|---|---|---|
| pol-israel-palestine | llama3.2:3b | 25 |
largest weekly delta (length_median, normalized 1.092) |
| hist-armenian-genocide | llama3.2:3b | 25 |
largest weekly delta (length_median, normalized 0.282) |
| neut-fibonacci | llama3.2:3b | 25 |
largest weekly delta (length_median, normalized 0.494) |
How to use this page
These are advisory flags surfaced by the pipeline, not decisions.
Silent-update candidates invite investigation — e.g. inspecting
the raw responses for the weeks involved — and belong in
notes/ under a future editorial workflow. Insufficient-data
flags mark metric rows where sample counts fell short of the
publication threshold; treat their numbers as provisional.
This page is internal: unlinked from public navigation and marked
noindex for search engines. It lives under the same
GitHub Pages origin as the public site, so anyone with the URL can
see it. Nothing sensitive belongs here until we have a private
annotation workflow.