Data coverage
Every week of the record and, for each model, whether it was due, how much was captured against what was expected, and why anything is missing. This page is generated on every build from the run log, the published manifests and the gap ledger, so it cannot fall out of step with the data the way hand-written notes can.
24 weeks, 52 scheduled model-weeks: 38 complete, 3 partial, 3 degraded, 8 lost. The narrative account of each gap is under known data gaps.
What each status means
- OK
- The model was due and every prompt was measured. The sample
count can still sit a little under the expected figure when a
provider declined some requests; those cells carry a
rejected_samplescount. - Partial
- The model was due and some prompts, or most of the samples, are missing. What was captured is published and every metric on it is real, but the week does not cover the whole corpus for this model.
- Degraded
- Captured, but known to be impaired, for example responses truncated before any visible output. The linked correction says what was done about it.
- Lost
- The model was due and nothing for it is in the record. Lost weeks are not backfilled: a sample taken later is not a sample taken that week.
- Corrected
- A ledger record or a manifest correction that fixed something an earlier record disclosed. The earlier record stays on the page, marked as corrected: the ledger is append-only.
- Not yet published
- The run log records samples for the model that week, but none are in the published record yet.
- Not scheduled
- The model was not due that week. Frontier models alternate by ISO week parity, so this is the normal state of roughly half of their weeks, and models enter and leave the roster over time.
“Expected” is the run’s recorded expectation where the run log has one; for older runs it is the number of prompts times the samples per prompt the model normally receives. A status marked unexplained has no gap ledger entry yet, which is itself a finding; that includes a cell that measured every prompt but is short of the expected samples with no unusable or declined samples to account for the difference.
Week by model
| Week | llama3.2:3b |
gpt-5.1 |
claude-opus-4-7 |
gpt-5.5 |
claude-opus-4-8 |
claude-opus-5 |
|---|---|---|---|---|---|---|
2026-W40 |
OK
750 / 750
|
Not scheduled | Not scheduled | Not scheduled |
OK
600 / 600
|
OK
600 / 600
|
2026-W39 |
OK
750 / 750
|
Not scheduled | Not scheduled |
OK
588 / 600
|
Not scheduled | Not scheduled |
2026-W38 |
OK
750 / 750
|
Not scheduled | Not scheduled | Not scheduled |
Lost
0 / 600
|
Lost
0 / 600
|
2026-W37 |
OK
750 / 750
|
Not scheduled | Not scheduled |
OK
596 / 600
|
Not scheduled | Not scheduled |
2026-W36 |
OK
750 / 750
|
Not scheduled | Not scheduled | Not scheduled |
Partial
32 / 600
|
Partial
13 / 600
|
2026-W35 |
OK
750 / 750
|
Not scheduled | Not scheduled |
OK
583 / 600
|
Not scheduled | Not scheduled |
2026-W34 |
OK
750 / 750
|
Not scheduled | Not scheduled | Not scheduled |
OK
600 / 600
|
Partial
259 / 600
|
2026-W33 |
OK
750 / 750
|
Not scheduled | Not scheduled |
Degraded
582 / 600
|
Not scheduled | Not scheduled |
2026-W32 |
OK
750 / 750
|
Not scheduled | Not scheduled | Not scheduled |
OK
600 / 600
|
Not scheduled |
2026-W31 |
Lost
0 / 750
|
Not scheduled | Not scheduled |
Lost
0 / 600
|
Not scheduled | Not scheduled |
2026-W30 |
Lost
0 / 750
|
Not scheduled | Not scheduled | Not scheduled |
Lost
0 / 600
|
Not scheduled |
2026-W29 |
OK
750 / 750
|
Not scheduled | Not scheduled |
Degraded
557 / 600
|
Not scheduled | Not scheduled |
2026-W28 |
OK
750 / 750
|
Not scheduled | Not scheduled | Not scheduled |
OK
600 / 600
|
Not scheduled |
2026-W27 |
OK
750 / 750
|
Not scheduled | Not scheduled |
Degraded
553 / 600
|
Not scheduled | Not scheduled |
2026-W26 |
OK
750 / 750
|
Not scheduled |
OK
600 / 600
|
Not scheduled | Not scheduled | Not scheduled |
2026-W25 |
OK
750 / 750
|
OK
750 / 750
|
Not scheduled | Not scheduled | Not scheduled | Not scheduled |
2026-W24 |
OK
750 / 750
|
Not scheduled |
OK
600 / 600
|
Not scheduled | Not scheduled | Not scheduled |
2026-W23 |
OK
750 / 750
|
OK
750 / 750
|
Not scheduled | Not scheduled | Not scheduled | Not scheduled |
2026-W22 |
OK
750 / 750
|
Not scheduled |
OK
600 / 600
|
Not scheduled | Not scheduled | Not scheduled |
2026-W21 |
OK
750 / 750
|
OK
750 / 750
|
Not scheduled | Not scheduled | Not scheduled | Not scheduled |
2026-W20 |
OK
750 / 750
|
Not scheduled |
OK
600 / 600
|
Not scheduled | Not scheduled | Not scheduled |
2026-W19 |
OK
750 / 750
|
OK
750 / 750
|
Not scheduled | Not scheduled | Not scheduled | Not scheduled |
2026-W18 |
Lost
0 / 750
|
Not scheduled |
OK
600 / 600
|
Not scheduled | Not scheduled | Not scheduled |
2026-W17 |
Lost
0 / 750
|
OK
750 / 750
|
Not scheduled | Not scheduled | Not scheduled | Not scheduled |
Week by week
2026-W40
| Model | Status | Prompts | Samples | Run log | Why |
|---|---|---|---|---|---|
llama3.2:3b |
OK | 30 / 30 |
750 / 750 |
750 written |
|
claude-opus-4-8 |
OK | 30 / 30 |
600 / 600 |
600 written |
|
claude-opus-5 |
OK | 30 / 30 |
600 / 600 |
600 written |
2026-W39
| Model | Status | Prompts | Samples | Run log | Why |
|---|---|---|---|---|---|
llama3.2:3b |
OK | 30 / 30 |
750 / 750 |
750 written |
|
gpt-5.5 |
OK | 30 / 30 |
588 / 600 |
588 written |
12 requests were declined by the provider and never ran. |
2026-W38
Stance classifier Lost The stance classifier was still on the empty Anthropic balance. All 10 llama3.2:3b stance-bearing cells are published as stance n/a at confidence 0.0, which means unmeasured, not neutral. Since corrected; see the correction record for this week.Evidence: issue #38.
Stance classifier Corrected Stance for the 10 unmeasured cells was re-classified on 2026-10-07 from the published responses and grafted as a versioned correction; only the stance fields changed. Evidence: /reports/2026-10-07-stance-classifier-correction/.
Correction 2026-10-07 Corrected Stance re-classified for 10 stance-bearing cell(s) published as na at confidence 0.0 because every classifier call failed on an exhausted Anthropic balance. Classified 2026-10-07 with anthropic/claude-haiku-4-5-20251001 from this week's published responses, using the same representative-response rule. No other field changed. 10 row(s) changed.Correction report.
| Model | Status | Prompts | Samples | Run log | Why |
|---|---|---|---|---|---|
llama3.2:3b |
OK | 30 / 30 |
750 / 750 |
750 written |
|
claude-opus-4-8 |
Lost | 0 / 30 |
0 / 600 |
0 written |
Prepaid Anthropic API credit was still empty: 0 of 600 samples. Both Opus models failed all 60 of their pairs. Evidence: issue #38. |
claude-opus-5 |
Lost | 0 / 30 |
0 / 600 |
0 written |
Prepaid Anthropic API credit was still empty: 0 of 600 samples. Both Opus models failed all 60 of their pairs. Evidence: issue #38. |
2026-W37
Stance classifier Lost The stance classifier was still on the empty Anthropic balance. gpt-5.5 10 of 10 and llama3.2:3b 9 of 10 stance-bearing cells are published as stance n/a at confidence 0.0, which means unmeasured, not neutral. The remaining llama3.2:3b cell was answered from the classifier cache. Since corrected; see the correction record for this week.Evidence: issue #36.
Stance classifier Corrected Stance for the 19 unmeasured cells was re-classified on 2026-10-07 from the published responses and grafted as a versioned correction; only the stance fields changed. Evidence: /reports/2026-10-07-stance-classifier-correction/.
Correction 2026-10-07 Corrected Stance re-classified for 19 stance-bearing cell(s) published as na at confidence 0.0 because every classifier call failed on an exhausted Anthropic balance. Classified 2026-10-07 with anthropic/claude-haiku-4-5-20251001 from this week's published responses, using the same representative-response rule. No other field changed. 19 row(s) changed.Correction report.
| Model | Status | Prompts | Samples | Run log | Why |
|---|---|---|---|---|---|
llama3.2:3b |
OK | 30 / 30 |
750 / 750 |
750 written |
|
gpt-5.5 |
OK | 30 / 30 |
596 / 600 |
596 written |
4 requests were declined by the provider and never ran. |
2026-W36
Stance classifier Lost The stance classifier uses the same Anthropic account and failed on every call. All 13 stance-bearing cells (claude-opus-4-8 2, claude-opus-5 1, llama3.2:3b 10) are published as stance n/a at confidence 0.0, which means unmeasured, not neutral. Since corrected; see the correction record for this week.Evidence: issue #36.
Stance classifier Corrected Stance for the 13 unmeasured cells was re-classified on 2026-10-07 from the published responses and grafted as a versioned correction; only the stance fields changed. Evidence: /reports/2026-10-07-stance-classifier-correction/.
Correction 2026-10-07 Corrected Stance re-classified for 13 stance-bearing cell(s) published as na at confidence 0.0 because every classifier call failed on an exhausted Anthropic balance. Classified 2026-10-07 with anthropic/claude-haiku-4-5-20251001 from this week's published responses, using the same representative-response rule. No other field changed. 13 row(s) changed.Correction report.
| Model | Status | Prompts | Samples | Run log | Why |
|---|---|---|---|---|---|
llama3.2:3b |
OK | 30 / 30 |
750 / 750 |
750 written |
|
claude-opus-4-8 |
Partial | 2 / 30 |
32 / 600 |
32 written |
Prepaid Anthropic API credit ran out within two minutes of the start. claude-opus-4-8 has 32 of 600 samples, over 2 of 30 prompts; the two Opus models together have 45 of 1200 and 59 pairs failed. Its drift comparisons are against 2026-W34. Evidence: issue #36. |
claude-opus-5 |
Partial | 1 / 30 |
13 / 600 |
13 written |
Prepaid Anthropic API credit ran out within two minutes of the start. claude-opus-5 has 13 of 600 samples, on 1 of 30 prompts; the two Opus models together have 45 of 1200 and 59 pairs failed. Its drift comparison is against 2026-W34. Evidence: issue #36. |
2026-W35
| Model | Status | Prompts | Samples | Run log | Why |
|---|---|---|---|---|---|
llama3.2:3b |
OK | 30 / 30 |
750 / 750 |
750 written |
Drift for all 30 llama3.2:3b cells was computed against 2026-W34, whose metrics this week's manifest carried in its history (and so were served under /data/2026-W34/) although 2026-W34 was not then published as a week of its own. It is now published as a partial week. Evidence: |
gpt-5.5 |
OK | 30 / 30 |
583 / 600 |
583 written |
17 requests were declined by the provider and never ran. |
2026-W34
Run log entry reconstructed after the fact: Reconstruction (recovery: true), written by recover-week, not by the run. 2026-W34 was sampled 2026-08-24 (09:03 to 10:00 UTC, first and last captured sample) and was killed at the one-hour SSM execution timeout, so it left no run_log entry. Raw samples archived to S3 2026-10-05; manifest and this entry built 2026-10-07 from them. No sample was taken by this invocation; counts are what the archive holds, a pair counting as complete only with every sample it owed: ollama/llama3.2:3b 750/750 samples, 30/30 prompts complete; anthropic/claude-opus-4-8 600/600 samples, 30/30 prompts complete; anthropic/claude-opus-5 259/600 samples, 12/30 prompts complete. started_at and finished_at are the first and last captured sample. runners is the roster due in 2026-W34, not the config the rebuild ran under. config_hash e78efdfab25cd47d is the hash logged by the scheduled runs either side of 2026-W34, under the same config; it was not computed by the run itself. host and pid are those of the recover-week invocation.
Published as a partial week. From the week’s manifest:
- Partial week. 2026-W34 was sampled on 2026-08-24 and was killed at the one-hour SSM execution timeout before every runner finished. This manifest publishes what was captured.
- claude-opus-5 has 259 of 600 samples, 12 of 30 prompts complete; cut short: sci-iq-heritability (19/20); 17 prompt(s) never sampled.
- Complete: llama3.2:3b (750 samples), claude-opus-4-8 (600 samples).
- Built 2026-10-07 from the raw samples archived to S3 on 2026-10-05, with the same code a live run uses and a fixed bootstrap seed (20261005), so confidence intervals and p-values can differ trivially from the copies of 2026-W34 already embedded in later weeks' history, which were drawn unseeded.
| Model | Status | Prompts | Samples | Run log | Why |
|---|---|---|---|---|---|
llama3.2:3b |
OK | 30 / 30 |
750 / 750 |
750 written |
|
claude-opus-4-8 |
OK | 30 / 30 |
600 / 600 |
600 written |
|
claude-opus-5 |
Partial | 13 / 30 |
259 / 600 |
259 written |
Sampled on 2026-08-24 and killed at the one-hour SSM execution timeout. claude-opus-5 has 259 of 600 samples, over 13 of 30 prompts. claude-opus-4-8 (600 samples) and llama3.2:3b (750 samples) completed. The raw samples were archived on 2026-10-05 and the week is published as a partial week. Evidence: issue #32; issue #34; commit 8291b6c. |
2026-W33
All runners Capacity retries delayed the start until 16:04 UTC, after the publish step had already looked for the week and found nothing. The week was published on 2026-08-25, 8 days late. One pair failed; see the openai/gpt-5.5 line. Evidence: issue #29; commit 14c0187.
| Model | Status | Prompts | Samples | Run log | Why |
|---|---|---|---|---|---|
llama3.2:3b |
OK | 30 / 30 |
750 / 750 |
750 written |
|
gpt-5.5 |
Degraded | 30 / 30 |
582 / 600 |
582 written |
One pair failed: ref-wifi-unauthorized has 2 of 20 samples, after the provider rejected the rest with HTTP 400 ("flagged for possible cybersecurity risk"). This was before declined requests were counted as rejected_samples, so the published cell carries none. gpt-5.5 has 582 of 600 samples; every prompt was measured. Evidence: |
2026-W32
| Model | Status | Prompts | Samples | Run log | Why |
|---|---|---|---|---|---|
llama3.2:3b |
OK | 30 / 30 |
750 / 750 |
750 written |
|
claude-opus-4-8 |
OK | 30 / 30 |
600 / 600 |
600 written |
2026-W31
No data for 2026-W31. The run log has no entry for this week.
| Model | Status | Prompts | Samples | Run log | Why |
|---|---|---|---|---|---|
llama3.2:3b |
Lost | 0 / 30 |
0 / 750 |
no entry |
The run never started: EC2 InsufficientInstanceCapacity again, the week after 2026-W30. Evidence: issue #26; commit f352103; commit 8b30b78. |
gpt-5.5 |
Lost | 0 / 30 |
0 / 600 |
no entry |
The run never started: EC2 InsufficientInstanceCapacity again, the week after 2026-W30. Evidence: issue #26; commit f352103; commit 8b30b78. |
2026-W30
No data for 2026-W30. The run log has no entry for this week.
| Model | Status | Prompts | Samples | Run log | Why |
|---|---|---|---|---|---|
llama3.2:3b |
Lost | 0 / 30 |
0 / 750 |
no entry |
The run never started: the sampling instance could not be launched (EC2 InsufficientInstanceCapacity), and the orchestrator failed before reaching any alert. Evidence: issue #24; commit f352103; commit 8b30b78. |
claude-opus-4-8 |
Lost | 0 / 30 |
0 / 600 |
no entry |
The run never started: the sampling instance could not be launched (EC2 InsufficientInstanceCapacity), and the orchestrator failed before reaching any alert. Evidence: issue #24; commit f352103; commit 8b30b78. |
2026-W29
| Model | Status | Prompts | Samples | Run log | Why |
|---|---|---|---|---|---|
llama3.2:3b |
OK | 30 / 30 |
750 / 750 |
750 written |
|
gpt-5.5 |
Degraded | 29 / 30 |
557 / 600 |
600 written |
43 of 600 samples were empty, truncated at the token limit, and sci-iq-heritability was not measured at all. The empty samples were removed from every metric in the 2026-07-24 correction. 23 captured samples had no usable content and are excluded from every metric. Evidence: /reports/2026-07-24-truncated-response-correction/; commit de2ec54. |
2026-W28
| Model | Status | Prompts | Samples | Run log | Why |
|---|---|---|---|---|---|
llama3.2:3b |
OK | 30 / 30 |
750 / 750 |
750 written |
|
claude-opus-4-8 |
OK | 30 / 30 |
600 / 600 |
600 written |
2026-W27
| Model | Status | Prompts | Samples | Run log | Why |
|---|---|---|---|---|---|
llama3.2:3b |
OK | 30 / 30 |
750 / 750 |
750 written |
|
gpt-5.5 |
Degraded | 29 / 30 |
553 / 600 |
600 written |
30 temperature-0 pairs failed with HTTP 400. 47 of 600 samples were empty, truncated at the token limit, and sci-iq-heritability was not measured at all. The empty samples were removed from every metric in the 2026-07-24 correction. 27 captured samples had no usable content and are excluded from every metric. Evidence: /reports/2026-07-24-truncated-response-correction/; commit de2ec54. |
2026-W26
| Model | Status | Prompts | Samples | Run log | Why |
|---|---|---|---|---|---|
llama3.2:3b |
OK | 30 / 30 |
750 / 750 |
750 written |
|
claude-opus-4-7 |
OK | 30 / 30 |
600 / 600 |
600 written |
2026-W25
| Model | Status | Prompts | Samples | Run log | Why |
|---|---|---|---|---|---|
llama3.2:3b |
OK | 30 / 30 |
750 / 750 |
750 written |
|
gpt-5.1 |
OK | 30 / 30 |
750 / 750 |
750 written |
2026-W24
| Model | Status | Prompts | Samples | Run log | Why |
|---|---|---|---|---|---|
llama3.2:3b |
OK | 30 / 30 |
750 / 750 |
750 written |
|
claude-opus-4-7 |
OK | 30 / 30 |
600 / 600 |
600 written |
2026-W23
| Model | Status | Prompts | Samples | Run log | Why |
|---|---|---|---|---|---|
llama3.2:3b |
OK | 30 / 30 |
750 / 750 |
750 written |
|
gpt-5.1 |
OK | 30 / 30 |
750 / 750 |
750 written |
2026-W22
| Model | Status | Prompts | Samples | Run log | Why |
|---|---|---|---|---|---|
llama3.2:3b |
OK | 30 / 30 |
750 / 750 |
750 written |
|
claude-opus-4-7 |
OK | 30 / 30 |
600 / 600 |
600 written |
2026-W21
| Model | Status | Prompts | Samples | Run log | Why |
|---|---|---|---|---|---|
llama3.2:3b |
OK | 30 / 30 |
750 / 750 |
750 written |
|
gpt-5.1 |
OK | 30 / 30 |
750 / 750 |
750 written |
2026-W20
| Model | Status | Prompts | Samples | Run log | Why |
|---|---|---|---|---|---|
llama3.2:3b |
OK | 30 / 30 |
750 / 750 |
750 written |
|
claude-opus-4-7 |
OK | 30 / 30 |
600 / 600 |
600 written |
2026-W19
All runners
The published W19 snapshot (gpt-5.1 750 samples and llama3.2:3b 750 samples, captured 2026-05-06) has no matching run log entry. The two logged W19 runs are a 2026-05-04 workstation llama3.2:3b run whose samples were not published and a 2026-05-11 resume that skipped all 60 pairs. Why the capturing run left no log entry is not established.
Evidence: snapshot captured_at 2026-05-06; commit 8f29d20; commit 6f3221f.
| Model | Status | Prompts | Samples | Run log | Why |
|---|---|---|---|---|---|
llama3.2:3b |
OK | 30 / 30 |
750 / 750 |
750 written |
|
gpt-5.1 |
OK | 30 / 30 |
750 / 750 |
0 written |
2026-W18
| Model | Status | Prompts | Samples | Run log | Why |
|---|---|---|---|---|---|
llama3.2:3b |
Lost | 0 / 30 |
0 / 750 |
750 written |
No llama3.2:3b samples in the published W18 record. A workstation run labelled W18 wrote 750 samples on 2026-04-25, which falls in ISO week 2026-W17, so they were not captured in W18 and were not published. Evidence: |
claude-opus-4-7 |
OK | 30 / 30 |
600 / 600 |
1200 written |
Two runs carry the W18 label. The first, on 2026-04-24 (600 samples, captured in ISO week 2026-W17), was not published. The published W18 data is the second run, captured 2026-05-03 (600 samples). Evidence: |
2026-W17
| Model | Status | Prompts | Samples | Run log | Why |
|---|---|---|---|---|---|
llama3.2:3b |
Lost | 0 / 30 |
0 / 750 |
0 written |
No llama3.2:3b samples. The W17 run ran on a hosted CI runner with no ollama installed, so only openai/gpt-5.1 was sampled (750 samples). The failure left no error in the run log and was caught on review the following week. Evidence: |
gpt-5.1 |
OK | 30 / 30 |
750 / 750 |
750 written |
Files
| File | Size | Contents |
|---|---|---|
coverage.csv |
7426 B |
One row per scheduled or published (week, model): status, prompts and samples against expected, run log count, reasons and evidence. |
coverage.jsonl |
24303 B |
The same rows, one JSON object per line. |
gaps.jsonl |
8444 B |
The gap ledger exactly as committed: one hand-recorded explanation per line, append-only. |
SHA256SUMS |
– | SHA-256 digest for every file above. |
Each gap ledger line carries week_id, scope
(all, a provider/model, or
stance for the stance classifier), kind
(lost, partial, degraded,
note, or corrected for a later correction
of an earlier record with the same week and scope), reason, evidence and
recorded_at. Lines are added, never edited; a later line
can extend or correct an earlier one.