Data coverage

Every week of the record and, for each model, whether it was due, how much was captured against what was expected, and why anything is missing. This page is generated on every build from the run log, the published manifests and the gap ledger, so it cannot fall out of step with the data the way hand-written notes can.

24 weeks, 52 scheduled model-weeks: 38 complete, 3 partial, 3 degraded, 8 lost. The narrative account of each gap is under known data gaps.

What each status means

OK
The model was due and every prompt was measured. The sample count can still sit a little under the expected figure when a provider declined some requests; those cells carry a rejected_samples count.
Partial
The model was due and some prompts, or most of the samples, are missing. What was captured is published and every metric on it is real, but the week does not cover the whole corpus for this model.
Degraded
Captured, but known to be impaired, for example responses truncated before any visible output. The linked correction says what was done about it.
Lost
The model was due and nothing for it is in the record. Lost weeks are not backfilled: a sample taken later is not a sample taken that week.
Corrected
A ledger record or a manifest correction that fixed something an earlier record disclosed. The earlier record stays on the page, marked as corrected: the ledger is append-only.
Not yet published
The run log records samples for the model that week, but none are in the published record yet.
Not scheduled
The model was not due that week. Frontier models alternate by ISO week parity, so this is the normal state of roughly half of their weeks, and models enter and leave the roster over time.

“Expected” is the run’s recorded expectation where the run log has one; for older runs it is the number of prompts times the samples per prompt the model normally receives. A status marked unexplained has no gap ledger entry yet, which is itself a finding; that includes a cell that measured every prompt but is short of the expected samples with no unusable or declined samples to account for the difference.

Week by model

Status and samples published against samples expected, newest week first.
Week llama3.2:3b gpt-5.1 claude-opus-4-7 gpt-5.5 claude-opus-4-8 claude-opus-5
2026-W40 OK
750 / 750
Not scheduled Not scheduled Not scheduled OK
600 / 600
OK
600 / 600
2026-W39 OK
750 / 750
Not scheduled Not scheduled OK
588 / 600
Not scheduled Not scheduled
2026-W38 OK
750 / 750
Not scheduled Not scheduled Not scheduled Lost
0 / 600
Lost
0 / 600
2026-W37 OK
750 / 750
Not scheduled Not scheduled OK
596 / 600
Not scheduled Not scheduled
2026-W36 OK
750 / 750
Not scheduled Not scheduled Not scheduled Partial
32 / 600
Partial
13 / 600
2026-W35 OK
750 / 750
Not scheduled Not scheduled OK
583 / 600
Not scheduled Not scheduled
2026-W34 OK
750 / 750
Not scheduled Not scheduled Not scheduled OK
600 / 600
Partial
259 / 600
2026-W33 OK
750 / 750
Not scheduled Not scheduled Degraded
582 / 600
Not scheduled Not scheduled
2026-W32 OK
750 / 750
Not scheduled Not scheduled Not scheduled OK
600 / 600
Not scheduled
2026-W31 Lost
0 / 750
Not scheduled Not scheduled Lost
0 / 600
Not scheduled Not scheduled
2026-W30 Lost
0 / 750
Not scheduled Not scheduled Not scheduled Lost
0 / 600
Not scheduled
2026-W29 OK
750 / 750
Not scheduled Not scheduled Degraded
557 / 600
Not scheduled Not scheduled
2026-W28 OK
750 / 750
Not scheduled Not scheduled Not scheduled OK
600 / 600
Not scheduled
2026-W27 OK
750 / 750
Not scheduled Not scheduled Degraded
553 / 600
Not scheduled Not scheduled
2026-W26 OK
750 / 750
Not scheduled OK
600 / 600
Not scheduled Not scheduled Not scheduled
2026-W25 OK
750 / 750
OK
750 / 750
Not scheduled Not scheduled Not scheduled Not scheduled
2026-W24 OK
750 / 750
Not scheduled OK
600 / 600
Not scheduled Not scheduled Not scheduled
2026-W23 OK
750 / 750
OK
750 / 750
Not scheduled Not scheduled Not scheduled Not scheduled
2026-W22 OK
750 / 750
Not scheduled OK
600 / 600
Not scheduled Not scheduled Not scheduled
2026-W21 OK
750 / 750
OK
750 / 750
Not scheduled Not scheduled Not scheduled Not scheduled
2026-W20 OK
750 / 750
Not scheduled OK
600 / 600
Not scheduled Not scheduled Not scheduled
2026-W19 OK
750 / 750
OK
750 / 750
Not scheduled Not scheduled Not scheduled Not scheduled
2026-W18 Lost
0 / 750
Not scheduled OK
600 / 600
Not scheduled Not scheduled Not scheduled
2026-W17 Lost
0 / 750
OK
750 / 750
Not scheduled Not scheduled Not scheduled Not scheduled

Week by week

2026-W40

Data for 2026-W40.

Models due in 2026-W40.
Model Status Prompts Samples Run log Why
llama3.2:3b OK 30 / 30 750 / 750 750 written
claude-opus-4-8 OK 30 / 30 600 / 600 600 written
claude-opus-5 OK 30 / 30 600 / 600 600 written

2026-W39

Data for 2026-W39.

Models due in 2026-W39.
Model Status Prompts Samples Run log Why
llama3.2:3b OK 30 / 30 750 / 750 750 written
gpt-5.5 OK 30 / 30 588 / 600 588 written

12 requests were declined by the provider and never ran.

2026-W38

Data for 2026-W38.

Stance classifier Lost The stance classifier was still on the empty Anthropic balance. All 10 llama3.2:3b stance-bearing cells are published as stance n/a at confidence 0.0, which means unmeasured, not neutral. Since corrected; see the correction record for this week.Evidence: issue #38.

Stance classifier Corrected Stance for the 10 unmeasured cells was re-classified on 2026-10-07 from the published responses and grafted as a versioned correction; only the stance fields changed. Evidence: /reports/2026-10-07-stance-classifier-correction/.

Correction 2026-10-07 Corrected Stance re-classified for 10 stance-bearing cell(s) published as na at confidence 0.0 because every classifier call failed on an exhausted Anthropic balance. Classified 2026-10-07 with anthropic/claude-haiku-4-5-20251001 from this week's published responses, using the same representative-response rule. No other field changed. 10 row(s) changed.Correction report.

Models due in 2026-W38.
Model Status Prompts Samples Run log Why
llama3.2:3b OK 30 / 30 750 / 750 750 written
claude-opus-4-8 Lost 0 / 30 0 / 600 0 written

Prepaid Anthropic API credit was still empty: 0 of 600 samples. Both Opus models failed all 60 of their pairs.

Evidence: issue #38.

claude-opus-5 Lost 0 / 30 0 / 600 0 written

Prepaid Anthropic API credit was still empty: 0 of 600 samples. Both Opus models failed all 60 of their pairs.

Evidence: issue #38.

2026-W37

Data for 2026-W37.

Stance classifier Lost The stance classifier was still on the empty Anthropic balance. gpt-5.5 10 of 10 and llama3.2:3b 9 of 10 stance-bearing cells are published as stance n/a at confidence 0.0, which means unmeasured, not neutral. The remaining llama3.2:3b cell was answered from the classifier cache. Since corrected; see the correction record for this week.Evidence: issue #36.

Stance classifier Corrected Stance for the 19 unmeasured cells was re-classified on 2026-10-07 from the published responses and grafted as a versioned correction; only the stance fields changed. Evidence: /reports/2026-10-07-stance-classifier-correction/.

Correction 2026-10-07 Corrected Stance re-classified for 19 stance-bearing cell(s) published as na at confidence 0.0 because every classifier call failed on an exhausted Anthropic balance. Classified 2026-10-07 with anthropic/claude-haiku-4-5-20251001 from this week's published responses, using the same representative-response rule. No other field changed. 19 row(s) changed.Correction report.

Models due in 2026-W37.
Model Status Prompts Samples Run log Why
llama3.2:3b OK 30 / 30 750 / 750 750 written
gpt-5.5 OK 30 / 30 596 / 600 596 written

4 requests were declined by the provider and never ran.

2026-W36

Data for 2026-W36.

Stance classifier Lost The stance classifier uses the same Anthropic account and failed on every call. All 13 stance-bearing cells (claude-opus-4-8 2, claude-opus-5 1, llama3.2:3b 10) are published as stance n/a at confidence 0.0, which means unmeasured, not neutral. Since corrected; see the correction record for this week.Evidence: issue #36.

Stance classifier Corrected Stance for the 13 unmeasured cells was re-classified on 2026-10-07 from the published responses and grafted as a versioned correction; only the stance fields changed. Evidence: /reports/2026-10-07-stance-classifier-correction/.

Correction 2026-10-07 Corrected Stance re-classified for 13 stance-bearing cell(s) published as na at confidence 0.0 because every classifier call failed on an exhausted Anthropic balance. Classified 2026-10-07 with anthropic/claude-haiku-4-5-20251001 from this week's published responses, using the same representative-response rule. No other field changed. 13 row(s) changed.Correction report.

Models due in 2026-W36.
Model Status Prompts Samples Run log Why
llama3.2:3b OK 30 / 30 750 / 750 750 written
claude-opus-4-8 Partial 2 / 30 32 / 600 32 written

Prepaid Anthropic API credit ran out within two minutes of the start. claude-opus-4-8 has 32 of 600 samples, over 2 of 30 prompts; the two Opus models together have 45 of 1200 and 59 pairs failed. Its drift comparisons are against 2026-W34.

Evidence: issue #36.

claude-opus-5 Partial 1 / 30 13 / 600 13 written

Prepaid Anthropic API credit ran out within two minutes of the start. claude-opus-5 has 13 of 600 samples, on 1 of 30 prompts; the two Opus models together have 45 of 1200 and 59 pairs failed. Its drift comparison is against 2026-W34.

Evidence: issue #36.

2026-W35

Data for 2026-W35.

Models due in 2026-W35.
Model Status Prompts Samples Run log Why
llama3.2:3b OK 30 / 30 750 / 750 750 written

Drift for all 30 llama3.2:3b cells was computed against 2026-W34, whose metrics this week's manifest carried in its history (and so were served under /data/2026-W34/) although 2026-W34 was not then published as a week of its own. It is now published as a partial week.

Evidence: data/manifests/2026-W35.json history.

gpt-5.5 OK 30 / 30 583 / 600 583 written

17 requests were declined by the provider and never ran.

2026-W34

Data for 2026-W34.

Run log entry reconstructed after the fact: Reconstruction (recovery: true), written by recover-week, not by the run. 2026-W34 was sampled 2026-08-24 (09:03 to 10:00 UTC, first and last captured sample) and was killed at the one-hour SSM execution timeout, so it left no run_log entry. Raw samples archived to S3 2026-10-05; manifest and this entry built 2026-10-07 from them. No sample was taken by this invocation; counts are what the archive holds, a pair counting as complete only with every sample it owed: ollama/llama3.2:3b 750/750 samples, 30/30 prompts complete; anthropic/claude-opus-4-8 600/600 samples, 30/30 prompts complete; anthropic/claude-opus-5 259/600 samples, 12/30 prompts complete. started_at and finished_at are the first and last captured sample. runners is the roster due in 2026-W34, not the config the rebuild ran under. config_hash e78efdfab25cd47d is the hash logged by the scheduled runs either side of 2026-W34, under the same config; it was not computed by the run itself. host and pid are those of the recover-week invocation.

Published as a partial week. From the week’s manifest:

  • Partial week. 2026-W34 was sampled on 2026-08-24 and was killed at the one-hour SSM execution timeout before every runner finished. This manifest publishes what was captured.
  • claude-opus-5 has 259 of 600 samples, 12 of 30 prompts complete; cut short: sci-iq-heritability (19/20); 17 prompt(s) never sampled.
  • Complete: llama3.2:3b (750 samples), claude-opus-4-8 (600 samples).
  • Built 2026-10-07 from the raw samples archived to S3 on 2026-10-05, with the same code a live run uses and a fixed bootstrap seed (20261005), so confidence intervals and p-values can differ trivially from the copies of 2026-W34 already embedded in later weeks' history, which were drawn unseeded.
Models due in 2026-W34.
Model Status Prompts Samples Run log Why
llama3.2:3b OK 30 / 30 750 / 750 750 written
claude-opus-4-8 OK 30 / 30 600 / 600 600 written
claude-opus-5 Partial 13 / 30 259 / 600 259 written

Sampled on 2026-08-24 and killed at the one-hour SSM execution timeout. claude-opus-5 has 259 of 600 samples, over 13 of 30 prompts. claude-opus-4-8 (600 samples) and llama3.2:3b (750 samples) completed. The raw samples were archived on 2026-10-05 and the week is published as a partial week.

Evidence: issue #32; issue #34; commit 8291b6c.

2026-W33

Data for 2026-W33.

All runners Capacity retries delayed the start until 16:04 UTC, after the publish step had already looked for the week and found nothing. The week was published on 2026-08-25, 8 days late. One pair failed; see the openai/gpt-5.5 line. Evidence: issue #29; commit 14c0187.

Models due in 2026-W33.
Model Status Prompts Samples Run log Why
llama3.2:3b OK 30 / 30 750 / 750 750 written
gpt-5.5 Degraded 30 / 30 582 / 600 582 written

One pair failed: ref-wifi-unauthorized has 2 of 20 samples, after the provider rejected the rest with HTTP 400 ("flagged for possible cybersecurity risk"). This was before declined requests were counted as rejected_samples, so the published cell carries none. gpt-5.5 has 582 of 600 samples; every prompt was measured.

Evidence: run_log 2026-08-17T16:04:33+00:00; data/manifests/2026-W33.json.

2026-W32

Data for 2026-W32.

Models due in 2026-W32.
Model Status Prompts Samples Run log Why
llama3.2:3b OK 30 / 30 750 / 750 750 written
claude-opus-4-8 OK 30 / 30 600 / 600 600 written

2026-W31

No data for 2026-W31. The run log has no entry for this week.

Models due in 2026-W31.
Model Status Prompts Samples Run log Why
llama3.2:3b Lost 0 / 30 0 / 750 no entry

The run never started: EC2 InsufficientInstanceCapacity again, the week after 2026-W30.

Evidence: issue #26; commit f352103; commit 8b30b78.

gpt-5.5 Lost 0 / 30 0 / 600 no entry

The run never started: EC2 InsufficientInstanceCapacity again, the week after 2026-W30.

Evidence: issue #26; commit f352103; commit 8b30b78.

2026-W30

No data for 2026-W30. The run log has no entry for this week.

Models due in 2026-W30.
Model Status Prompts Samples Run log Why
llama3.2:3b Lost 0 / 30 0 / 750 no entry

The run never started: the sampling instance could not be launched (EC2 InsufficientInstanceCapacity), and the orchestrator failed before reaching any alert.

Evidence: issue #24; commit f352103; commit 8b30b78.

claude-opus-4-8 Lost 0 / 30 0 / 600 no entry

The run never started: the sampling instance could not be launched (EC2 InsufficientInstanceCapacity), and the orchestrator failed before reaching any alert.

Evidence: issue #24; commit f352103; commit 8b30b78.

2026-W29

Data for 2026-W29.

Models due in 2026-W29.
Model Status Prompts Samples Run log Why
llama3.2:3b OK 30 / 30 750 / 750 750 written
gpt-5.5 Degraded 29 / 30 557 / 600 600 written

43 of 600 samples were empty, truncated at the token limit, and sci-iq-heritability was not measured at all. The empty samples were removed from every metric in the 2026-07-24 correction.

23 captured samples had no usable content and are excluded from every metric.

Evidence: /reports/2026-07-24-truncated-response-correction/; commit de2ec54.

2026-W28

Data for 2026-W28.

Models due in 2026-W28.
Model Status Prompts Samples Run log Why
llama3.2:3b OK 30 / 30 750 / 750 750 written
claude-opus-4-8 OK 30 / 30 600 / 600 600 written

2026-W27

Data for 2026-W27.

Models due in 2026-W27.
Model Status Prompts Samples Run log Why
llama3.2:3b OK 30 / 30 750 / 750 750 written
gpt-5.5 Degraded 29 / 30 553 / 600 600 written

30 temperature-0 pairs failed with HTTP 400. 47 of 600 samples were empty, truncated at the token limit, and sci-iq-heritability was not measured at all. The empty samples were removed from every metric in the 2026-07-24 correction.

27 captured samples had no usable content and are excluded from every metric.

Evidence: /reports/2026-07-24-truncated-response-correction/; commit de2ec54.

2026-W26

Data for 2026-W26.

Models due in 2026-W26.
Model Status Prompts Samples Run log Why
llama3.2:3b OK 30 / 30 750 / 750 750 written
claude-opus-4-7 OK 30 / 30 600 / 600 600 written

2026-W25

Data for 2026-W25.

Models due in 2026-W25.
Model Status Prompts Samples Run log Why
llama3.2:3b OK 30 / 30 750 / 750 750 written
gpt-5.1 OK 30 / 30 750 / 750 750 written

2026-W24

Data for 2026-W24.

Models due in 2026-W24.
Model Status Prompts Samples Run log Why
llama3.2:3b OK 30 / 30 750 / 750 750 written
claude-opus-4-7 OK 30 / 30 600 / 600 600 written

2026-W23

Data for 2026-W23.

Models due in 2026-W23.
Model Status Prompts Samples Run log Why
llama3.2:3b OK 30 / 30 750 / 750 750 written
gpt-5.1 OK 30 / 30 750 / 750 750 written

2026-W22

Data for 2026-W22.

Models due in 2026-W22.
Model Status Prompts Samples Run log Why
llama3.2:3b OK 30 / 30 750 / 750 750 written
claude-opus-4-7 OK 30 / 30 600 / 600 600 written

2026-W21

Data for 2026-W21.

Models due in 2026-W21.
Model Status Prompts Samples Run log Why
llama3.2:3b OK 30 / 30 750 / 750 750 written
gpt-5.1 OK 30 / 30 750 / 750 750 written

2026-W20

Data for 2026-W20.

Models due in 2026-W20.
Model Status Prompts Samples Run log Why
llama3.2:3b OK 30 / 30 750 / 750 750 written
claude-opus-4-7 OK 30 / 30 600 / 600 600 written

2026-W19

Data for 2026-W19.

All runners The published W19 snapshot (gpt-5.1 750 samples and llama3.2:3b 750 samples, captured 2026-05-06) has no matching run log entry. The two logged W19 runs are a 2026-05-04 workstation llama3.2:3b run whose samples were not published and a 2026-05-11 resume that skipped all 60 pairs. Why the capturing run left no log entry is not established. Evidence: snapshot captured_at 2026-05-06; commit 8f29d20; commit 6f3221f.

Models due in 2026-W19.
Model Status Prompts Samples Run log Why
llama3.2:3b OK 30 / 30 750 / 750 750 written
gpt-5.1 OK 30 / 30 750 / 750 0 written

2026-W18

Data for 2026-W18.

Models due in 2026-W18.
Model Status Prompts Samples Run log Why
llama3.2:3b Lost 0 / 30 0 / 750 750 written

No llama3.2:3b samples in the published W18 record. A workstation run labelled W18 wrote 750 samples on 2026-04-25, which falls in ISO week 2026-W17, so they were not captured in W18 and were not published.

Evidence: run_log 2026-04-25T03:43:33+00:00; commit f6489cd.

claude-opus-4-7 OK 30 / 30 600 / 600 1200 written

Two runs carry the W18 label. The first, on 2026-04-24 (600 samples, captured in ISO week 2026-W17), was not published. The published W18 data is the second run, captured 2026-05-03 (600 samples).

Evidence: run_log 2026-04-24T22:39:32+00:00; run_log 2026-05-03T01:40:54+00:00; commit f6489cd.

2026-W17

Data for 2026-W17.

Models due in 2026-W17.
Model Status Prompts Samples Run log Why
llama3.2:3b Lost 0 / 30 0 / 750 0 written

No llama3.2:3b samples. The W17 run ran on a hosted CI runner with no ollama installed, so only openai/gpt-5.1 was sampled (750 samples). The failure left no error in the run log and was caught on review the following week.

Evidence: run_log 2026-04-24T05:49:07+00:00; commit 488ac4e.

gpt-5.1 OK 30 / 30 750 / 750 750 written

Files

Everything published under /data/coverage/.
FileSizeContents
coverage.csv 7426 B One row per scheduled or published (week, model): status, prompts and samples against expected, run log count, reasons and evidence.
coverage.jsonl 24303 B The same rows, one JSON object per line.
gaps.jsonl 8444 B The gap ledger exactly as committed: one hand-recorded explanation per line, append-only.
SHA256SUMS – SHA-256 digest for every file above.

Each gap ledger line carries week_id, scope (all, a provider/model, or stance for the stance classifier), kind (lost, partial, degraded, note, or corrected for a later correction of an earlier record with the same week and scope), reason, evidence and recorded_at. Lines are added, never edited; a later line can extend or correct an earlier one.