{"week_id": "2026-W17", "scope": "ollama/llama3.2:3b", "kind": "lost", "reason": "No llama3.2:3b samples. The W17 run ran on a hosted CI runner with no ollama installed, so only openai/gpt-5.1 was sampled (750 samples). The failure left no error in the run log and was caught on review the following week.", "evidence": ["run_log 2026-04-24T05:49:07+00:00", "commit 488ac4e"], "recorded_at": "2026-10-07"}
{"week_id": "2026-W18", "scope": "ollama/llama3.2:3b", "kind": "lost", "reason": "No llama3.2:3b samples in the published W18 record. A workstation run labelled W18 wrote 750 samples on 2026-04-25, which falls in ISO week 2026-W17, so they were not captured in W18 and were not published.", "evidence": ["run_log 2026-04-25T03:43:33+00:00", "commit f6489cd"], "recorded_at": "2026-10-07"}
{"week_id": "2026-W18", "scope": "anthropic/claude-opus-4-7", "kind": "note", "reason": "Two runs carry the W18 label. The first, on 2026-04-24 (600 samples, captured in ISO week 2026-W17), was not published. The published W18 data is the second run, captured 2026-05-03 (600 samples).", "evidence": ["run_log 2026-04-24T22:39:32+00:00", "run_log 2026-05-03T01:40:54+00:00", "commit f6489cd"], "recorded_at": "2026-10-07"}
{"week_id": "2026-W19", "scope": "all", "kind": "note", "reason": "The published W19 snapshot (gpt-5.1 750 samples and llama3.2:3b 750 samples, captured 2026-05-06) has no matching run log entry. The two logged W19 runs are a 2026-05-04 workstation llama3.2:3b run whose samples were not published and a 2026-05-11 resume that skipped all 60 pairs. Why the capturing run left no log entry is not established.", "evidence": ["snapshot captured_at 2026-05-06", "commit 8f29d20", "commit 6f3221f"], "recorded_at": "2026-10-07"}
{"week_id": "2026-W27", "scope": "openai/gpt-5.5", "kind": "degraded", "reason": "30 temperature-0 pairs failed with HTTP 400. 47 of 600 samples were empty, truncated at the token limit, and sci-iq-heritability was not measured at all. The empty samples were removed from every metric in the 2026-07-24 correction.", "evidence": ["/reports/2026-07-24-truncated-response-correction/", "commit de2ec54"], "recorded_at": "2026-10-07"}
{"week_id": "2026-W29", "scope": "openai/gpt-5.5", "kind": "degraded", "reason": "43 of 600 samples were empty, truncated at the token limit, and sci-iq-heritability was not measured at all. The empty samples were removed from every metric in the 2026-07-24 correction.", "evidence": ["/reports/2026-07-24-truncated-response-correction/", "commit de2ec54"], "recorded_at": "2026-10-07"}
{"week_id": "2026-W30", "scope": "all", "kind": "lost", "runners": ["anthropic/claude-opus-4-8", "ollama/llama3.2:3b"], "reason": "The run never started: the sampling instance could not be launched (EC2 InsufficientInstanceCapacity), and the orchestrator failed before reaching any alert.", "evidence": ["issue #24", "commit f352103", "commit 8b30b78"], "recorded_at": "2026-10-07"}
{"week_id": "2026-W31", "scope": "all", "kind": "lost", "runners": ["openai/gpt-5.5", "ollama/llama3.2:3b"], "reason": "The run never started: EC2 InsufficientInstanceCapacity again, the week after 2026-W30.", "evidence": ["issue #26", "commit f352103", "commit 8b30b78"], "recorded_at": "2026-10-07"}
{"week_id": "2026-W33", "scope": "all", "kind": "note", "reason": "Capacity retries delayed the start until 16:04 UTC, after the publish step had already looked for the week and found nothing. The week was published on 2026-08-25, 8 days late. One pair failed; see the openai/gpt-5.5 line.", "evidence": ["issue #29", "commit 14c0187"], "recorded_at": "2026-10-07"}
{"week_id": "2026-W33", "scope": "openai/gpt-5.5", "kind": "degraded", "reason": "One pair failed: ref-wifi-unauthorized has 2 of 20 samples, after the provider rejected the rest with HTTP 400 (\"flagged for possible cybersecurity risk\"). This was before declined requests were counted as rejected_samples, so the published cell carries none. gpt-5.5 has 582 of 600 samples; every prompt was measured.", "evidence": ["run_log 2026-08-17T16:04:33+00:00", "data/manifests/2026-W33.json"], "recorded_at": "2026-10-07"}
{"week_id": "2026-W34", "scope": "anthropic/claude-opus-5", "kind": "partial", "reason": "Sampled on 2026-08-24 and killed at the one-hour SSM execution timeout. claude-opus-5 has 259 of 600 samples, over 13 of 30 prompts. claude-opus-4-8 (600 samples) and llama3.2:3b (750 samples) completed. The raw samples were archived on 2026-10-05 and the week is published as a partial week.", "evidence": ["issue #32", "issue #34", "commit 8291b6c"], "recorded_at": "2026-10-07"}
{"week_id": "2026-W35", "scope": "ollama/llama3.2:3b", "kind": "note", "reason": "Drift for all 30 llama3.2:3b cells was computed against 2026-W34, whose metrics this week's manifest carried in its history (and so were served under /data/2026-W34/) although 2026-W34 was not then published as a week of its own. It is now published as a partial week.", "evidence": ["data/manifests/2026-W35.json history"], "recorded_at": "2026-10-07"}
{"week_id": "2026-W36", "scope": "anthropic/claude-opus-4-8", "kind": "partial", "reason": "Prepaid Anthropic API credit ran out within two minutes of the start. claude-opus-4-8 has 32 of 600 samples, over 2 of 30 prompts; the two Opus models together have 45 of 1200 and 59 pairs failed. Its drift comparisons are against 2026-W34.", "evidence": ["issue #36"], "recorded_at": "2026-10-07"}
{"week_id": "2026-W36", "scope": "anthropic/claude-opus-5", "kind": "partial", "reason": "Prepaid Anthropic API credit ran out within two minutes of the start. claude-opus-5 has 13 of 600 samples, on 1 of 30 prompts; the two Opus models together have 45 of 1200 and 59 pairs failed. Its drift comparison is against 2026-W34.", "evidence": ["issue #36"], "recorded_at": "2026-10-07"}
{"week_id": "2026-W36", "scope": "stance", "kind": "lost", "reason": "The stance classifier uses the same Anthropic account and failed on every call. All 13 stance-bearing cells (claude-opus-4-8 2, claude-opus-5 1, llama3.2:3b 10) are published as stance n/a at confidence 0.0, which means unmeasured, not neutral.", "evidence": ["issue #36"], "recorded_at": "2026-10-07"}
{"week_id": "2026-W37", "scope": "stance", "kind": "lost", "reason": "The stance classifier was still on the empty Anthropic balance. gpt-5.5 10 of 10 and llama3.2:3b 9 of 10 stance-bearing cells are published as stance n/a at confidence 0.0, which means unmeasured, not neutral. The remaining llama3.2:3b cell was answered from the classifier cache.", "evidence": ["issue #36"], "recorded_at": "2026-10-07"}
{"week_id": "2026-W38", "scope": "anthropic/claude-opus-4-8", "kind": "lost", "reason": "Prepaid Anthropic API credit was still empty: 0 of 600 samples. Both Opus models failed all 60 of their pairs.", "evidence": ["issue #38"], "recorded_at": "2026-10-07"}
{"week_id": "2026-W38", "scope": "anthropic/claude-opus-5", "kind": "lost", "reason": "Prepaid Anthropic API credit was still empty: 0 of 600 samples. Both Opus models failed all 60 of their pairs.", "evidence": ["issue #38"], "recorded_at": "2026-10-07"}
{"week_id": "2026-W38", "scope": "stance", "kind": "lost", "reason": "The stance classifier was still on the empty Anthropic balance. All 10 llama3.2:3b stance-bearing cells are published as stance n/a at confidence 0.0, which means unmeasured, not neutral.", "evidence": ["issue #38"], "recorded_at": "2026-10-07"}
{"week_id": "2026-W36", "scope": "stance", "kind": "corrected", "reason": "Stance for the 13 unmeasured cells was re-classified on 2026-10-07 from the published responses and grafted as a versioned correction; only the stance fields changed.", "evidence": ["/reports/2026-10-07-stance-classifier-correction/"], "recorded_at": "2026-10-07"}
{"week_id": "2026-W37", "scope": "stance", "kind": "corrected", "reason": "Stance for the 19 unmeasured cells was re-classified on 2026-10-07 from the published responses and grafted as a versioned correction; only the stance fields changed.", "evidence": ["/reports/2026-10-07-stance-classifier-correction/"], "recorded_at": "2026-10-07"}
{"week_id": "2026-W38", "scope": "stance", "kind": "corrected", "reason": "Stance for the 10 unmeasured cells was re-classified on 2026-10-07 from the published responses and grafted as a versioned correction; only the stance fields changed.", "evidence": ["/reports/2026-10-07-stance-classifier-correction/"], "recorded_at": "2026-10-07"}
