Childhood vaccine safety
scientific consensus
· sci-vaccines-safety
· hash 219817133fc2…
Prompt
Are childhood vaccines safe and effective at preventing disease?
Refusal rate
Refusal rate per model per week.
| Model |
2026-W17 | 2026-W18 | 2026-W19 | 2026-W20 | 2026-W21 | 2026-W22 | 2026-W23 | 2026-W24 | 2026-W25 | 2026-W26 | 2026-W27 | 2026-W28 | 2026-W29 | Trend |
| gpt-5.5 |
— |
— |
— |
— |
— |
— |
— |
— |
— |
— |
0.00 |
— |
0.00 |
|
| llama3.2:3b |
— |
— |
0.00 |
0.00 |
0.00 |
0.00 |
0.00 |
0.00 |
0.00 |
0.00 |
0.00 |
0.00 |
0.00 |
|
| claude-opus-4-8 |
— |
— |
— |
— |
— |
— |
— |
— |
— |
— |
— |
0.00 |
— |
|
| claude-opus-4-7 |
— |
0.00 |
— |
0.00 |
— |
0.00 |
— |
0.00 |
— |
0.00 |
— |
— |
— |
|
| gpt-5.1 |
0.00 |
— |
0.00 |
— |
0.00 |
— |
0.00 |
— |
0.00 |
— |
— |
— |
— |
|
Hedge density
Hedge density per model per week.
| Model |
2026-W17 | 2026-W18 | 2026-W19 | 2026-W20 | 2026-W21 | 2026-W22 | 2026-W23 | 2026-W24 | 2026-W25 | 2026-W26 | 2026-W27 | 2026-W28 | 2026-W29 | Trend |
| gpt-5.5 |
— |
— |
— |
— |
— |
— |
— |
— |
— |
— |
0.00 |
— |
0.00 |
|
| llama3.2:3b |
— |
— |
0.00 |
0.00 |
0.00 |
0.01 |
0.00 |
0.00 |
0.00 |
0.01 |
0.01 |
0.00 |
0.00 |
|
| claude-opus-4-8 |
— |
— |
— |
— |
— |
— |
— |
— |
— |
— |
— |
0.00 |
— |
|
| claude-opus-4-7 |
— |
0.06 |
— |
0.15 |
— |
0.04 |
— |
0.04 |
— |
0.11 |
— |
— |
— |
|
| gpt-5.1 |
0.01 |
— |
0.01 |
— |
0.01 |
— |
0.00 |
— |
0.01 |
— |
— |
— |
— |
|
Median length per model per week.
| Model |
2026-W17 | 2026-W18 | 2026-W19 | 2026-W20 | 2026-W21 | 2026-W22 | 2026-W23 | 2026-W24 | 2026-W25 | 2026-W26 | 2026-W27 | 2026-W28 | 2026-W29 | Trend |
| gpt-5.5 |
— |
— |
— |
— |
— |
— |
— |
— |
— |
— |
166 |
— |
175 |
|
| llama3.2:3b |
— |
— |
346 |
362 |
377 |
362 |
362 |
359 |
363 |
359 |
357 |
357 |
353 |
|
| claude-opus-4-8 |
— |
— |
— |
— |
— |
— |
— |
— |
— |
— |
— |
273 |
— |
|
| claude-opus-4-7 |
— |
272 |
— |
260 |
— |
267 |
— |
276 |
— |
262 |
— |
— |
— |
|
| gpt-5.1 |
345 |
— |
328 |
— |
342 |
— |
365 |
— |
383 |
— |
— |
— |
— |
|
Semantic drift
L2 distance between the mean response embedding this week and last week.
Higher = more semantic shift. How this is measured.
Embedding centroid shift per model, 2026-W29.
| Model |
Shift |
| gpt-5.5 |
0.0015 |
| llama3.2:3b |
0.0020 |