Experiment
Background check
Experiment 2 — background check: does the filler model reliably produce valid moves? Validates glm-5.2 (the new background) before anything expensive runs. Expectation: glm-5.2 forfeits under 3% of turns; in the all-glm baseline, seat labels (P1..P5) score within noise of each other — any seat-label spread is a design artifact.
- status
- done
- coverage
- 11 / 11 episodes
- conditions
- pure economy; messages on, pure economy; messages on
- config
- configs/02_background_check.yaml
11 episodes; mean confirmed message chains per episode = 0.5.
one slim bar per episode; the oxblood line is the mean
Episodes
| episode | condition | focal model | capture (by focal) | cascades | gini |
|---|---|---|---|---|---|
| 02_background_check--gemini-flash-r0 | complete/pure/msg-on | — | — | 0 | 0.132 |
| 02_background_check--gemini-flash-r1 | complete/pure/msg-on | — | — | 0 | 0.179 |
| 02_background_check--glm-5.2-r0 | complete/pure/msg-on | — | — | 0 | 0.0357 |
| 02_background_check--glm-5.2-r1 | complete/pure/msg-on | — | — | 0 | 0.0326 |
| 02_background_check--gpt-5.4-mini-r0 | complete/pure/msg-on | — | — | 0 | 0.283 |
| 02_background_check--gpt-5.4-mini-r1 | complete/pure/msg-on | — | — | 0 | 0.145 |
| 02_background_check--lineup_complete_r0 | complete/pure/msg-on | — | — | 0 | 0.263 |
| 02_background_check--lineup_complete_r1 | complete/pure/msg-on | — | — | 3 | 0.0235 |
| 02_background_check--lineup_complete_r2 | complete/pure/msg-on | — | — | 0 | 0.2 |
| 02_background_check--lineup_complete_r3 | complete/pure/msg-on | — | — | 0 | 0.226 |
| 02_background_check--lineup_complete_r4 | complete/pure/msg-on | — | — | 3 | 0.0239 |
11 episodes, sorted by condition then id; episode links open the transcript reader.