The official all-False baseline is reproduced at 67.86% (19/28) with full coverage, but the reported 92.86% intervention is withdrawn: it is an in-sample model-name rule chosen after reading all 28 labels, the CV loops never fit inside training folds, and the scrambled-text surrogate keeps the same 92.86%.
hunts/r_662b12 (Record 52 of 98 in chronological sequence)
Guiding Question
AIMO Interpretability 2026: reproduce the official baseline and test one structure-matched robustness signal
Method & Verification
The public val-sample of 28 rows was inventoried and the all-False constant was scored through the official ingestion shape; an independent audit then showed the frontier rule hard-codes gpt-5.2 and glm-5.1 after seeing those labels, so LOOCV, LOPO, and 5-fold only rescore a frozen in-sample rule.
Lineage & Relationships
Primary Sources (at pin 8fa46e134)
Editorial Notes
STALE HANDBACK: HANDBACK.json still claims a +25 pp GO. RESULTS.md (correction 2026-08-22) and AUDIT.md reject the intervention. Only the 67.86% baseline and the 19/9 sample census stand. Prize disposition is NO-GO on this evidence. Disposition is partly settled, not completed, because the baseline is retained and the intervention is withdrawn.
Date Provenance
commit 649b68ff4836a3adf87b1af7b7a36b407f7a7f1a, hunts/r_662b12/RESULTS.md, author 2026-08-21T23:30:47-05:00