The AIMO clear-cut protocol is reconstructed from public data as minority-robust, a legal model-identity prior beats the constant leave-one-problem-out (26/28 vs 19/28), the Small-track entry is the always-non-robust constant, and the learned-method GPU arm stays killed.
hunts/aimo2 (Record 55 of 98 in chronological sequence)
Guiding Question
Can we (a) reproduce, from public data alone, a validation-design finding about the AIMO sample worth a technical-report submission, and (b) build a legal Small/Main-track method that beats the better constant baseline out-of-fold under a preregistered leakage-controlled protocol?
Method & Verification
val-sample's 9 robust rows were matched to sample-full pairs with base accuracy 1.0 and zero detrimental perturbations, a per-model prior was scored leave-one-problem-out on public sets, and the official container path was replayed at starter e46be92.
Lineage & Relationships
Primary Sources (at pin 8fa46e134)
Editorial Notes
Successor to r_662b12; the withdrawn +25 pp claim stays withdrawn and is re-measured here as a model-identity base rate. Status is report-ready with the learned arm killed, so disposition is qualified rather than completed. Grade is measured on public data. No GPU spent. Hidden-set model identifiers and the hidden robust fraction remain unsettled.
Date Provenance
commit 9d9e697909058c5ab9c2078b2465cdd8ad3b9d50, hunts/aimo2/RESULTS.md, author 2026-08-22T21:01:20-05:00