Basically I didnโt match Muon and AdamW up together. In the morning I can get the actual details. At just 3B tokens, Boris-2-0917, as Iโve taken to calling it, has already beaten the 30B checkpoint on all benchmarks.
bench-labs/cagliostro-v3 just hit an Intelligence Index of 26.13 on the AxiomicLabs/Open_SLM_Leaderboard a 146M-param model trained completely from scratch on a single consumer GPU. That's 2nd place overall, and as far as I can tell, the most capable SLM trained on consumer hardware to date. Beating SmolLM-135m on 1/8th of the data is just silly levels of efficiency.