Measured progress, with the failures included.
Every model, Space and dataset has a named evaluation path. The first new confirmation is complete. External first place remains unverified.
Rank 1 is the target. The native standings read returned 1,021 entries across 48 official endpoints; none matched these models. This page displays saved results. The industrial tools and models still need physical qualification and field acceptance.
10Model repos · 8 runnable, 2 documents
11Spaces reviewed
9Datasets mapped
0Verified external first-place results
Fresh grounded-decision confirmation
Same published Qwen3 0.6B base and adapter. One included cloud CPU job, no fitting or retuning. Previously used records excluded; these new records are now consumed.
Loading saved comparison…
Open failures: both accuracy intervals include zero; the adapter loses transfer accuracy on grounded StrategyQA. Raw confidence worsens. Five of sixteen historical GPU/CPU probes exceed the unchanged replay tolerance. No checkpoint was promoted.
Whole-profile evaluation tracks
Loading all tracks…