97c0eeb0550d8609cef2358f2430280a2c8e6c4a
Add a second pure transform to the benchmark core: a `--leaderboard` mode that ranks per-skill results models into an index leaderboard, one row per skill carrying both verdicts, links out to each per-skill report, and sorts any red first (regressions above efficacy failures), then fragile-but-passing, then clean green. The fragile tier reuses the 0006 per-badge chips: the efficacy chip now carries its win count so the leaderboard reads it against the pass floor without recomputing margins, and fragility is scoped to the passing tier so a red row never carries a chip. The runner's SKILL.md gains the no-argument batch flow and the index invocation. The fixture test covers the tiered sort, the not-applicable Regression cell, and the per-skill links.
skills
My personal skills packaged through Nix.
Languages
Nix
97.7%
Shell
2.3%