Turn a benchmark run into a three-arm experiment: alongside the new-skill and no-skill arms, add a previous-version arm materialized from the default branch's HEAD, drawn automatically whenever the skill's directory differs from HEAD and degrading to the two-arm efficacy-only shape otherwise. Two blind head-to-heads now fall out per trial — efficacy (new-vs-no-skill) and regression (new-vs-old). The deterministic core dispatches a pass rule per comparison (efficacy at wins>=3 and losses<=1, regression at losses<=1 with no wins floor), gates only Efficacy on the hard assertion, and yields a second skill-level Regression verdict, green when no case regressed and not-applicable on a two-arm run. The report gains a second badge, a headline reading both verdicts together, a previous-version cost row with a two-ratio footnote, and per-case side-by-side comparisons that auto-expand on any failure or flagged loss to show the losing-trial output pair. The fixture test covers both shapes in one test: a three-arm run asserting per-arm metrics, both per-case and skill-level verdicts, the net margins, the three-row cost table and losing-trial evidence, and the two-arm run retained as the degenerate no-previous-version case.
13 KiB
13 KiB