feat: add efficacy-benchmark tracer bullet (task 0004)
Introduce the thinnest complete path that benchmarks a skill's efficacy against a no-skill baseline and renders an HTML report. - tests/ convention: a top-level tree mirroring skills/ by full path, with case directories holding a case.md (CASE-FORMAT.md) and an optional fixture. - benchmark-skill: a distributed, non-model-invocable runner (/benchmark-skill <name>) that orchestrates force-invoked new-skill and skill-absent baseline arms plus a blind per-trial judge as in-session subagents. - Deterministic core (Python): a pure transform over the collected run data that sums per-arm usage, collapses each trial to WIN/TIE/LOSS, applies the efficacy pass rule (wins >= 3, losses <= 1), and renders a self-contained report. Parameterized over arms and comparisons; two-arm here. - Fixture-driven no-LLM check (checks/benchmark-core.nix) wired into flake.nix. - Real tests tree for axi-review; a live run produced its efficacy report.
This commit is contained in:
@@ -84,6 +84,13 @@
|
||||
shell-hook = import ./checks/shell-hook.nix {
|
||||
inherit pkgs mkSkill mkSkillsShellHook;
|
||||
};
|
||||
|
||||
# Proves the benchmark skill's deterministic core at its single seam:
|
||||
# it feeds committed fixtures to the core and asserts the metrics, the
|
||||
# Efficacy verdict, and the rendered report.
|
||||
benchmark-core = import ./checks/benchmark-core.nix {
|
||||
inherit pkgs;
|
||||
};
|
||||
};
|
||||
}
|
||||
);
|
||||
|
||||
Reference in New Issue
Block a user