refactor: tighten benchmark-skill instructions per craft-skill audit

Collapse duplication so each rule has one home: the pass-rule threshold
and cost footnote defer to the core, and the case-contract detail defers
to CASE-FORMAT.md. Drop editor-facing sediment from the arms step.

Fixes surfaced by a verify walk-through: define the hard-assertion gate
as passing only when every predicate exits zero (an earlier failure was
maskable by a later success), state that in-session arms run at the
session default temperature, name the fixture directory, and note the
steps run from the repo root.
This commit was merged in pull request #4.
This commit is contained in:
2026-07-24 15:01:56 -04:00
parent 3e96db664c
commit 5db98dcab9
2 changed files with 15 additions and 19 deletions

View File

@@ -19,6 +19,7 @@ The authoring contract for a test case is [`CASE-FORMAT.md`](CASE-FORMAT.md).
Read it before you read the target skill's tests.
Report artifacts live under a git-ignored `tests/.reports/` directory at the repo root, flat and keyed by skill name.
Run from the repo root: every path in the steps below is relative to it.
## 1. Resolve the target and validate its tests tree
@@ -29,7 +30,7 @@ If the skill exists but has no tests directory or no case directories, report th
Validate the shape of the tests tree at run time:
- Each case directory has a `case.md` that parses per `CASE-FORMAT.md` — a `description`, a `## Prompt`, at least one `## Soft criteria` entry, and optional `## Seed` and `## Hard assertions` blocks.
- Each case directory has a `case.md` that parses per `CASE-FORMAT.md`.
- Any fixture a case references exists.
- The test directory maps to a real skill.
@@ -40,12 +41,10 @@ Done when every case parses, its fixture exists, and you have the case list in s
## 2. Fix the arms and per-run settings
This slice is **two-arm efficacy only**: a new-skill arm and a no-skill baseline arm.
(The previous-version regression arm is a later slice.
The core already handles extra arms, so do not remove the two-arm shape.)
Run **two arms**: a new-skill arm and a no-skill baseline arm.
Choose the trial temperature: a realistic temperature (around 1.0) by default, so the result reflects whether the skill *reliably* helps across variance.
A skill may pin a temperature-zero override when it wraps a genuinely mechanical task — honor that override if the skill declares one.
Arms run at the session's realistic default temperature (around 1.0), so the result reflects whether the skill *reliably* helps across variance rather than helping once by luck.
Per-arm temperature is not settable through the in-session workflow surface, so a skill's temperature-zero override for a genuinely mechanical task is a documented knob for the future headless path, not something to set here.
Each case runs **5 paired trials**.
Trial *i*'s new-skill output is judged against trial *i*'s no-skill output.
@@ -82,11 +81,8 @@ Have it return a single winner: A, B, or tie.
## 4. Run the hard-assertion gate
Run the case's `## Hard assertions` against the **new arm only**, once per trial.
Bind `$OUTPUT` to the path of that trial's new-arm final message and `$WORLD` to that trial's new-arm world (its fixture copy, which is the arm's working directory).
Run each predicate.
A non-zero exit fails the assertion.
Record pass/fail per trial — a failed hard assertion fails the case outright regardless of the head-to-head.
Run the case's `## Hard assertions` against the **new arm only**, once per trial, with `$OUTPUT` and `$WORLD` bound as `CASE-FORMAT.md` defines — that trial's new-arm final message and its world.
Record each trial's pass or fail into the bundle for the core to score.
A case with no `## Hard assertions` block simply has no gate.
## 5. Collect the run data
@@ -146,9 +142,9 @@ python3 skills/benchmark-skill/core/benchmark_core.py \
--html tests/.reports/<name>.html
```
The core collapses each efficacy trial to WIN, TIE, or LOSS for the new arm (a tie is a non-win), applies the pass rule — **efficacy passes for a case when wins ≥ 3 and losses ≤ 1** flags any loss for human review, and yields a skill-level **Efficacy verdict** that is green only when every case passes.
It renders a self-contained HTML report: the Efficacy badge and run metadata, a two-row per-arm cost table (footnoted that absolute cost is inflated by shared-context cache overhead, so the trustworthy signal is the new-vs-no-skill ratio), and the cases in stable authored order.
Cost is reported alongside quality but never gates the verdict.
The core collapses each trial to a per-arm WIN, TIE, or LOSS, applies the efficacy pass rule, flags any loss for human review, and yields the skill-level **Efficacy verdict**.
It renders the self-contained HTML report the verdict badge, run metadata, the per-arm cost table, and the cases in stable authored order — alongside the machine-readable model.
Cost is reported but never gates the verdict.
Clean up the scratch when done: remove the `tests/.reports/.work/` directory and every per-arm world and skill materialization you created under the system temp path.