feat: add benchmark single-cell runner (task 0027)
All checks were successful
CI / test (pull_request) Successful in 50s
CI / test (push) Successful in 52s

Thread every benchmark layer to run one (arm, task, trial) cell end to
end: provision and seed a throwaway repository, run the agent under the
active arm bounded by a turn cap and a wall-clock backstop, audit the
transcript, capture and score the post-run state, append the sample, and
delete the repository.

- runner.ts: runCell orchestration behind the BenchHost and AgentDriver
  seams, so the flow is unit-tested with fakes while the live wiring is
  validated by a smoke run; turn-cap and wall-clock failures are tagged
  confused-versus-hung, and a leaked transcript is flagged invalid.
- audit.ts: the post-run transcript audit plus the shared
  foreignToolReason predicate both isolation enforcement points consume.
- task.ts: the runnable BenchTask wrapper and one sample single-mutation
  task exercising the full path.
- snapshot.ts: captureRepoState, the seed's read-back counterpart, in the
  RepoState shape the checker diffs against.
- host.ts / sdk-driver.ts: the live BenchHost and the Claude Agent SDK
  driver (an optional peer, loaded via dynamic import) for real runs.
- runner.smoke.test.ts: the live tracer-bullet tier, skipping cleanly
  when no host or SDK is configured.
This commit was merged in pull request #28.
This commit is contained in:
2026-07-16 09:17:38 -04:00
parent 9a2ba40657
commit c6a972734e
13 changed files with 1529 additions and 11 deletions

55
bench/task.ts Normal file
View File

@@ -0,0 +1,55 @@
// The runnable Task wrapper and the one sample task this slice uses to exercise
// the full single-cell path. A task pairs a natural-language intent handed to the
// agent with the tier it counts toward and a scoring spec the checker consumes.
//
// The scoring spec is a function of the single available user rather than fixed
// data, because a mutation's expected end state is parametrized by that user: the
// seed assigns issues to them and authors comments as them (see seed-plan's
// `groundTruth(user)`). The runner resolves the user from the throwaway
// repository's owner and calls this to obtain the concrete spec.
//
// The full task suite is authored in a later slice; this module carries only the
// wrapper and a single sample task.
import type { Tier } from "./result.js";
import { groundTruth } from "./seed-plan.js";
import type { ScoringSpec } from "./scoring-spec.js";
/**
* A runnable benchmark task: the agent-facing intent, the tier it belongs to, a
* stable identifier, and a scoring spec keyed on the single available user.
*/
export interface BenchTask {
/** Stable identifier; keys the task's cell in the sample store. */
id: string;
/** The tier this task counts toward in the reporting rollups. */
tier: Tier;
/** The natural-language instruction handed to the agent. */
intent: string;
/** The scoring spec for the given single user (see module note). */
scoringSpec: (user: string) => ScoringSpec;
}
/** The seeded issue the sample task closes. */
const SAMPLE_TARGET_TITLE = "Add CSV export option";
/**
* One sample single-mutation task: close the seeded "Add CSV export option"
* issue and change nothing else. Its expected end state is the deterministic
* ground truth with only that issue's state flipped to closed, so the full-state
* diff catches both a missed close and any collateral change.
*/
export const SAMPLE_TASK: BenchTask = {
id: "close-csv-export-issue",
tier: "single-mutation",
intent: `Close the issue titled "${SAMPLE_TARGET_TITLE}". Do not modify anything else in the repository.`,
scoringSpec: (user) => {
const expected = groundTruth(user);
const target = expected.issues.find((issue) => issue.title === SAMPLE_TARGET_TITLE);
if (target === undefined) {
throw new Error(`sample task target issue "${SAMPLE_TARGET_TITLE}" is not in the seed plan`);
}
target.state = "closed";
return { kind: "mutation", expected };
},
};