feat: add benchmark single-cell runner (task 0027)
Thread every benchmark layer to run one (arm, task, trial) cell end to end: provision and seed a throwaway repository, run the agent under the active arm bounded by a turn cap and a wall-clock backstop, audit the transcript, capture and score the post-run state, append the sample, and delete the repository. - runner.ts: runCell orchestration behind the BenchHost and AgentDriver seams, so the flow is unit-tested with fakes while the live wiring is validated by a smoke run; turn-cap and wall-clock failures are tagged confused-versus-hung, and a leaked transcript is flagged invalid. - audit.ts: the post-run transcript audit plus the shared foreignToolReason predicate both isolation enforcement points consume. - task.ts: the runnable BenchTask wrapper and one sample single-mutation task exercising the full path. - snapshot.ts: captureRepoState, the seed's read-back counterpart, in the RepoState shape the checker diffs against. - host.ts / sdk-driver.ts: the live BenchHost and the Claude Agent SDK driver (an optional peer, loaded via dynamic import) for real runs. - runner.smoke.test.ts: the live tracer-bullet tier, skipping cleanly when no host or SDK is configured.
This commit was merged in pull request #28.
This commit is contained in:
@@ -73,8 +73,12 @@ async function requireOk(res: Response, method: string, path: string): Promise<R
|
||||
return res;
|
||||
}
|
||||
|
||||
/** Issue a request and require a 2xx, returning the parsed JSON body. */
|
||||
async function send<T>(
|
||||
/**
|
||||
* Issue a request and require a 2xx, returning the parsed JSON body. Exported so
|
||||
* the post-run snapshot capture (snapshot.ts) reads the live repository through
|
||||
* the same authenticated round-trip the seed writes through.
|
||||
*/
|
||||
export async function send<T>(
|
||||
access: BenchAccess,
|
||||
method: string,
|
||||
path: string,
|
||||
|
||||
Reference in New Issue
Block a user