docs: add benchmark harness spec, ADRs, and task breakdown
Add the benchmark-harness spec, three supporting ADRs (cost-equivalent token metric, single-user seed, guard-based tool isolation), and the 0022-0030 task breakdown that slices the harness into foundational seams, an integrating single-cell runner, and the reporting layer.
This commit was merged in pull request #22.
This commit is contained in:
22
.claude/adr/0015-bench-single-user-seed.md
Normal file
22
.claude/adr/0015-bench-single-user-seed.md
Normal file
@@ -0,0 +1,22 @@
|
||||
# Single-user seed and its constraints on the task surface
|
||||
|
||||
The benchmark seeds a throwaway repository to a known state before each trial and scores tasks against that ground truth.
|
||||
Seeding realistic author and assignee variety would require several Gitea accounts.
|
||||
The maintainer prefers to run the benchmark under their single existing account rather than provision additional accounts.
|
||||
|
||||
## Considered Options
|
||||
|
||||
**Provision throwaway collaborator accounts** (rejected) — Multiple accounts would restore the author and assignee dimensions and enable non-self pull-request approvals, but they add account lifecycle and credential handling that the maintainer explicitly declined for this benchmark.
|
||||
|
||||
**Keep multi-user tasks and let arms fail** (rejected) — Retaining tasks that need distinct authors or a non-self reviewer would make those tasks impossible under one account for every arm uniformly, producing no comparative signal while consuming budget.
|
||||
|
||||
**Single-user seed with a redesigned surface** (chosen) — All seed content is authored by the one account, and the discriminating dimensions become label, state, assignee presence (assigned-to-self versus unassigned), and title keyword instead of author.
|
||||
Tasks that assumed author or assignee variety are recast onto these axes: reassignment becomes assign-to-self or unassign, and author-filtered bulk mutation becomes a filter on assignee presence.
|
||||
|
||||
## Consequences
|
||||
|
||||
- Author filtering leaves the scored suite; it carries little signal with one author anyway.
|
||||
- Review tasks in the scored suite use comment-type reviews, which a user may leave on their own pull request.
|
||||
Whether the host permits self-approval and self-request-changes is verified during implementation; if permitted, those two tasks are promoted from comment reviews to approve and request-changes, otherwise approve and request-changes move to the bonus table as two-account scenarios.
|
||||
- Non-self approval, distinct-author filtering, and distinct-assignee tasks are out of scope for the scored suite and belong to the multi-account bonus scenarios.
|
||||
- The seed stays small and fully deterministic, which keeps the full-state diff used for collateral-damage checking cheap to compute and reason about.
|
||||
Reference in New Issue
Block a user