diff --git a/.claude/adr/0008-errors-as-values-at-cli-boundary.md b/.claude/adr/0008-errors-as-values-at-cli-boundary.md index c66a84f..7512c46 100644 --- a/.claude/adr/0008-errors-as-values-at-cli-boundary.md +++ b/.claude/adr/0008-errors-as-values-at-cli-boundary.md @@ -21,6 +21,7 @@ It keeps the bin layer consistent with core's total-function stance, so the whol The possibility of failure becomes visible in a function's signature instead of hidden behind a `throw`. The success value cannot be read without first handling the error case, which removes a class of mistakes at compile time. A tagged-union error type gives exhaustive handling: a new failure mode is a new tag that every match must account for. +Because failure is a returned value rather than a side effect, error propagation can be asserted in-process by the integration tier, without spawning the binary (see ADR 0009). ## Consequences diff --git a/.claude/adr/0009-testing-tiers-and-boundaries.md b/.claude/adr/0009-testing-tiers-and-boundaries.md new file mode 100644 index 0000000..800de30 --- /dev/null +++ b/.claude/adr/0009-testing-tiers-and-boundaries.md @@ -0,0 +1,44 @@ +# ADR 0009 — Testing Tiers and Boundaries + +**Status**: Accepted + +## Context + +The test suite accumulated overlapping files — unit, integration, smoke, and subprocess — with no crisp definition of what each was responsible for. +The result was duplication and confusion: the same behaviour asserted in more than one tier, and no rule for where a given test belonged. +Both specs already referred to unit, integration, and smoke tests, but none of them pinned the boundaries. + +## Decision + +Recognise three test tiers, each defined by the seam it exercises. + +**Unit** — one component in isolation, a pure function, asserted by its return value. +`parse` in core and `render` in bin are the unit seams. + +**Integration** — several components composed in-process, asserted by the value the composed function returns. +The `view` command function — file read, then `parse`, then `render`, returning a `Result` — is the integration seam. + +**E2E** — the binary as a black box. +It is spawned as a subprocess and asserted on its exit code and its stdout and stderr, with no knowledge of the internal structure. + +The line between what is unit- or integration-testable and what is e2e-only is whether a function returns a value or performs a process-level side effect. +A function that returns a value can be asserted in-process. +A function that calls `process.exit` or `process.stdout.write` can only be observed by spawning the binary. +So all logic is pushed into value-returning functions, and the entry shell is kept as thin as possible, because it is the one part reachable only through a subprocess. + +Core exposes a single public seam, `parse`, so it has unit tests only. +A whole-fixture parse test is still a unit test on that same seam with a broad input — a corpus test — not a separate tier. + +## Rationale + +Precise, non-overlapping definitions prevent the duplication that a vague unit-integration-smoke split produced. +Classifying by seam matches what is actually cheap or expensive to test. +Value-returning code runs fast and is visible to coverage in-process, while process-side-effecting code needs a subprocess and is invisible to coverage. +Concentrating the process-boundary surface in one thin shell keeps the amount of e2e-only code to a minimum. + +## Consequences + +Each behaviour is tested at exactly one tier: logic at unit or integration in-process, the process boundary at e2e. +The e2e tier is deliberately minimal — it verifies wiring such as exit codes and stream routing, not content already proven in-process. +Tests are co-located as `{module}_test.ts`, so the command-function tests live beside the entry module and are integration tests despite the file name. +When a command function shares a file with the top-level `program.parse()`, that call is guarded with `import.meta.main`, so importing the module for a test does not run the CLI.