feat(skills): adapt planning skills to artifact workflow

This commit is contained in:
2026-07-31 15:13:31 -04:00
parent 181dcc7a9e
commit 80cecf3350
6 changed files with 264 additions and 204 deletions

View File

@@ -1,7 +1,6 @@
# Logic Prototype
A tiny interactive terminal app that lets the user drive a state model by hand.
Use this when the question is about **business logic, state transitions, or data shape** — the kind of thing that looks reasonable on paper but only feels wrong once you push it through real cases.
A tiny interactive terminal app that lets the user drive a state model by hand. Use this when the question is about **business logic, state transitions, or data shape** — the kind of thing that looks reasonable on paper but only feels wrong once you push it through real cases.
## When this is the right shape
@@ -10,105 +9,89 @@ Use this when the question is about **business logic, state transitions, or data
- "I want to feel out what the API should look like before writing it."
- Anything where the user wants to **press buttons and watch state change**.
If the question is "what should this look like" — wrong branch.
Use [UI.md](UI.md).
If the question is "what should this look like" — wrong branch. Use [UI.md](UI.md).
## Process
### 1. State the question
Before writing code, write down what state model and what question you're prototyping.
One paragraph, in the prototype's README or a comment at the top of the file.
A logic prototype that answers the wrong question is pure waste make the question explicit so it can be checked later, whether the user is watching now or returning to it AFK.
Use one paragraph in the prototype's README or a comment at the top of the file.
A logic prototype that answers the wrong question is pure waste, so make the question explicit enough to check later whether the user is watching now or returning to it AFK.
Done when the prototype states one concrete logic question and the model being tested.
### 2. Pick the language
Use whatever the host project uses.
If the project has no obvious runtime (e.g. a docs repo), ask.
Use whatever the host project uses. If the project has no obvious runtime (e.g. a docs repo), ask.
Match the project's existing conventions for tooling — don't add a new package manager or runtime just for the prototype.
Match the project's existing conventions for tooling.
Don't add a new package manager or runtime just for the prototype.
Done when the prototype has a runnable host-project language and toolchain without introducing a new runtime convention.
### 3. Isolate the logic in a portable module
Put the actual logic — the bit that's answering the question — behind a small, pure interface that could be lifted out and dropped into the real codebase later.
The TUI around it is throwaway.
The logic module shouldn't be.
Put the actual logic — the bit that's answering the question — behind a small, pure interface that could be lifted out and dropped into the real codebase later. The TUI around it is throwaway; the logic module shouldn't be.
The right shape depends on the question:
- **A pure reducer** — `(state, action) => state`.
Good when actions are discrete events and state is a single value.
- **A state machine** — explicit states and transitions.
Good when "which actions are even legal right now" is part of the question.
- **A small set of pure functions** over a plain data type.
Good when there's no implicit current state — just transformations.
- **A pure reducer** — `(state, action) => state`. Good when actions are discrete events and state is a single value.
- **A state machine** — explicit states and transitions. Good when "which actions are even legal right now" is part of the question.
- **A small set of pure functions** over a plain data type. Good when there's no implicit current state — just transformations.
- **A class or module with a clear method surface** when the logic genuinely owns ongoing internal state.
Pick whichever shape best fits the question being asked, *not* whichever is easiest to wire to a TUI.
Keep it pure: no I/O, no terminal code, no `console.log` for control flow.
The TUI imports it and calls into it.
Nothing flows the other direction.
Pick whichever shape best fits the question being asked, *not* whichever is easiest to wire to a TUI. Keep it pure: no I/O, no terminal code, no `console.log` for control flow. The TUI imports it and calls into it; nothing flows the other direction.
This is what makes the prototype useful past its own lifetime: when the question's been answered, the validated reducer / machine / function set can be lifted into the real module on its own.
This is what makes the prototype useful past its own lifetime.
When the question is answered, the validated reducer, machine, or function set can be lifted into the real module on its own.
Done when all tested logic lives behind one portable, pure interface and the TUI depends on it in only one direction.
### 4. Build the smallest TUI that exposes the state
Build it as a **lightweight TUI** — on every tick, clear the screen (`console.clear()` / `print("\033[2J\033[H")` / equivalent) and re-render the whole frame.
The user should always see one stable view, not an ever-growing scrollback.
Build it as a **lightweight TUI** — on every tick, clear the screen (`console.clear()` / `print("\033[2J\033[H")` / equivalent) and re-render the whole frame. The user should always see one stable view, not an ever-growing scrollback.
Each frame has two parts, in this order:
1. **Current state**, pretty-printed and diff-friendly (one field per line, or formatted JSON).
Use **bold** for field names or section headers and **dim** for less important context (timestamps, IDs, derived values).
Native ANSI escape codes are fine — `\x1b[1m` bold, `\x1b[2m` dim, `\x1b[0m` reset.
No need to pull in a styling library unless one is already in the project.
2. **Keyboard shortcuts**, listed at the bottom: `[a] add user [d] delete user [t] tick clock [q] quit`.
Bold the key, dim the description, or vice-versa — whatever reads cleanly.
1. **Current state**, pretty-printed and diff-friendly (one field per line, or formatted JSON). Use **bold** for field names or section headers and **dim** for less important context (timestamps, IDs, derived values). Native ANSI escape codes are fine — `\x1b[1m` bold, `\x1b[2m` dim, `\x1b[0m` reset. No need to pull in a styling library unless one is already in the project.
2. **Keyboard shortcuts**, listed at the bottom: `[a] add user [d] delete user [t] tick clock [q] quit`. Bold the key, dim the description, or vice-versa — whatever reads cleanly.
Behaviour:
1. **Initialise state** — a single in-memory object/struct.
Render the first frame on start.
1. **Initialise state** — a single in-memory object/struct. Render the first frame on start.
2. **Read one keystroke (or one line)** at a time, dispatch to a handler that mutates state.
3. **Re-render** the full frame after every action — don't append, replace.
4. **Loop until quit.**
The whole frame should fit on one screen.
Done when every available action re-renders a complete one-screen view of the current state and shortcuts.
### 5. Make it runnable in one command
Add a script to the project's existing task runner (`package.json` scripts, `Makefile`, `justfile`, `pyproject.toml`).
The user should run `pnpm run <prototype-name>` or equivalent — never need to remember a path.
Add a script to the project's existing task runner (`package.json` scripts, `Makefile`, `justfile`, `pyproject.toml`). The user should run `pnpm run <prototype-name>` or equivalent — never need to remember a path.
If the host project has no task runner, just put the command at the top of the prototype's README.
If the host project has no task runner, put the command at the top of the prototype's README.
Done when a fresh user can launch the prototype with one documented command.
### 6. Hand it over
Give the user the run command.
They'll drive it themselves.
The interesting moments are when they say "wait, that shouldn't be possible" or "huh, I assumed X would be different" those are the bugs in the _idea_, which is the whole point.
If they want new actions added, add them.
Prototypes evolve.
They drive it themselves.
The interesting moments are when they say "wait, that shouldn't be possible" or "huh, I assumed X would be different" because those expose bugs in the idea.
Add actions when the feedback needs them.
### 7. Capture the answer and the prototype
Done when the user can exercise the model and the prototype exposes every state transition needed to reach a verdict.
Once the prototype has answered its question, capture the answer, then capture the prototype the way the [SKILL](SKILL.md) describes.
When the caller permits implementation, the validated reducer, machine, or function set lifts into the real module as the absorbed decision.
A planning-only caller leaves the real module unchanged.
The TUI shell rides along to the throwaway branch that keeps the prototype as a primary source.
## Production mapping
When the shared [SKILL](SKILL.md) permits production work, lift the validated reducer, machine, or function set into the real module.
Keep the TUI shell on the throwaway branch.
## Anti-patterns
- **Don't add tests.**
A prototype that needs tests is no longer a prototype.
- **Don't wire it to the real database.**
Use an in-memory store unless the question is specifically about persistence.
- **Don't generalise.**
No "what if we wanted to support X later."
The prototype answers one question.
- **Don't blur the logic and the TUI together.**
If the reducer / state machine references `console.log`, prompts, or terminal escape codes, it's no longer portable.
Keep the TUI as a thin shell over a pure module.
- **Don't ship the TUI shell into production.**
The shell is optimised for being driven by hand from a terminal.
The logic module behind it is the bit worth keeping.
- **Don't generalise.** No "what if we wanted to support X later." The prototype answers one question.
- **Don't blur the logic and the TUI together.** If the reducer / state machine references `console.log`, prompts, or terminal escape codes, it's no longer portable. Keep the TUI as a thin shell over a pure module.
- **Don't ship the TUI shell into production.** The shell is optimised for being driven by hand from a terminal. The logic module behind it is the bit worth keeping.

View File

@@ -5,43 +5,70 @@ description: Build a throwaway prototype to answer a design question. Use when t
# Prototype
A prototype is **throwaway code that answers a question**.
The question decides the shape.
A prototype is **throwaway code that answers one question**.
The question decides the branch.
## Pick a branch
## 1. Pick a branch
Identify which question is being answered — from the user's prompt, the surrounding code, or by asking if the user is around:
Identify the question from the user's prompt and surrounding code.
Ask when it remains genuinely ambiguous and the user is reachable.
- **"Does this logic / state model feel right?"** → [LOGIC.md](LOGIC.md).
Build a tiny interactive terminal app that pushes the state machine through cases that are hard to reason about on paper.
- **"What should this look like?"** → [UI.md](UI.md).
Generate several radically different UI variations on a single route, switchable via a URL search param and a floating bottom bar.
- **Does this logic or state model feel right?**
Follow [`LOGIC.md`](LOGIC.md) to build a tiny interactive terminal app that pushes the model through hard-to-reason-about cases.
- **What should this look like?**
Follow [`UI.md`](UI.md) to build several radically different UI variants on one route with a URL-controlled switcher.
The two branches produce very different artifacts — getting this wrong wastes the whole prototype.
If the question is genuinely ambiguous and the user isn't reachable, default to whichever branch better matches the surrounding code (a backend module → logic, a page or component → UI) and state the assumption at the top of the prototype.
When the user is unavailable, default to Logic for a backend module and UI for a page or component, then state the assumption in the prototype.
## Rules that apply to both
Done when exactly one branch and one design question govern the prototype.
1. **Throwaway from day one, and clearly marked as such.**
Locate the prototype code close to where it will actually be used (next to the module or page it's prototyping for) so context is obvious — but name it so a casual reader can see it's a prototype, not production.
For throwaway UI routes, obey whatever routing convention the project already uses.
Don't invent a new top-level structure.
2. **One command to run.**
Whatever the project's existing task runner supports — `pnpm <name>`, `python <path>`, `bun <path>`, etc.
The user must be able to start it without thinking.
3. **No persistence by default.**
State lives in memory.
Persistence is the thing the prototype is _checking_, not something it should depend on.
If the question explicitly involves a database, hit a scratch DB or a local file with a clear "PROTOTYPE — wipe me" name.
4. **Skip the polish.**
No tests, no error handling beyond what makes the prototype _runnable_, no abstractions.
The point is to learn something fast.
5. **Surface the state.**
After every action (logic) or on every variant switch (UI), print or render the full relevant state so the user can see what changed.
6. **Capture it when done.**
When the caller permits implementation, fold any validated decision into the real code.
A planning-only caller such as Wayfinder stops at the verdict.
Commit the prototype itself to a throwaway branch, out of main, as a **primary source**.
Write a Prototype artifact under `$(xdg-user-dir DOCUMENTS)/ai-artifacts/projects/<project>/`, where `<project>` is the lowercase basename of the current working directory, following the vault's `AGENTS.md`.
Include the question, a context pointer to the branch, run instructions, and the verdict, plus screenshots and useful code snippets where they help preserve the result.
The main branch keeps only a validated decision that the caller permitted the skill to fold in.
## Common rules
- **Throwaway from day one.**
Locate the code close to where it would be used, but name it so nobody mistakes it for production.
Follow the project's routing and source-layout conventions rather than inventing a new top-level structure.
- **One command to run.**
Use the project's existing task runner so the user does not need to remember a path or setup sequence.
- **No persistence by default.**
Keep state in memory unless persistence is the question being tested.
Use an unmistakably disposable database or local file when that question requires one.
- **Skip polish.**
Add no tests, production-grade error handling, speculative abstractions, or unrelated cleanup.
- **Surface state.**
Show the full relevant state after every Logic action or UI variant switch.
## 2. Build and reach a verdict
Follow the selected branch through its handover step and iterate on the prototype in response to the user's feedback.
Do not treat a runnable prototype as the result.
The result is the verdict that answers the design question.
Done when the user has reached an explicit verdict or stated that the prototype did not resolve the question.
## 3. Capture the primary source
Commit the complete prototype to a throwaway branch outside main.
The branch is the primary source.
Resolve the AI artifacts vault through `$(xdg-user-dir DOCUMENTS)/ai-artifacts` and read its `AGENTS.md` before writing.
Use the lowercase basename of the current working directory as the project.
When the caller provides an allocated filename and `parent`, use them exactly and do not advance `.counter`.
Create only the Prototype artifact and leave the parent artifact unchanged.
Otherwise, allocate the next vault-sequence identifier and name the artifact `<NNN>-<project>-<subject-slug>-prototype.md`.
Include `parent` only when an earlier artifact directly caused the prototype.
The Prototype artifact links the throwaway branch and preserves the question, run instructions, verdict, and branch-appropriate evidence:
- UI evidence uses screenshots.
- Logic evidence uses useful code snippets and, where needed, a short interaction transcript.
Done when the complete prototype is committed outside main and exactly one Prototype artifact preserves the result according to the vault convention.
## 4. Fold in the decision when permitted
A planning-only caller such as Wayfinder stops after the verdict and leaves production code unchanged.
Otherwise, fold the validated decision into production only when the caller permits implementation.
Follow the selected branch's **Production mapping** and keep all other throwaway code out of main.
Done when production is unchanged for a planning-only run, or contains only the permitted validated decision for an implementation run.

View File

@@ -1,10 +1,8 @@
# UI Prototype
Generate **several radically different UI variations** on a single route, switchable from a floating bottom bar.
The user flips between variants in the browser, picks one (or steals bits from each), then throws the rest away.
Generate **several radically different UI variations** on a single route, switchable from a floating bottom bar. The user flips between variants in the browser, picks one (or steals bits from each), then throws the rest away.
If the question is about logic/state rather than what something looks like — wrong branch.
Use [LOGIC.md](LOGIC.md).
If the question is about logic/state rather than what something looks like — wrong branch. Use [LOGIC.md](LOGIC.md).
## When this is the right shape
@@ -15,32 +13,21 @@ Use [LOGIC.md](LOGIC.md).
## Two sub-shapes — strongly prefer sub-shape A
A UI prototype is much easier to judge when it's **butting up against the rest of the app** — real header, real sidebar, real data, real density.
A throwaway route on its own is a vacuum: every variant looks fine in isolation.
Default to sub-shape A whenever there's a plausible existing page to host the variants.
Only reach for sub-shape B if the prototype genuinely has no nearby home.
A UI prototype is much easier to judge when it's **butting up against the rest of the app** — real header, real sidebar, real data, real density. A throwaway route on its own is a vacuum: every variant looks fine in isolation. Default to sub-shape A whenever there's a plausible existing page to host the variants. Only reach for sub-shape B if the prototype genuinely has no nearby home.
### Sub-shape A — adjustment to an existing page (preferred)
The route already exists.
Variants are rendered **on the same route**, gated by a `?variant=` URL search param.
The existing data fetching, params, and auth all stay — only the rendering swaps.
This is the default.
Pick it unless there's a specific reason not to.
The route already exists. Variants are rendered **on the same route**, gated by a `?variant=` URL search param. The existing data fetching, params, and auth all stay — only the rendering swaps. This is the default; pick it unless there's a specific reason not to.
If the prototype is for something that doesn't yet have a page but *would naturally live inside one* (a new section of the dashboard, a new card on the settings screen, a new step in an existing flow) — that's still sub-shape A.
Mount the variants inside the host page.
If the prototype is for something that doesn't yet have a page but *would naturally live inside one* (a new section of the dashboard, a new card on the settings screen, a new step in an existing flow) — that's still sub-shape A. Mount the variants inside the host page.
### Sub-shape B — a new page (last resort)
Only use this when the thing being prototyped genuinely has no existing page to live inside — e.g. an entirely new top-level surface, or a flow that can't be embedded anywhere sensible.
Create a **throwaway route** following whatever routing convention the project already uses — don't invent a new top-level structure.
Name it so it's obviously a prototype (e.g. include the word `prototype` in the path or filename).
Same `?variant=` pattern.
Create a **throwaway route** following whatever routing convention the project already uses — don't invent a new top-level structure. Name it so it's obviously a prototype (e.g. include the word `prototype` in the path or filename). Same `?variant=` pattern.
Before committing to sub-shape B, sanity-check: is there really no existing page this could be embedded in?
An empty route hides design problems that a populated one would expose.
Before committing to sub-shape B, sanity-check: is there really no existing page this could be embedded in? An empty route hides design problems that a populated one would expose.
In both sub-shapes the floating bottom bar is identical.
@@ -48,8 +35,7 @@ In both sub-shapes the floating bottom bar is identical.
### 1. State the question and pick N
Default to **3 variants**.
More than 5 stops being radically different and starts being noise — cap there.
Default to **3 variants**. More than 5 stops being radically different and starts being noise — cap there.
Write down the plan in one line, in the prototype's location or a top-of-file comment:
@@ -57,18 +43,19 @@ Write down the plan in one line, in the prototype's location or a top-of-file co
This works whether the user is here to push back or not.
Done when the prototype states one concrete UI question, its host route, and a variant count from three through five.
### 2. Generate radically different variants
Draft each variant.
Hold each one to:
Draft each variant. Hold each one to:
- The page's purpose and the data it has access to.
- The project's component library / styling system (TailwindCSS, shadcn, MUI, plain CSS, whatever).
- A clear exported component name, e.g. `VariantA`, `VariantB`, `VariantC`.
Variants must be **structurally different** — different layout, different information hierarchy, different primary affordance, not just different colours.
Three slightly-tweaked card grids isn't a UI prototype, it's wallpaper.
If two drafts come out too similar, redo one with explicit "do not use a card grid" guidance.
Variants must be **structurally different** — different layout, different information hierarchy, different primary affordance, not just different colours. Three slightly-tweaked card grids isn't a UI prototype, it's wallpaper. If two drafts come out too similar, redo one with explicit "do not use a card grid" guidance.
Done when every variant materially differs in layout, information hierarchy, and primary affordance while using the project's existing design system.
### 3. Wire them together
@@ -87,11 +74,12 @@ return (
);
```
For sub-shape A (existing page): keep all the existing data fetching above the switcher.
Only the rendered subtree changes per variant.
For sub-shape A (existing page): keep all the existing data fetching above the switcher; only the rendered subtree changes per variant.
For sub-shape B (new page): the throwaway route under `/prototype/<name>` mounts the same switcher.
Done when one route renders every variant from the URL parameter without duplicating data loading.
### 4. Build the floating switcher
A small fixed-position bar at the bottom-centre of the screen with three pieces:
@@ -103,43 +91,31 @@ A small fixed-position bar at the bottom-centre of the screen with three pieces:
Behaviour:
- Clicking an arrow updates the URL search param (use the framework's router — `router.replace` on Next, `navigate` on React Router, etc) so the variant is shareable and reload-stable.
- Keyboard: `←` and `→` arrow keys also cycle.
Don't intercept arrow keys when an `<input>`, `<textarea>`, or `[contenteditable]` is focused.
- Keyboard: `←` and `→` arrow keys also cycle. Don't intercept arrow keys when an `<input>`, `<textarea>`, or `[contenteditable]` is focused.
- Visually distinct from the page (e.g. high-contrast pill, subtle shadow) so it's obviously not part of the design being evaluated.
- Hidden in production builds — gate on `process.env.NODE_ENV !== 'production'` or an equivalent check, so a stray prototype merge can't ship the bar to users.
Put the switcher in a single shared component so both sub-shapes can reuse it.
Locate it wherever shared UI lives in the project.
Put the switcher in a single shared component so both sub-shapes can reuse it. Locate it wherever shared UI lives in the project.
Done when mouse and keyboard controls cycle through every shareable variant without intercepting text-editing keys, and the switcher cannot render in production.
### 5. Hand it over
Surface the URL (and the `?variant=` keys).
The user will flip through whenever they get to it.
The interesting feedback is usually **"I want the header from B with the sidebar from C"** — that's the actual design they want.
Surface the URL and the `?variant=` keys.
The user flips through the variants and may combine elements rather than choosing one unchanged.
### 6. Capture the answer and clean up
Done when the user can compare every variant in its host context and the prototype exposes enough contrast to reach a verdict.
Once a variant has won, capture the answer — which variant and why — then capture the prototype the way the [SKILL](SKILL.md) describes.
When the caller permits implementation, fold the winner into the real code and move the rest onto the throwaway branch, not into main:
## Production mapping
When the shared [SKILL](SKILL.md) permits production work, keep the full variant set on the throwaway branch and apply the verdict as follows:
- **Sub-shape A** — fold the winner into the existing page and drop the losing variants and switcher from main.
- **Sub-shape B** — promote the winning variant to a real route and drop the throwaway route and switcher from main.
A planning-only caller leaves the real route unchanged.
The full set of variants is the primary source, so it lands on the throwaway branch, not the bin — variant components and the switcher left in the main branch rot fast and confuse the next reader.
- **Sub-shape B** — promote the winner to a real route and drop the throwaway route and switcher from main.
## Anti-patterns
- **Variants that differ only in colour or copy.**
That's a tweak, not a prototype.
Real variants disagree about structure.
- **Sharing too much code between variants.**
A shared `<Header>` is fine.
A shared `<Layout>` defeats the point.
Each variant should be free to throw out the layout.
- **Wiring variants to real mutations.**
Read-only prototypes are fine.
If a variant needs to mutate, point it at a stub — the question is "what should this look like", not "does the backend work".
- **Promoting the prototype directly to production.**
The variant code was written under prototype constraints (no tests, minimal error handling).
Rewrite it properly when you fold it in.
- **Variants that differ only in colour or copy.** That's a tweak, not a prototype. Real variants disagree about structure.
- **Sharing too much code between variants.** A shared `<Header>` is fine; a shared `<Layout>` defeats the point. Each variant should be free to throw out the layout.
- **Wiring variants to real mutations.** Read-only prototypes are fine. If a variant needs to mutate, point it at a stub — the question is "what should this look like", not "does the backend work".
- **Promoting the prototype directly to production.** The variant code was written under prototype constraints (no tests, minimal error handling). Rewrite it properly when you fold it in.

View File

@@ -1,14 +1,46 @@
---
name: research
description: Investigate a question against high-trust primary sources and capture the findings as a Markdown file in the AI artifacts vault. Use when the user wants a topic researched, docs or API facts gathered, or reading legwork delegated to a background agent.
description: Investigate a question against high-trust primary sources and capture the cited findings as a Research artifact in the AI artifacts vault. Use when a topic needs documentation, API, specification, source-code, or other reading legwork.
---
Spin up a **background agent** to do the research, so you keep working while it reads.
# Research
Its job:
Investigate one question and preserve the findings in one cited Research artifact.
Run in the current process.
Isolation and concurrency belong to the caller.
1. Investigate the question against **primary sources** — official docs, source code, specs, first-party APIs — not a secondary write-up of them.
Follow every claim back to the source that owns it.
2. Write the findings to a single Markdown file, citing each claim's source.
3. Save it under `$(xdg-user-dir DOCUMENTS)/ai-artifacts/projects/<project>/`, where `<project>` is the lowercase basename of the current working directory.
Read the vault's `AGENTS.md` and follow its artifact conventions.
## 1. Resolve the artifact
Resolve the vault through `$(xdg-user-dir DOCUMENTS)/ai-artifacts` and read its `AGENTS.md` before writing.
Use the lowercase basename of the current working directory as the project and create its flat `projects/<project>/` directory only when needed.
When the caller provides an allocated filename and `parent`, use them exactly and do not advance `.counter`.
Create only the Research artifact and leave the parent artifact unchanged.
Otherwise, allocate the next vault-sequence identifier through `.counter` and name the artifact `<NNN>-<project>-<subject-slug>-research.md`.
Include `parent` only when an earlier artifact directly caused the research.
Done when one authoritative output path and its metadata are settled according to the vault convention.
## 2. Investigate the question
Use primary sources such as official documentation, specifications, source code, and first-party APIs rather than relying on secondary accounts.
Follow every substantive claim back to the primary source that owns it.
Use secondary material only to discover primary sources.
When no primary source establishes a needed claim, record that limitation instead of presenting the claim as settled.
Done when the question is answered as far as primary evidence permits and every substantive claim has an owning source or an explicit evidence gap.
## 3. Write the Research artifact
Write the findings to the resolved Markdown file and follow the vault's artifact conventions.
Keep the question, findings, limitations, and citations sufficient for a future reader to evaluate the result without reconstructing the research session.
Do not create a source dump or research log.
Done when exactly one Research artifact exists at the resolved path and every substantive claim in it cites its source.
## 4. Return the result
Report the artifact path and a concise statement of what the research established or could not establish.
Done when the caller can locate the artifact and understand whether the question was resolved.

View File

@@ -8,7 +8,6 @@ Create `projects/<project>/` when a new project first needs a map.
Keep the project directory flat.
Allocate every new artifact through the vault-root `.counter`.
The coordinating Wayfinder agent reconciles duplicate identifiers after concurrent workers return.
## Names
@@ -29,6 +28,12 @@ The map is the effort's root artifact and has no `parent`.
It is an index rather than the store for ticket resolutions.
```markdown
---
status: open
tags:
- wayfinder/map
---
# <effort name>
## Destination
@@ -56,6 +61,8 @@ It is an index rather than the store for ticket resolutions.
<work consciously ruled beyond the destination>
```
Map status is `open` while any live ticket or fog remains and `complete` when neither remains.
The Frontier is a derived navigation index.
Ticket metadata is authoritative.
Repair the Frontier whenever it is missing, stale, or inconsistent with ticket state.
@@ -118,15 +125,24 @@ Use `PI_SESSION_ID` when available and an equivalent harness session identifier
Claims do not expire automatically.
The acting agent uses the available context to recover an abandoned claim.
Only resolved tickets appear under Resolutions so far.
An out-of-scope ticket is closed and linked from Out of scope with the reason it lies beyond the destination.
Only resolved tickets appear under **Resolutions so far**.
An out-of-scope ticket is closed and linked from **Out of scope** with the reason it lies beyond the destination.
## Results
A Grill or Task ticket stores its canonical result under a `## Resolution` section in that ticket.
Research and Prototype tickets leave their question in the ticket and store results in one or more child artifacts whose `parent` points to the ticket.
The called skill creates those result artifacts but does not edit the ticket or map.
The coordinating Wayfinder agent integrates the artifacts, marks the ticket resolved, and updates the map.
Research and Prototype tickets leave their question in the ticket and store the result in a child artifact whose `parent` points to the ticket.
When invoking `research` or `prototype`, provide the project artifact directory, allocated filename, and ticket wikilink that the result must use as its `parent`.
The called skill creates the result artifact but does not edit the ticket or map.
The coordinating Wayfinder agent validates the returned artifact, marks the ticket resolved, and updates the map.
If a called skill cannot honor this artifact contract, leave the ticket unresolved and record the incompatibility instead of silently storing the result elsewhere.
Navigate the artifact journey forward by finding every note whose `parent` links to the current artifact.
Do not duplicate those relationships through per-artifact Next sections.
## Concurrent writes
Re-read every shared artifact immediately before editing it.
After concurrent workers return, detect duplicate identifiers, preserve pre-existing artifacts, renumber current outputs, update their wikilinks, and advance `.counter` as required by the vault convention.
Recompute the Frontier only after returned artifacts and ticket states have been reconciled.

View File

@@ -1,36 +1,46 @@
---
name: wayfinder
description: Plan work too large and uncertain for one agent session as a durable map of decision tickets, then resolve one frontier ticket at a time until the route to the destination is clear.
description: Plan a huge chunk of work that exceeds one agent session as a durable map of decision tickets, then resolve them one at a time until the way to the destination is clear.
disable-model-invocation: true
---
# wayfinder
# Wayfinder
A loose idea has arrived that is too large for one agent session and too foggy to plan directly.
Wayfinding charts the decisions needed to reach a **destination**, then works those decisions one at a time until the route is clear.
A loose idea has arrived that is too large for one agent session and wrapped in fog.
Wayfinding charts the way to a **destination** rather than charging at it.
It creates a durable map of questions whose resolutions are decisions, findings, prototypes, or completed prerequisites rather than slices of the destination work.
Read [`ARTIFACTS.md`](ARTIFACTS.md) before charting or working a map.
It is the sole source for how maps, tickets, claims, blocking, results, and the Frontier live in the AI artifacts vault.
It is the single source of truth for how maps, tickets, claims, blocking, resolutions, and the Frontier live in the AI artifacts vault.
## Plan, don't do
Wayfinder plans by default.
The map is complete when nothing remains to decide before someone performs the destination work.
The urge to implement the destination usually marks the edge of the map.
An effort may explicitly permit execution in its Notes, but otherwise preserve resolutions and hand off rather than deliver the destination.
The urge to implement the destination usually marks the edge of the map and the time to hand off.
An effort may explicitly permit execution in its Notes, but otherwise preserve resolutions rather than deliver the destination.
The destination varies by effort and shapes every ticket.
It may be a spec to hand off, a decision to lock before planning, or a change whose route must be understood before implementation.
## Refer by name
Refer to every map and ticket by its human-readable title as a wikilink, never by a bare identifier, filename, or slug.
The artifact identifier remains inside the wikilink without standing in for the name.
## Ticket types
Every ticket is either **HITL**, worked through a live exchange with the human, or **AFK**, driven by the agent.
A HITL ticket never resolves by having the agent speak for the human.
A HITL ticket only resolves through that exchange.
The agent never speaks for the human's side.
- **Research** (AFK): Investigate knowledge outside the current working directory through the `research` skill.
- **Prototype** (HITL): Create concrete Logic or UI code to react to through the `prototype` skill.
Within Wayfinder, stop at the verdict rather than folding the result into production.
- **Grill** (HITL): Resolve a decision through the `grill` skill.
- **Research** (AFK): Investigate documentation, third-party APIs, or resources outside the current working directory through `research`.
The called skill creates a Research artifact and Wayfinder integrates it.
- **Prototype** (HITL): Raise the fidelity of a logic, state-model, or UI decision through `prototype`.
The called skill creates a Prototype artifact and Wayfinder integrates it after the human reaches a verdict.
- **Grill** (HITL): Resolve a decision through `grill`.
This is the default ticket type.
- **Task** (AFK or HITL): Perform prerequisite work that must happen before a decision can be made.
The required capability is specific to the action.
The agent performs it where possible and otherwise gives the human a precise checklist.
A Task earns its place by unblocking a decision, not by delivering part of the destination.
@@ -38,13 +48,21 @@ A Task earns its place by unblocking a decision, not by delivering part of the d
## Fog of war
The map is deliberately incomplete.
**Not yet specified** holds in-scope questions that are visible but cannot yet be stated precisely enough to ticket.
Create a ticket as soon as its question is precise, even when it is blocked and cannot yet be answered.
A fog entry may graduate into several tickets or disappear when an earlier resolution changes the route.
Beyond its tickets lies the **fog of war**, where in-scope questions are visible but cannot yet be stated precisely because they depend on unresolved questions.
Resolving a ticket clears the fog ahead of it and graduates newly precise questions into tickets.
Use this test:
- Create a ticket when the question is precise now, even if it is blocked.
- Keep an entry under **Not yet specified** when the question cannot yet be phrased precisely.
Do not pre-slice fog into speculative tickets.
One fog entry may become several tickets or disappear as the frontier advances.
The destination fixes scope.
Work beyond it belongs in **Out of scope**, never in fog.
If an existing ticket proves to be beyond the destination, close it as out of scope and link it from that section rather than recording it as a resolution.
Work beyond it belongs under **Out of scope**, never under **Not yet specified**.
When an existing ticket proves to be beyond the destination, mark it out of scope and link it from that section with the reason.
Do not record a scope boundary as a resolution on the route.
## Select the mode
@@ -56,42 +74,50 @@ Never resolve more than one non-Research ticket in a session.
## Chart the map
1. **Name the destination.**
Invoke `grill` to settle what reaching the end of this effort looks like.
Done when the destination states the spec, decision, or change the map is finding its way toward and fixes its scope.
2. **Map breadth-first.**
Invoke `grill` again to fan out across the whole space, identifying precise open questions, blocking relationships, and fog without drilling into any one answer.
If this surfaces no fog and the route fits one session, stop and ask how the user wants to proceed instead of creating a map.
Done when every visible in-scope uncertainty is either a precise ticket question or an honest fog entry.
Invoke `grill` to settle what this map is finding its way toward.
Done when the destination names the spec, decision, or change at the end of the effort and fixes its scope.
2. **Map the frontier breadth-first.**
Invoke `grill` again to fan out across the whole space without resolving any one branch in depth.
Surface every currently precise question, its blocking relationships, and the remaining fog.
If no fog remains and the whole route fits one session, stop and ask how the user wants to proceed instead of creating a map.
Done when every visible in-scope uncertainty has exactly one home as a precise ticket question or an honest fog entry.
3. **Create the map and tickets.**
Create the map first, then every currently precise ticket, then wire `blocked-by` relationships in a second pass according to `ARTIFACTS.md`.
Create the map first, then every currently precise ticket, then wire blocking relationships in a second pass according to [`ARTIFACTS.md`](ARTIFACTS.md).
Done when the map is the effort root, every precise question has one ticket, every known blocking edge is represented, and the Frontier is current.
4. **Dispatch Research.**
Invoke `research` for each Research ticket using whatever isolation or concurrency the caller provides.
Let the coordinating Wayfinder agent integrate returned Research artifacts according to `ARTIFACTS.md`.
Done when every dispatched Research result has been integrated or every Research ticket that could not run remains open with the reason visible.
5. Stop without resolving a HITL ticket.
Integrate each returned Research artifact according to [`ARTIFACTS.md`](ARTIFACTS.md).
Leave a ticket open with the reason visible when its Research run cannot complete.
Done when every dispatched result is integrated or every incomplete Research ticket records why it remains open.
5. **Stop.**
Stop without resolving a HITL ticket.
Done when charting has created and dispatched the visible route without consuming its human decision work.
## Work through the map
1. **Orient.**
Read the map at low resolution and reconcile its derived Frontier against ticket metadata.
Done when the destination, standing Notes, prior resolutions, fog, scope boundary, and current Frontier agree with the artifacts.
Read the map at low resolution rather than loading every ticket.
Reconcile its derived Frontier against ticket metadata.
Done when the destination, Notes, prior resolutions, fog, scope boundary, and current Frontier agree with the artifacts.
2. **Claim one ticket.**
Use the user-named ticket when it is actionable, otherwise claim the first Frontier ticket.
Persist the claim before doing its work.
Done when exactly one open, unblocked ticket records this session's claim.
Use the user-named ticket when it is actionable.
Otherwise take the first Frontier ticket in artifact-identifier order.
Persist the claim before doing any work.
Done when exactly one unblocked ticket records this session's claim with `status: claimed`.
3. **Resolve by type.**
Invoke `research`, `prototype`, or `grill` for those ticket types.
Invoke `research`, `prototype`, or `grill` for the corresponding ticket type.
Perform a Task through the capability or human checklist it requires.
Zoom into related artifacts only as needed rather than loading the whole effort.
Done when the ticket's question has a resolution or the prerequisite Task is complete.
Load related artifacts only when needed.
Done when the question has a resolution or the prerequisite Task is complete.
4. **Record the resolution.**
Persist the result, resolve the ticket, and add its gist and links under the map's Resolutions so far.
Persist the canonical result, resolve the ticket, and append its one-line gist and artifact links under the map's **Resolutions so far** according to [`ARTIFACTS.md`](ARTIFACTS.md).
Done when the resolution lives in exactly one canonical place and the map points to it without restating it.
5. **Advance the frontier.**
Create tickets surfaced by the resolution according to `ARTIFACTS.md`.
Create tickets surfaced by the resolution and wire their blockers.
Graduate newly precise fog, remove invalidated tickets, move beyond-destination work out of scope, and recompute the Frontier.
Done when every newly visible question has exactly one home and the map matches all current ticket metadata.
Expect concurrent sessions to edit the same effort.
Re-read shared artifacts before each write and reconcile collisions through the vault convention.
Re-read shared artifacts before each write because other sessions may edit the effort concurrently.
Done when every newly visible question has exactly one home and the map agrees with all current ticket metadata.
6. **Complete or stop.**
When no unresolved tickets or fog remain, mark the map complete and stop for an explicit handoff instruction.
Otherwise stop.
Done when the map records its current lifecycle state and no destination work has begun without permission.