← all pages

Reference · 03

Evidence and dossiers

A tracker where agents do most of the writing drowns in report-dumps unless legibility is engineered. Ours is engineered around three ideas: the glance, the artefact, and evidence named in advance.

The glance

Every note an agent leaves on a ticket is written to be read in one breath, the way an engineer reports to their CEO: "this is done", "blocked on X, diagnosis attached", "filed Y as a new issue". Detail lives in the linked artifact and never in the comment itself. An unattended run additionally opens with a fixed marker — "Generated by AI during triage." — so a reader always knows which voice is speaking.

The test a comment must pass: a colleague who was not in the session reads it once and knows what happened and what, if anything, they owe. A wall of text fails even when every sentence in it is true.

The artefact

Long-form records — grilling records, specs, plans, accumulated evidence — do not go in comments. Each effort has one accreting page, its artefact, named after the effort and never after a ticket: tickets are plural and get renumbered, the effort is one thing. The top of the page is the current contract — what to build, the evidence, the plan — and below it sits one section per stage, each with a stable anchor so a ticket can link the part that concerns it. A stage already written is a record of what was believed at the time, so corrections are dated entries and never edits to history. The tracker keeps the glance plus the link.

Where the page lives varies per repo, and the repo's own seeded instructions say which: a file in the repo, for work read from a checkout; a Claude artefact, for a link shared in a hurry; or a published site, for colleagues who need a URL that renders. Redkale's is an internal, access-walled dossier site. What does not vary is that it is HTML — the reader is a person, and a rendered page is read in a fraction of the time a markdown wall is.

The split is an audience seam, and it decides format. Artifacts agents consume — specs, decision records, glossaries — are markdown in the repo: versioned, diffable, readable by a session with no browser. Artifacts humans read long-form are HTML at the artefact. The process itself never cares what evidence is marked up in; that doctrine is seeded into each repo, read once at the start of a session rather than carried around as a shared skill every session re-reads.

Evidence, named in advance

Every gate in the machine names its evidence kind before the work starts, and there are five kinds to name: a screenshot or GIF, a terminal transcript, a query and its rows, a test run, a live URL. A spec's validation gate says which of them a human will look at to say "yes, it works", and a passing test suite on its own is deliberately not an answer. Done's evidence is the canary: proof the deployed change works in production, not proof the deploy pipeline went green.

Evidence is the closing act of implementing, but it is produced in a sub-agent with fresh context that was not part of the session that wrote the code, so the judge sits outside the system under test. Where a named kind cannot be produced, the comment says which and why rather than substituting another. At spec level the same rule takes a shape: every breakdown of a spec ends with a verify ticket, blocked by all the others, whose whole deliverable is the feature's evidence posted on the spec. If that proof cannot be produced, the spec is not done — and because human review happens on the ticket a human opened, the sub-issues go straight to Done.

Repos whose merge changes no production system declare, in their own delivery facts, what shipping means there — one repo here ships by a one-line team announcement posted only after its canary, because the message claims the change is live and the claim has to be true first.

Why this is a process concern at all

Principle two says the tracker is the system of record; this page is what makes that record worth reading. A record only humans can navigate defeats the agents; a record only agents can parse defeats the humans. The seam — glance for the board, markdown for the sessions, HTML for the long read — is how one system of record serves both.