Ideas the project already ruled out
Ten prompts that re-propose a rejected idea, run with kadence, without it, and with a hand-kept CLAUDE.md. Graded blind.
A decision is worth recording if it stops the same idea coming back. This measures exactly that: when a prompt proposes something the project already rejected, does the agent push back with the reason?
The result
| Pushed back, with the reason | |
|---|---|
| Without kadence: the same repository, no journal, no hooks | 6 of 10 |
| With kadence: journal and hooks | 9 of 10 |
| No journal, but a hand-kept CLAUDE.md that restates most of the decisions | 9 of 10 |
kadence matched a memory file that someone writes and keeps current by hand, without anyone writing it. It did not beat it. That file restates the decisions behind 7 of the 10 prompts; two of the other three it caught from the code comments, which carry many of this repository's reasons.
The hook on its own
A separate, earlier measurement looks only at the prompt hook, not at what the agent then said: on the same kind of prompt, did the decision that rejected the idea arrive before the agent answered? It did for 8 of 10, with no false alarms on ten ordinary code prompts (kadence 0.8.1). That is the figure on the front page. The table above is the stricter test: whether the agent, with that decision in hand, actually pushed back.
How it was run
- Three local clones of the kadence repository at one commit. With: the
journal kept and
kadence initrun for its session and prompt hooks, the decision sections removed from CLAUDE.md. Without: the journal and those sections removed. Hand-kept: the journal removed, CLAUDE.md left whole. - Ten prompts, each re-proposing an alternative the journal records as rejected: an async core, committing the cache, a delete that erases, and so on. Each ends "tell me briefly how you would go about it", with file edits disallowed.
- One Claude Sonnet session per prompt and arm, 30 sessions. Read, search and the kadence CLI allowed in every arm.
- Graded blind by a separate agent: the three answers to each prompt shuffled and unlabelled, the grader given only the prompt and the recorded reason.
Read before quoting it
- Ten prompts, one model. One different answer moves an arm by ten points. An earlier run of the same setup gave without 5 of 10.
- This repository is not typical. Its code comments carry many of the reasons, and the agents without kadence found them by reading the code. A project whose reasons live in people's heads and closed pull requests would be a harder test, and that has not been run.
- The one kadence missed: auto-committing the journal. The decision is in the journal; the prompt hook did not bring it up. That is a retrieval gap in the product, not in the test.
- The first run was unfair to kadence. The sessions could not run the CLI the way the repository's own instructions say to, so several knew a decision existed but could not read its reason. The table above is the second run, with that fixed in every arm.