kadence
Proof

Ideas the project already ruled out

Ten prompts that re-propose a rejected idea, run with kadence, without it, and with a hand-kept CLAUDE.md. Graded blind.

A decision is worth recording if it stops the same idea coming back. This measures exactly that: when a prompt proposes something the project already rejected, does the agent push back with the reason?

The result

Pushed back, with the reason
Without kadence: the same repository, no journal, no hooks6 of 10
With kadence: journal and hooks9 of 10
No journal, but a hand-kept CLAUDE.md that restates most of the decisions9 of 10

kadence matched a memory file that someone writes and keeps current by hand, without anyone writing it. It did not beat it. That file restates the decisions behind 7 of the 10 prompts; two of the other three it caught from the code comments, which carry many of this repository's reasons.

The hook on its own

A separate, earlier measurement looks only at the prompt hook, not at what the agent then said: on the same kind of prompt, did the decision that rejected the idea arrive before the agent answered? It did for 8 of 10, with no false alarms on ten ordinary code prompts (kadence 0.8.1). That is the figure on the front page. The table above is the stricter test: whether the agent, with that decision in hand, actually pushed back.

How it was run

  • Three local clones of the kadence repository at one commit. With: the journal kept and kadence init run for its session and prompt hooks, the decision sections removed from CLAUDE.md. Without: the journal and those sections removed. Hand-kept: the journal removed, CLAUDE.md left whole.
  • Ten prompts, each re-proposing an alternative the journal records as rejected: an async core, committing the cache, a delete that erases, and so on. Each ends "tell me briefly how you would go about it", with file edits disallowed.
  • One Claude Sonnet session per prompt and arm, 30 sessions. Read, search and the kadence CLI allowed in every arm.
  • Graded blind by a separate agent: the three answers to each prompt shuffled and unlabelled, the grader given only the prompt and the recorded reason.

Read before quoting it

  • Ten prompts, one model. One different answer moves an arm by ten points. An earlier run of the same setup gave without 5 of 10.
  • This repository is not typical. Its code comments carry many of the reasons, and the agents without kadence found them by reading the code. A project whose reasons live in people's heads and closed pull requests would be a harder test, and that has not been run.
  • The one kadence missed: auto-committing the journal. The decision is in the journal; the prompt hook did not bring it up. That is a retrieval gap in the product, not in the test.
  • The first run was unfair to kadence. The sessions could not run the CLI the way the repository's own instructions say to, so several knew a decision existed but could not read its reason. The table above is the second run, with that fixed in every arm.

On this page