← writing

I told my coding agent to search before grep. 0 of 10 did

A line in CLAUDE.md did not change what agents did first. Putting the records next to the prompt did. Ten questions, one repository, and every caveat.

I build a tool that keeps a team's tasks, decisions and notes in the git repository, so a coding agent can read what was already decided instead of guessing. It has a search command for exactly that. The obvious question was whether agents use it.

They don't. And telling them to made it worse.

The setup

Ten questions about this repository's own history, the kind a new teammate asks: why is a limit what it is, what was tried before, which approach was rejected. Claude Sonnet subagents, each loading the repository's CLAUDE.md. The prompt was the question and nothing else. It never mentioned search.

The answer to every question was written down in the journal. The agent only had to look there.

An instruction line

Before any change, 1 of 10 agents used the search. It got there at step nine, after grepping the whole tree.

So I added one line to the section of CLAUDE.md the tool maintains: search the journal before grep. Same ten questions.

0 of 10.

Every agent opened with grep over the whole repository. Correct answers stayed at 9 of 10, and the context they burned barely moved. The line was in the context of every run, and it did not change the first move.

There was a quieter failure in the same runs. Four answers out of twenty came from documents that exist on my disk and are not in git. They were right on my machine and would be unanswerable in a teammate's clone. The agent had no way to know which files the team shares and which are mine.

Records next to the question

What changed the behaviour was moving the information, not the instruction. Claude Code runs a UserPromptSubmit hook before the agent sees a prompt. I made that hook run the search on the prompt's text and put up to three records, 1 KB at most, next to it. In a demo repository, asked "why are sessions in redis?", it prints:

kadence: if this is about why or how something was decided, the journal may already say:
  DEC-1  [decision] Keep sessions in Redis
         “Keep sessions in Redis Revocation must be instant JWT: cannot revoke before expiry”
More: kadence search "…" --json · open one: task show / decision show / doc show

Same ten questions:

instruction linerecords in the prompt
Used the journal0 of 109 of 10 (6 as the first step)
Right answer9 of 1010 of 10
Context added, median~10.2k tokens~7.9k tokens
Tool calls, median6.53
Answered from files a clone lacks1–30

On seven of the ten, the hook's records named the one that answered the question. On the other three, the agent still went to the journal first and found it there. Once the journal was in view, it became the place to look.

What it costs when the question has nothing to do with it

A hook that runs on every prompt also runs on "rename this function". So I ran ten ordinary code tasks, with and without it. The hook spoke on all ten, 240 to 800 bytes each. No agent opened a journal record it did not need. Answers were the same, time was the same, and the paired median cost was about 180 tokens per task.

The obvious fix for that noise was a confidence gate: only speak when the match is strong. I measured it and threw it out. At the threshold where it went quiet on code tasks (5 of 30 still got records), it also went quiet on the questions it exists for: it spoke on 3 of 75. A short coding prompt shares more words with the journal than a paraphrased question does, so match strength cannot tell the two apart. The hook now says what it found, once per session per record, and frames it as an offer: if this is about why or how something was decided.

The limits

  • One repository, one model, ten questions. Treat it as a signal and check it on your own repository.
  • The hook was simulated. Subagents do not fire hooks, so its real output was placed in front of the prompt by hand. That measures what the records do once they arrive, not the hook firing in a live session.
  • Claude Code only. Cursor and Codex read the same AGENTS.md section. I have not measured them, and after the result above I would not assume the instruction does anything there either.
  • The search is by words. "Why did we choose Postgres" finds nothing if the record says "Sessions live in Postgres" and never uses the word choose. The hook stays silent rather than guess.

What to take from it, whatever you use

If you want an agent to use something, an instruction to go and look for it is the weakest lever you have. The first move is decided by what is already in front of the model. A hook that puts the relevant lines into the prompt, even a grep over your decision records, will do more than a paragraph in CLAUDE.md asking the agent to go and find them.

It also changes what "the context" means for a team. Whatever the hook reads is what every teammate's agent sees, so it should read the files in git, which everyone has.

If you want to try it

kadence is MIT and runs offline, with no server and no account. kadence init installs both hooks: prime at session start, and the search on each prompt. --no-hooks skips them.

npm install -g kadence
kadence init

Nobody outside this repository uses it yet. If your team hands work between people and agent sessions, I would like to hear how that goes today, and where this would break. The issues are open.

Where the numbers come from

  • Natural-mode runs, ten questions before and ten after the instruction line, Claude Sonnet subagents that load CLAUDE.md: search used by 1 of 10 at step 9, then 0 of 10; correct 9 of 10 both times; 4 of 20 answered from documents a clone does not have — kadence journal, notes on KAD-53 (2026-09-28); CHANGELOG 0.8.0
  • The prompt hook prints up to three records, 1 KB at most — kadence CHANGELOG 0.8.0
  • The same ten questions with the prompt hook's real output prepended: journal used by 9 of 10, first by 6; correct 10 of 10; median context 7.9k vs 10.9k tokens; median tool calls 3 vs 6.5 — kadence journal, notes on KAD-53; CHANGELOG 0.8.0
  • Ten ordinary code tasks with and without the hook: it spoke on 10 of 10 (240 to 800 bytes), no agent opened a journal record, paired median +180 tokens — kadence journal, notes on KAD-56 (2026-09-29)
  • Coverage gate at 0.7: the hook speaks on 3 of 75 answerable questions and 5 of 30 code prompts — kadence journal, notes on KAD-56; decision DEC-36