← writing

Grep finds the right document. Then the agent has to read the others

We asked five real questions of one repository's docs. Search found every answer, in a pile of 42 to 64 matching files. What a link saves.

Ask how to point a coding agent at the right documentation and the first page of results has good answers. Put an AGENTS.md at the root that tells the agent where the docs live. Reference the file directly in the prompt; VS Code's guide shows #file for exactly this. Build an index the agent reads first. Put retrieval in front of a big docs site.

All of that works. What none of those pages measures is what happens in between: the agent knows roughly what it needs, searches for it, and gets a pile. So we measured the pile.

The question, made concrete

Take one repository with a real docs/ folder: decisions, designs, research notes. Write down five questions that someone picking up work there actually faces, and the one document that answers each:

QuestionSearch term
What must not break in the --json contract?json
May the tool add a folder to the user's repository?git
Why is board --json a problem at scale?board
Which invariants govern folding the journal into state?state
Why is conflict-freedom not the headline?conflict

Then do what an agent without a pointer does. Search docs/, the spec and the README for the natural term, case-insensitively. Count the files that match, and the bytes it would take to open them all and find out which one matters.

What it found, twice

The first run was on 2026-09-08. Search found the answer every time, among 10 to 35 matching files. Opening all of them against opening only the five answers came to 1.3 MB → 39 KB, a gap of 34×.

We ran it again today, on the same repository. Before trusting the method we checked it: replayed against the repository as it stood on the day of the first run, it reproduced all ten recorded figures exactly. Then we ran it on the documents as they stand now:

TermFiles matching (of 69)Bytes in those filesThe answer alone
json52692,8607,241
git64782,0348,648
board42621,2806,656
state47685,6236,395
conflict52704,8656,136
Total3,486,66235,076

Every answer is in its pile. The pile is about 99 times its size.

Do not read the change from 34× to 99× as a trend. Between the two runs the documents were translated from Ukrainian to English, so English search terms now match text they could not match before. Two runs on two different corpora are not a growth curve. What today's run does show is simpler, and it does not need the comparison. In a folder of a few dozen documents about one project, a natural search term is barely a filter. git matches 64 of 69 files. Search did not fail. It returned nearly everything.

Search saves the finding, not the sifting

This is the part the advice pages skip. The agent can find the answer; the cost is in telling it apart from the other matches. That cost lands in the context window, where every file opened by mistake crowds out the code the agent was meant to work on.

And the case above is the easy one, where the agent knows which word to search for. The harder case needs no numbers at all: grep needs a term, and an agent new to a task often does not know which decision constrains it. It cannot search for a document it does not know exists.

A link is the obvious fix, and most of the advice already contains it in some form. A #file reference in a prompt is a link. An index file that maps topics to documents is a table of links.

The difference is where the link lives and who pays for it:

  • In the prompt, it is made again every session, by whoever is writing the prompt, and only if they remember the document exists
  • In an index, it is made once, but the agent still has to read the index and match the task against it, which is a smaller version of the same sifting
  • On the task itself, it is made once, by whoever knew the document mattered, and it arrives with the task for whoever picks it up next. Nothing to search, nothing to match

The third is the one that holds up when the person who knew is not the person starting the next session.

Where the numbers overstate it

These limits come with the measurement:

  • The repository is heavy on documents. A project with a thin docs/ folder has a smaller pile, and the gap shrinks with it
  • "Open every candidate" is an upper bound. A capable agent discards files by name, and these names are descriptive. How many it would really skip could not be modelled honestly, so the true cost is lower than the table, by an amount we do not know
  • The questions were written by someone who knew the answers. That is the most serious bias here, and re-running the method does not remove it
  • Today's run is on the working tree, not a commit. The documents have since been taken out of git tracking, so the second run cannot be replayed from history the way the first one could

Allow for all of them and the core finding still stands: search that returns most of a folder has not narrowed anything down.


kadence lets a task or a decision point at the document that explains it, so the document arrives with the work. The original probe is where the first run's per-question figures come from.

Where the numbers come from

  • Probe run 2026-09-28: `how do I point an AI coding agent at the right docs for a task` returned code.visualstudio.com, nextjs.org, document360.com, blog.jenuel.dev, blog.tedivm.com, builder.io, aishippingblog.com, dacharycarey.com and a goodreads author blog: nine written pages and no repository
  • Probe D, 2026-09-08: five questions, grep for the natural term across docs/, SPEC.md and README.md, all candidates opened. Candidates 27, 35, 18, 14 and 10; 1,321,537 bytes in candidates against 38,701 in the five answering documents, 34× — kadence docs/research/probe-d-docs-linkage.md
  • Re-measurement run 2026-09-28. First reproduced on the product repository at the parent of commit 5ac5cfc with `git grep -l -i`: all ten recorded figures matched exactly. Then run with `grep -rli` on the product working tree, 69 files searched (66 markdown documents, one CSV, SPEC.md, README.md), against 42 files at the probe. Candidates: json 52 (692,860 bytes), git 64 (782,034 bytes), board 42 (621,280 bytes), state 47 (685,623 bytes), conflict 52 (704,865 bytes). Total 3,486,662 bytes in candidates against 35,076 in the five answers, about 99×. The answers today, by `stat`: json → 7,241 bytes; git → 8,648 bytes; board → 6,656 bytes; state → 6,395 bytes; conflict → 6,136 bytes. Each answer was among its candidates
  • Between the two runs the product's working documents were translated from Ukrainian to English and taken out of git tracking; the answers are kadence docs/decisions/009-the-agent-contract.md, docs/decisions/007-what-goes-into-git.md, docs/research/probe-c-agent-cost.md, docs/design/state-machine.md and docs/research/probe-a-results.md, identified by their exact byte sizes at commit 5ac5cfc