What an AI agent's context actually costs
Everyone measures the context window. Almost nobody measures what fills it before the agent reads a word — so we measured our own, and gave a number back.
Everybody quotes the size of the context window. Almost nobody measures what is already sitting in it before the agent has read a word of your question.
That second number is the one that decides whether a tool is usable, and it splits into three parts that behave completely differently:
- Ambient cost — loaded every session, whether or not it is used. Tool definitions, instruction files, system prompts.
- The answer — what one question costs to ask.
- The search — what the agent spends finding out which question to ask.
The first is a tax. The third is where the real money goes. And almost every comparison you will read only discusses the second.
Measure bytes, and say so
No tokenizer runs offline, and adding one to a project just to produce a marketing number is a dependency bought with a measurement. So everything below is bytes, and where a token figure helps it is bytes ÷ 3.5 — the usual ratio for punctuation-heavy JSON — and it is marked as an estimate.
Every comparison here is a ratio, so the divisor cancels out. That is the only reason the shortcut is defensible, and it is worth saying rather than hoping nobody asks.
The ambient tax: the number we borrowed, then had to give back
The most-cited figure in this whole conversation is how much a large MCP server costs before a session starts. We cited it too. Then we looked at where it comes from.
For the same server, the published counts are 17,600 tokens (GitHub's own figure), ~42,000 from one independent measurement, and 55,000 across 93 tool definitions from another. Those are not small disagreements. They are not necessarily wrong either — they measure different versions, different tool sets, different tokenizers — which is precisely the problem with quoting any of them about your project.
So we measured ours. Thirteen commands, generated into tool definitions from our own contract: 3,058 bytes of ambient context, against 591 bytes for the plain instruction section that does the same job. The difference is about ~700 tokens a session.
That is a real saving and it is nothing like the number we had been repeating. We had been using somebody else's evidence for a decision it did not support. The decision survived on its own merits — an instruction file works for agents with no client at all — but the argument for it had to be thrown away.
If you take one thing from this post: the ambient cost of a tool is a property of that tool, and borrowing a headline figure from a bigger one tells you nothing. Thirteen commands do not cost what ninety-three do.
The answer: what should be constant, and usually is not
Here is the part worth designing for. We built four repositories at 10, 50, 200 and 1,000 tasks, each with real movement, comments and logged hours, then measured every path an agent can take through them.
One task, read as an agent reads it, is 982 bytes — at 10 tasks and at 1,000. The history behind that answer grows from 5 KB → 528 KB over the same range.
The multiplier is not the point, and neither is our number. The shape is: the cost of asking does not grow with the history that makes the answer worth having. A store where it does is a store that gets slower to consult exactly as it becomes worth consulting — which is the failure mode you will not notice in a demo repository and cannot miss in a real one.
Test any tool you are considering this way. Ask it the same question at two very different project sizes and compare the bytes it hands back. It takes ten minutes and it is the most informative thing you can do.
The search: where it actually went wrong for us
Our own measurement found two failures, and both were in the third category.
A listing that did not scale. Asking for the whole board at 1,000 tasks returned 803 KB — larger than the history it was folded from, because every task was emitted in full, in every column. That is past the context window of most models, for the exact audience the flag existed to serve. The fix was letting the caller ask for the fields it reads: on a 200-task board, 130,799 → 11,015 bytes.
A truncation nobody's tests caught. Every response above 128 KiB came back cut off mid-string — but only when output went to a pipe, which is exactly how an agent reads it. Redirected to a file, the same command produced valid JSON. The suite was green and every fixture in it was small.
Both are unglamorous, and both are the kind of thing that only shows up when you generate a repository big enough to hurt and then measure it. Neither would have been found by reasoning about the design.
What none of this tells you
It measures what an answer costs. It does not measure whether anyone wants the answer. A cheap answer to a question nobody asks is worth nothing, and no amount of byte-counting closes that gap — it is a different kind of work, with users in it, and we have been honest on this site about not having done enough of it yet.
But if you are choosing between tools your agents will read every session, the three numbers to ask any of them for are: what does it cost me before I ask, what does one answer cost, and does that second number grow with my project.
Most tools cannot tell you. That is itself an answer.
The measurements above come from kadence's own probe — the full table, the four repositories and the raw data are at what an answer costs an AI agent. Every figure names the test that keeps it true.
Where the numbers come from
- Four repositories built through the real CLI at 10, 50, 200 and 1,000 tasks, then measured along every path an agent can take — kadence docs/research/probe-c-agent-cost.md, summarised at /docs/proof/agent-cost
- 591 bytes for the instruction section and 3,058 bytes for 13 tools generated from our own contract — kadence docs/research/probe-c-agent-cost.md
- 803 KB for a 1,000-task board, and responses above 128 KiB truncated mid-string when stdout was a pipe before --fields existed — kadence docs/research/probe-c-agent-cost.md
- Published counts for the same GitHub MCP server: 17,600 tokens (GitHub's own figure), ~42,000 (Nebulagg) and 55,000 across 93 tool definitions (Piotr Hajdas), collected in getunblocked.com/blog/github-mcp-token-cost and dev.to/kenimo49