Token Savings & Grounding · Research

Grounding is the flip side of token savings

The most common mistake in talking about token cost and grounding is to treat them as two goals — one about cost, one about trust — and to assume you have to trade between them. You usually do not. The same discipline that makes an agent cheaper — sending less of the right knowledge — also makes it more dependable. Grounding and savings are not two goals. They are one discipline seen from two sides.

This sounds like a claim about elegance. It is a claim about design. The agent that pays for less knowledge is the agent that is easier to ground — because grounding is mostly the discipline of knowing what knowledge the agent used, and an agent that is paying for less knowledge is an agent that has less wrong knowledge to hide behind. The cost comes down and the basis comes into view at the same time.

The shared root

The shared root is the knowledge layer, designed as a delivery surface. Send the agent the knowledge it needs for the action it is about to take, in the shape it can use, with the currentness and permission attached. Do that, and two things happen at once.

The cost comes down, because the agent is not paying to move and re-read the wrong knowledge. The trust comes up, because the agent is acting on knowledge you can point to — current, permitted, and in the shape the agent can use. The cheaper run and the more dependable run are the same run, designed with the knowledge layer in mind.

This is not a model trick. It is a knowledge discipline. The model is still doing the reasoning. The knowledge layer is doing the rest.

The two mistakes that look different but are the same

The mistake that shows up as cost is the bucket instinct — dump everything in and let the model sort it out. The agent pays for everything and acts on the wrong slice. The mistake that shows up as trust is the same thing, seen from the other side — the agent acts on a slice it cannot show the basis for, because the basis was never carried with the action.

Same root. Two symptoms. One fix.

The fix is not to spend less on tokens by choosing a cheaper model. It is to send less of the right knowledge, in the right shape, only when the agent needs it, with the basis attached. That is the discipline that reduces cost and increases trust at the same time, because it is the discipline of taking the knowledge layer seriously at the point where it meets the agent.

What the discipline looks like in practice

The discipline looks like three things, done together.

Send less of the right knowledge. For each action the agent might take, decide what knowledge it actually needs — not everything it might need, the knowledge it needs for that action — and pack only that. The rest can wait until it is needed, or it can stay out of the context entirely. The agent pays for less, and the basis is smaller, and the basis is easier to ground.

Make the basis retrievable. When the agent acts, it should be able to show what knowledge it used — in what shape, from where, at what version, with what permission. Not as a separate forensic step. As part of the action itself. The same discipline that makes the basis retrievable also makes the agent stop re-reading the same knowledge, because the basis is carried with the action, not re-derived from scratch every time.

Make currentness and permission explicit. The agent should be able to tell what is current and what it is allowed to act on — not by guessing, not by assuming the latest document is the latest truth. The basis includes the version, the date, the owner who says this is the one, and the explicit statement that this supersedes the previous version. The agent can show the basis, and the company can tell whether it was the right basis.

The boring payoff

The boring payoff is the best one: an agent that costs less to run, acts on more dependable knowledge, and can show its basis when asked. That is not a model achievement. It is a knowledge achievement, delivered through a context window designed as a delivery surface rather than a bucket.

The model is still doing the reasoning. The knowledge layer is doing the rest. The cost and the trust both come from how seriously the knowledge layer is taken at the point where it meets the agent.

The flip side is not a trade. It is the same thing, seen from the other side. Send less of the right knowledge, and the agent is cheaper. Make the basis retrievable, and the agent is more dependable. Do both, and the cheaper run is the more dependable run — because they are the same run, designed with the knowledge layer in mind.

Back to blogs · Token Savings & Grounding · Research

Knowledge Sidekick — research · standards · adoption. Not-for-profit. No ads, no tracking.