MCP's Fix for Bloated Tool Catalogues Has a Bill Attached — And It's in the Docs, Not the Roadmap

The MCP maintainers' current roadmap names a problem most people building agents have felt without measuring: a server's tool catalogue is charged to the model before anyone asks a question. Under Improved primitives, the post is blunt about it — "Connecting to a server with a hundred tools means the model pays for that entire surface before the user has asked a single question, and tool selection tends to get worse as the list grows."
What makes this worth reading is not the roadmap. It's that the fix is already written down, in the client documentation, alongside a warning that it can cost more than the problem it solves — and that warning gets far less attention than the fix it qualifies.
The number the docs put on it
The client best-practices page
that shipped with the 2026-07-28 specification describes progressive
discovery: the host still calls tools/list as normal, but defers injecting
those definitions into the model's context. Instead it hands the model one
lightweight search_tools meta-tool, and loads full schemas only for what comes
back.
The page's own diagram puts figures on the difference — roughly 150,000 tokens consumed by definitions alone when everything is loaded upfront, against about 2,000 under progressive discovery. That is the documentation's own illustration rather than a published benchmark, so treat it as the maintainers' order-of-magnitude claim, not a measurement you can reproduce. It is still a useful shape: two orders of magnitude, paid before the conversation starts.
The docs also give a threshold rather than leaving it to taste:
Implement a threshold as a percentage of the context window. For example, 1%-5%.
Below that, loading everything is fine. Above it, switch.
The trap
Here is the sentence that changes how you'd build this. Still on the same page, under Interaction with Prompt Caching:
Adding or removing tool definitions mid-conversation invalidates that cache, and the resulting miss can cost more tokens than the definitions you removed.
Most providers cache the prompt prefix, and the tools array sits near the front
of it. So the mechanism that saves you 148,000 tokens of upfront definitions works
by mutating the prefix — which is precisely what a prompt cache cannot tolerate.
A prefix cache is only valid up to the first thing that changed: edit the array at
turn six and everything cached after that point goes, which on a long conversation
is nearly all of it.
The docs' own mitigations are worth reading as design constraints rather than
tips: append new definitions strictly after the cache breakpoint rather than
re-sorting the tools array, or route every call through a single stable
call_tool({name, args}) meta-tool so the array never changes at all. And treat
disconnecting a server as a conversation boundary, not a per-turn operation.
Three caches, and only one is the protocol's
This is where it gets genuinely confusing, and it's worth separating the layers because they are easy to collapse into one.
The transport cache. tools/list results carry ttlMs and cacheScope
hints, defined in the specification's caching utility. A client that honours them
skips the HTTP round trip. This one is MCP's, in the sense that the protocol
specifies it.
The host-side memo. Advice rather than protocol — a line in the client documentation's implementation guidelines, recommending you memoise a fetched definition so re-injecting it later doesn't need another call. It describes what a host should do with its own state; nothing crosses the wire. The page is explicit about what it does not do: "This is separate from what's currently in the model's context."
The provider's prompt cache. Not MCP's at all. Owned by whoever serves your model, keyed on the prefix, and invalidated by exactly the thing progressive discovery does for a living.
We ran into the first of those the hard way while writing
Part 3 of the LangGraph series —
a client-side cache=True does nothing unless the server actually advertises a
TTL, and nothing warns you. Having now read this page properly, the more useful
lesson is that even getting that right buys you a round trip and nothing else.
The schemas still land in context. Those are different problems with different
fixes, and conflating them is the default mistake.
What this means if you're scoping tools by hand
Part 4 split one MCP server between a scheduler and a billing agent by filtering the catalogue per role — a static allow-list, decided before the run and fixed for its duration.
Read against these docs, that turns out to have a property worth naming: because
the list never changes mid-conversation, it doesn't touch the prompt prefix, so
it sidesteps the cache-invalidation problem entirely. It is a cruder instrument
than search_tools — you decide up front instead of letting the model search —
but for a small catalogue split across known roles, "crude and cache-stable" may
simply be the right trade. The docs say as much in the other direction: below the
1–5% threshold, loading everything is fine.
I would not generalise that further. We have not measured cache-miss cost against definition cost on the WEC Inference API, and the honest position is that the crossover depends on your catalogue size, your conversation length, and your provider's caching behaviour. What the docs establish is that a crossover exists.
On the roadmap having no dates
The roadmap names five priority areas and commits to no release date or version number anywhere. It's fair to ask why, and the fairest answer comes from the maintainers themselves. Soria Parra, in the March roadmap:
A release-oriented roadmap implies a level of predictability that open-standards work rarely has.
That's a reasonable position for a specification developed across working groups, and March's own record backs it up. "Transport Evolution and Scalability" was one of its four priority areas, and it named the problem precisely — running MCP at scale had surfaced "a consistent set of gaps: stateful sessions fight with load balancers, horizontal scaling requires workarounds" — while committing to no date for a fix. The July specification delivered that stateless core, five months later.
Caching is a different story. It appears nowhere in the March roadmap, and progressive discovery arrived as client documentation without having been on a roadmap at all — which is worth noting before treating either roadmap as a reliable index of what is coming.
The reception hasn't been uniformly warm. The Hacker News thread
on the roadmap ran to 270 points and 161 comments, and the criticism runs in three
directions rather than one. The cost of past churn, from colingauvin:
It's unreal how bad the initial rollout was between HTTP/streaming and stdio, bearer auth and OAuth. Virtually every client/MCP server pair had a different portion of that matrix implemented.
The design itself, from nprateem:
The real disaster was making it stateful. Need to get some adults in the room.
And whether the protocol earns its complexity at all, from zackify: "I think the spec
overcomplicates everything honestly."
The first of those is the one that bears on this post. It's a fair grievance about implementation drift, and it's the risk to watch with progressive discovery too: this is currently a client-side pattern with recommended strategies rather than a specified one, which means two hosts can both be reasonable and behave differently.
The short version
The roadmap tells you tool-catalogue bloat is on the maintainers' list. The documentation tells you what to do about it now, and — to its credit — tells you in the same breath that the fix has a bill. If you're loading a large catalogue into every conversation, read that page before you build a discovery layer, and check where your provider's cache breakpoint sits before you decide the token maths is in your favour.
📖 Sources: Model Context Protocol Blog — The New MCP Roadmap · MCP Docs — Client Best Practices · Model Context Protocol Blog — The 2026 MCP Roadmap · Hacker News — New MCP Roadmap
