LangChain Just Made MCP First-Class — We Ran It the Same Day, and Three Things Don't Work Yet

In July the MCP specification ripped out sessions. We wrote at the time that the change was infrastructure, not changelog — that a stateless core would let servers sit behind ordinary load balancers and survive redeploys, and that clients would need to catch up.
Today LangChain caught up. MCP support moved out of the separate
langchain-mcp-adapters package and into langchain itself, rebuilt on FastMCP,
with two features the old spec made impossible: elicitation as a LangGraph
interrupt, and a cacheable tool catalog.
We installed it the same afternoon and pointed it at a server we wrote against the new spec. The announcement is accurate. The ecosystem around it is not ready — and one of the gaps quietly converts a human-approval gate into a sentence the model invents.
What actually shipped
pip install "langchain[mcp]", requiring 1.4.0 or newer, and in beta — it says
so on import, which you can see in our captures further down. Python today,
TypeScript "soon to follow."
- One class, in the main package.
MultiServerMCPClientcollapses intoMCPAdapter.async with MCPAdapter(url) as adapter:thenawait adapter.list_tools(), and what comes back are ordinary LangChain tools that go anywhere tools go. - Built on FastMCP, so transports, auth (bearer, OAuth 2.1, machine-to-machine,
CIMD, or any
httpx2.Auth), connection management and protocol negotiation come from the client underneath. MCP now has two eras — the 2025-11-25 handshake protocol and the 2026-07-28 stateless one — and the FastMCP client picks one per connection: "it tries the new protocol and falls back to the handshake for a server that hasn't upgraded." - Elicitation via interrupts. When a tool on the MCP server can't finish without
asking a human, your agent's run pauses as a LangGraph
interrupt(), and you resume it with a structured answer. This is only possible because the stateless spec turned a mid-call question into a retry-able round rather than something held open on a socket. - Client-side caching.
fastmcp.Client(url, cache=True)honours the server's freshness hints so the tool catalog isn't re-fetched every run. ClientGroupfor several servers at once, each keeping its own era and credentials, with tool names prefixed by server —billing_searchanddocs_searchstay distinct.
The framing in the announcement is the same one we used in July: "a redeploy no longer kills live sessions, because there are none."
What the stateless spec actually removed
One tool call, before and after, is the clearest picture of what changed — so here it is, with a third panel for what we actually got:
The first two panels are the pitch, and the pitch is real. The third is what a brand-new server does on the afternoon the client ships.
Three things that don't work yet
We built an MCP server exposing a small booking-and-invoicing database, pointed a LangChain agent at it, and gated a refund behind a human. (The agent's own model runs on a self-hosted gateway — that sits between the agent and the LLM, not between the agent and MCP.) It works — that's the tutorial. Getting there surfaced three gaps between what's written and what runs.
1. The spec is stateless. Your server isn't, by default.
The first request to a brand-new FastMCP 4.0.2 server — here a plain curl asking
for tools/list, so nothing client-side can be blamed for it:

An error about a concept the specification deleted five weeks ago. The client is stateless; the server still defaults to the stateful transport.
The fix isn't a header, a client option, or anything you send. It's an argument on
the call that starts your server — the last line of the server file, where mcp
is your FastMCP instance:
from fastmcp import FastMCP
mcp = FastMCP("office-tools")
# ... your @mcp.tool functions ...
if __name__ == "__main__":
# Without stateless_http=True this server answers a 2026-07-28 client
# with "Bad Request: Missing session ID".
mcp.run(transport="http", host="127.0.0.1", port=8770, stateless_http=True)
Nothing in the announcement says so — reasonably enough, since the post is about the
client. But it means the client half of the upgrade is one pip install, and the
server half is a flag you have to know exists.
Two flags exist, and run_http_async (which mcp.run calls for HTTP) accepts both.
From its own docstring:
stateless_http: Whether to use stateless HTTP (defaults to settings.stateless_http)
stateless: Alias for stateless_http for CLI consistency
One switch, two names, and the setting it defaults to is off.
2. ctx.elicit is the old API — and the error goes to the model, not to you
Every elicitation example you can find calls await ctx.elicit(message, response_type)
inside the tool body, where ctx is the Context object FastMCP passes to your tool.
That is the handshake-era mechanism: it blocks mid-execution and
speaks over the session's back-channel — the back-channel a stateless connection
doesn't have. FastMCP's own docs are explicit that it's for connections
≤ 2025-11-25, and promise that calling it on a modern one "raises a clear era
error rather than failing obscurely."
It does raise one. That isn't the problem. We rebuilt the refund tool around
ctx.elicit and ran the same agent against it:

Read that in order. FastMCP raised the era error, exactly as documented. LangChain
caught it and turned it into a ToolMessage with status="error". That is deliberate:
langchain/mcp/tools.py wires a handler whose docstring says it exists to hand the
server's own error detail to the model "instead of ending the run." The model read
that error and told the user:
I have located Maria Alvarez and identified her open invoice (ID: 2) for $80.00. I have initiated the refund request for this invoice. Please note that a human must approve the amount and provide a reason to complete the refund.
No interrupt was raised. The process exited 0. Nothing is pending, nothing is
waiting for a human, and no one will ever be asked — and the invoice is still open,
so the refund didn't happen either. What you get is not a destructive action slipping
past a gate; it's an approval workflow that silently doesn't exist, described in
fluent English by a model that read the error and paraphrased it as progress.
The fix is that the modern pattern has the tool return an InputRequiredResult
describing what it needs, and exit. Your agent surfaces that as the LangGraph interrupt,
a human answers, and the MCP client re-issues the same tools/call with the answer
attached. But the failure mode is the story: an era error is a fine thing to raise
into a program, and a terrible thing to hand to a language model that is rewarded for
sounding helpful.
3. cache=True is necessary, not sufficient
The client cache respects the ttlMs and cacheScope hints a server attaches to
its tools/list response, and only against modern-era servers that send them. A default
FastMCP server sends neither — we dumped our own tools/list response and it carries no
cache hints at all. So you turn the cache on, call list_tools(cache_mode="use") twice
back to back, time both, and get:

You conclude the cache is broken. It isn't — there was nothing to cache. Worth noting
the cache belongs to the fastmcp.Client, not to MCPAdapter, and one client per
caller keeps catalogs from crossing between tenants.
And two smaller ones, in the announcement itself
The elicitation snippet raises AttributeError. The post reads the interrupt as
paused["__interrupt__"][0].value.requests[0] — attribute access. In
langchain/mcp/elicitation.py on 1.4.0:
class MCPElicitationInterrupt(TypedDict):
type: Literal["mcp_elicitation"]
tool_name: str
requests: list[MCPElicitationRequest]
A TypedDict is a dict at runtime, so .requests doesn't resolve. It's
value["requests"] — which the same snippet gets right a few lines later, reading
question["key"] by subscript.
The elicitation docs link 404s. The post closes the section by pointing at
docs.langchain.com/oss/python/langchain/mcp/elicitation for "declining a question, and
gating destructive tools behind the same approval flow." That page returns 404 as of
publication, while the parent page and its other children — mcp,
mcp/connections, mcp/tools, mcp/auth — all resolve. It is, inconveniently, the
one page that would have documented the approval flow gap 2 shows falling over.
Why this matters for you, specifically
If you consume MCP servers someone else runs, this is straightforwardly good news and you'll notice mostly the shorter import path.
If you write MCP servers, the July revision moved the ground and the tooling is
still settling. Three things are now yours to get right: your server does not become
stateless because the spec did — you set a flag; a tool that needs human input has to
be written to be re-entered rather than resumed — interrupt() unwinds the whole
call, so the tool body runs again from the top when you answer; and any
elicitation example predating August is teaching you an API whose failure lands in the
model's context instead of your logs.
That last one generalises past MCP. As frameworks get better at keeping agents alive through errors, the class of bug that ends a run is shrinking and the class that gets narrated to a user is growing. A gate that fails closed is a bug you find in testing. A gate that was never installed, described by a model as awaiting your approval, is one you find in an audit.
The direction is right. The stateless core is what makes an agent's human-approval pause survive a redeploy, and that's a real capability rather than a refactor. Just don't expect your first request to succeed.
📖 Sources: LangChain — MCP in LangChain: stateless protocol, elicitation, and more · MCP in LangChain docs · Migrating from langchain-mcp-adapters · FastMCP client documentation · FastMCP elicitation · MCP 2026-07-28 specification
Versions under test: langchain 1.4.0, fastmcp 4.0.2, mcp 2.1.1, model served over a self-hosted gateway. The ctx.elicit capture is a render of real captured output from that run, not a screen grab.
