Skip to main content

LangChain Just Made MCP First-Class — We Ran It the Same Day, and Three Things Don't Work Yet

· 10 min read
Rafael Fernandes
NLP Engineer & Tech Writer at WiLine
Share:
Agent Frameworks · AI NewsMCP support lands in the main LangChain package

In July the MCP specification ripped out sessions. We wrote at the time that the change was infrastructure, not changelog — that a stateless core would let servers sit behind ordinary load balancers and survive redeploys, and that clients would need to catch up.

Today LangChain caught up. MCP support moved out of the separate langchain-mcp-adapters package and into langchain itself, rebuilt on FastMCP, with two features the old spec made impossible: elicitation as a LangGraph interrupt, and a cacheable tool catalog.

We installed it the same afternoon and pointed it at a server we wrote against the new spec. The announcement is accurate. The ecosystem around it is not ready — and one of the gaps quietly converts a human-approval gate into a sentence the model invents.

What actually shipped​

pip install "langchain[mcp]", requiring 1.4.0 or newer, and in beta — it says so on import, which you can see in our captures further down. Python today, TypeScript "soon to follow."

  • One class, in the main package. MultiServerMCPClient collapses into MCPAdapter. async with MCPAdapter(url) as adapter: then await adapter.list_tools(), and what comes back are ordinary LangChain tools that go anywhere tools go.
  • Built on FastMCP, so transports, auth (bearer, OAuth 2.1, machine-to-machine, CIMD, or any httpx2.Auth), connection management and protocol negotiation come from the client underneath. MCP now has two eras — the 2025-11-25 handshake protocol and the 2026-07-28 stateless one — and the FastMCP client picks one per connection: "it tries the new protocol and falls back to the handshake for a server that hasn't upgraded."
  • Elicitation via interrupts. When a tool on the MCP server can't finish without asking a human, your agent's run pauses as a LangGraph interrupt(), and you resume it with a structured answer. This is only possible because the stateless spec turned a mid-call question into a retry-able round rather than something held open on a socket.
  • Client-side caching. fastmcp.Client(url, cache=True) honours the server's freshness hints so the tool catalog isn't re-fetched every run.
  • ClientGroup for several servers at once, each keeping its own era and credentials, with tool names prefixed by server — billing_search and docs_search stay distinct.

The framing in the announcement is the same one we used in July: "a redeploy no longer kills live sessions, because there are none."

What the stateless spec actually removed​

One tool call, before and after, is the clearest picture of what changed — so here it is, with a third panel for what we actually got:

Three sequence diagrams comparing a tool call. Under the 2025-11-25 handshake era the agent sends initialize, receives a session id, then sends tools/call with that id. Under the 2026-07-28 stateless spec there is no handshake and the agent sends tools/call directly. Against a default FastMCP server on the new spec, the same request is refused with Bad Request: Missing session ID.

The first two panels are the pitch, and the pitch is real. The third is what a brand-new server does on the afternoon the client ships.

Three things that don't work yet​

We built an MCP server exposing a small booking-and-invoicing database, pointed a LangChain agent at it, and gated a refund behind a human. (The agent's own model runs on a self-hosted gateway — that sits between the agent and the LLM, not between the agent and MCP.) It works — that's the tutorial. Getting there surfaced three gaps between what's written and what runs.

1. The spec is stateless. Your server isn't, by default.​

The first request to a brand-new FastMCP 4.0.2 server — here a plain curl asking for tools/list, so nothing client-side can be blamed for it:

A curl POST to the MCP endpoint returning Bad Request: Missing session ID with JSON-RPC error code -32600

An error about a concept the specification deleted five weeks ago. The client is stateless; the server still defaults to the stateful transport.

The fix isn't a header, a client option, or anything you send. It's an argument on the call that starts your server — the last line of the server file, where mcp is your FastMCP instance:

office_tools.py
from fastmcp import FastMCP

mcp = FastMCP("office-tools")

# ... your @mcp.tool functions ...

if __name__ == "__main__":
# Without stateless_http=True this server answers a 2026-07-28 client
# with "Bad Request: Missing session ID".
mcp.run(transport="http", host="127.0.0.1", port=8770, stateless_http=True)

Nothing in the announcement says so — reasonably enough, since the post is about the client. But it means the client half of the upgrade is one pip install, and the server half is a flag you have to know exists.

Two flags exist, and run_http_async (which mcp.run calls for HTTP) accepts both. From its own docstring:

stateless_http: Whether to use stateless HTTP (defaults to settings.stateless_http)
stateless: Alias for stateless_http for CLI consistency

One switch, two names, and the setting it defaults to is off.

2. ctx.elicit is the old API — and the error goes to the model, not to you​

Every elicitation example you can find calls await ctx.elicit(message, response_type) inside the tool body, where ctx is the Context object FastMCP passes to your tool. That is the handshake-era mechanism: it blocks mid-execution and speaks over the session's back-channel — the back-channel a stateless connection doesn't have. FastMCP's own docs are explicit that it's for connections ≤ 2025-11-25, and promise that calling it on a modern one "raises a clear era error rather than failing obscurely."

It does raise one. That isn't the problem. We rebuilt the refund tool around ctx.elicit and ran the same agent against it:

The agent run printing interrupt raised? False, three tool messages where issue_refund has status=error carrying 'elicitation via server-initiated requests is unavailable on 2026-07-28 connections', the model replying that it has initiated the refund and a human must approve it, and a sqlite3 query showing invoice 2 still open

Read that in order. FastMCP raised the era error, exactly as documented. LangChain caught it and turned it into a ToolMessage with status="error". That is deliberate: langchain/mcp/tools.py wires a handler whose docstring says it exists to hand the server's own error detail to the model "instead of ending the run." The model read that error and told the user:

I have located Maria Alvarez and identified her open invoice (ID: 2) for $80.00. I have initiated the refund request for this invoice. Please note that a human must approve the amount and provide a reason to complete the refund.

No interrupt was raised. The process exited 0. Nothing is pending, nothing is waiting for a human, and no one will ever be asked — and the invoice is still open, so the refund didn't happen either. What you get is not a destructive action slipping past a gate; it's an approval workflow that silently doesn't exist, described in fluent English by a model that read the error and paraphrased it as progress.

The fix is that the modern pattern has the tool return an InputRequiredResult describing what it needs, and exit. Your agent surfaces that as the LangGraph interrupt, a human answers, and the MCP client re-issues the same tools/call with the answer attached. But the failure mode is the story: an era error is a fine thing to raise into a program, and a terrible thing to hand to a language model that is rewarded for sounding helpful.

3. cache=True is necessary, not sufficient​

The client cache respects the ttlMs and cacheScope hints a server attaches to its tools/list response, and only against modern-era servers that send them. A default FastMCP server sends neither — we dumped our own tools/list response and it carries no cache hints at all. So you turn the cache on, call list_tools(cache_mode="use") twice back to back, time both, and get:

Two tool discoveries of five tools each, timed at 14.6 ms and 12.5 ms, showing no cache effect

You conclude the cache is broken. It isn't — there was nothing to cache. Worth noting the cache belongs to the fastmcp.Client, not to MCPAdapter, and one client per caller keeps catalogs from crossing between tenants.

And two smaller ones, in the announcement itself​

The elicitation snippet raises AttributeError. The post reads the interrupt as paused["__interrupt__"][0].value.requests[0] — attribute access. In langchain/mcp/elicitation.py on 1.4.0:

class MCPElicitationInterrupt(TypedDict):
type: Literal["mcp_elicitation"]
tool_name: str
requests: list[MCPElicitationRequest]

A TypedDict is a dict at runtime, so .requests doesn't resolve. It's value["requests"] — which the same snippet gets right a few lines later, reading question["key"] by subscript.

The elicitation docs link 404s. The post closes the section by pointing at docs.langchain.com/oss/python/langchain/mcp/elicitation for "declining a question, and gating destructive tools behind the same approval flow." That page returns 404 as of publication, while the parent page and its other children — mcp, mcp/connections, mcp/tools, mcp/auth — all resolve. It is, inconveniently, the one page that would have documented the approval flow gap 2 shows falling over.

Why this matters for you, specifically​

If you consume MCP servers someone else runs, this is straightforwardly good news and you'll notice mostly the shorter import path.

If you write MCP servers, the July revision moved the ground and the tooling is still settling. Three things are now yours to get right: your server does not become stateless because the spec did — you set a flag; a tool that needs human input has to be written to be re-entered rather than resumed — interrupt() unwinds the whole call, so the tool body runs again from the top when you answer; and any elicitation example predating August is teaching you an API whose failure lands in the model's context instead of your logs.

That last one generalises past MCP. As frameworks get better at keeping agents alive through errors, the class of bug that ends a run is shrinking and the class that gets narrated to a user is growing. A gate that fails closed is a bug you find in testing. A gate that was never installed, described by a model as awaiting your approval, is one you find in an audit.

The direction is right. The stateless core is what makes an agent's human-approval pause survive a redeploy, and that's a real capability rather than a refactor. Just don't expect your first request to succeed.


📖 Sources: LangChain — MCP in LangChain: stateless protocol, elicitation, and more · MCP in LangChain docs · Migrating from langchain-mcp-adapters · FastMCP client documentation · FastMCP elicitation · MCP 2026-07-28 specification

Versions under test: langchain 1.4.0, fastmcp 4.0.2, mcp 2.1.1, model served over a self-hosted gateway. The ctx.elicit capture is a render of real captured output from that run, not a screen grab.

Comments & questions

Hit an error, spotted a typo, or have a question? Leave a note below.