Book appointments and refund invoices from a LangGraph agent over MCP

- 1State that survives a restart
- 2Hand work between agents
- 3Governed tools an agent can call
- 🏆Engineer the context, not the prompt
Part 1 gave an agent state that survives a restart, and a pause where a human approves before it continues. Part 2 put a supervisor in front of a scheduler and a billing agent and routed work between them.
Neither of those agents scheduled or billed anything. They discussed it. Every tool they had was a function you wrote into the same process, and the interesting ones — look up a customer, take a slot, move money — didn't exist.
This part builds the toolbox those two roles need, over MCP: a calendar, a customer list, invoices, and a refund. To keep the moving parts down we point one agent at all five tools rather than rebuilding Part 2's supervisor — the agent discovers them at runtime instead of being wired to them, chains them to answer a question, and writes rows you can go and check in the database afterwards. Then the refund stops it dead and makes a human type the amount.
Splitting these tools back across the scheduler and the billing agent — so the scheduler
cannot see issue_refund at all — is the natural next step, and the last section says how.
In July the MCP specification dropped sessions entirely. On the day we ran this, LangChain shipped support for that revision in the main package. So this is a first run against a five-week-old protocol and a same-day client — which is why three things in here aren't in anyone's documentation yet.
One WEC Instance: 8 vCPU AMD EPYC 7601, 15 GB RAM, Ubuntu 22.04, Python 3.10. langchain
1.4.0, langchain-openai 1.6.0, langgraph 1.2.11, mcp 2.1.1, fastmcp 4.0.2 — all
installed on the day of writing. Models served through the LiteLLM gateway from
Part 1 of the gateway series, which forwards to
WEC Inference. langchain.mcp is in beta and warns on every import; the API may move.
What MCP actually is
A server publishes tools. A client asks it what tools exist. The model picks one and the client calls it. That's the whole protocol — the value is that "asks what tools exist" is standardised, so any client can use any server without glue written for the pair.
Three words you need:
Tool. A function the server exposes, with a description and a JSON schema for its arguments. The description is what the model reads to decide whether to call it; the schema is what it must fill in. Nothing else about your code reaches the model.
Discovery. The client asks tools/list and gets that catalog back. Your agent learns what
it can do at runtime instead of at the time you wrote it.
Elicitation. A tool that can't finish without asking a human something — confirming a delete, supplying a parameter the model left out. Under the new spec this is an ordinary request the client retries with an answer attached.
The round trip that replaced the session
Under the old spec a tool that needed input held the connection open and asked down a
back-channel. Under 2026-07-28 it returns, and the client calls it again with the answer
attached. Two complete requests instead of one long-lived one:
The tool runs twice. That's why its body starts with if ctx.input_responses is None — it
has to be written to be re-entered, not resumed. Nothing is held open in between, so the answer
can arrive after a redeploy, or at a different replica behind a load balancer.
Prerequisites
- The virtualenv and WEC Inference access from Part 1
- A LiteLLM gateway with one virtual key, or any OpenAI-compatible endpoint you can point a model at
- Python 3.10+
- Two terminal sessions on the box — the server blocks one of them
Step 1 — Install
MCP support now lives in the main langchain package. If you have langchain-mcp-adapters
installed from before, it is the thing this replaces.
mkdir -p ~/mcp-tools && cd ~/mcp-tools && python3 -m venv .venv && ./.venv/bin/pip install -q "langchain[mcp]" langchain-openai httpx && ./.venv/bin/pip list | grep -iE 'langchain|langgraph|^mcp |fastmcp'

langchain[mcp] needs 1.4.0 or newer — below that, MCPAdapter doesn't exist.
mcp[cli] is goneOlder guides tell you to pip install "mcp[cli]". On mcp 2.x that extra no longer exists,
and pip fails with a wall of does not provide the extra 'cli' across sixty versions before
giving up with ResolutionImpossible. It reads like a dependency conflict; it's a removed
extra.
Step 2 — Something for the tools to act on
The agents need data that exists. SQLite, seeded once — three customers, six slots (one already taken), four invoices.
import sqlite3
db = sqlite3.connect("office.db")
db.executescript("""
DROP TABLE IF EXISTS customers; DROP TABLE IF EXISTS slots; DROP TABLE IF EXISTS invoices;
CREATE TABLE customers(id INTEGER PRIMARY KEY, name TEXT, email TEXT);
CREATE TABLE slots(id INTEGER PRIMARY KEY, day TEXT, time TEXT, customer_id INTEGER);
CREATE TABLE invoices(id INTEGER PRIMARY KEY, customer_id INTEGER, amount REAL, status TEXT);
INSERT INTO customers(name,email) VALUES
('Maria Alvarez','maria@example.com'),
('Tomas Reis','tomas@example.com'),
('Priya Nair','priya@example.com');
INSERT INTO slots(day,time,customer_id) VALUES
('2026-09-04','09:00',NULL),('2026-09-04','11:00',NULL),('2026-09-04','14:00',2),
('2026-09-05','09:00',NULL),('2026-09-05','11:00',NULL),('2026-09-05','16:00',NULL);
INSERT INTO invoices(customer_id,amount,status) VALUES
(1,240.00,'paid'),(1,80.00,'open'),(2,150.00,'paid'),(3,410.00,'open');
""")
db.commit()
Two details are load-bearing. Maria has both a paid and an open invoice, so "refund Maria's invoice" is ambiguous and the agent has to resolve it. Tomas already holds Friday at 14:00, so availability is a real question rather than "everything".
Step 3 — The toolbox
Five tools: three reads, one write, one that cannot finish without a human.
import os, sqlite3
from fastmcp import FastMCP, Context
from mcp.types import InputRequiredResult, ElicitRequest, ElicitRequestFormParams
DB = os.path.join(os.path.dirname(os.path.abspath(__file__)), "office.db")
mcp = FastMCP("office-tools")
def q(sql, args=()):
con = sqlite3.connect(DB); con.row_factory = sqlite3.Row
try:
rows = con.execute(sql, args).fetchall(); con.commit()
return [dict(r) for r in rows]
finally:
con.close()
def w(sql, args=()):
con = sqlite3.connect(DB)
try:
cur = con.execute(sql, args); con.commit(); return cur.rowcount
finally:
con.close()
@mcp.tool
async def find_customer(query: str) -> list[dict]:
"""Find customers whose name or email matches the query."""
like = f"%{query}%"
return q("SELECT id, name, email FROM customers WHERE name LIKE ? OR email LIKE ?", (like, like))
@mcp.tool
async def list_availability(day: str) -> list[dict]:
"""List free appointment slots on a given day, formatted YYYY-MM-DD."""
return q("SELECT id, day, time FROM slots WHERE day = ? AND customer_id IS NULL ORDER BY time", (day,))
@mcp.tool
async def book_slot(slot_id: int, customer_id: int) -> str:
"""Book a free slot for a customer. Fails if the slot is already taken."""
n = w("UPDATE slots SET customer_id = ? WHERE id = ? AND customer_id IS NULL", (customer_id, slot_id))
if n == 0:
return f"Slot {slot_id} is already taken or does not exist."
row = q("SELECT day, time FROM slots WHERE id = ?", (slot_id,))[0]
return f"Booked slot {slot_id} ({row['day']} {row['time']}) for customer {customer_id}."
@mcp.tool
async def open_invoices(customer_id: int) -> list[dict]:
"""List a customer's invoices with their amounts and status."""
return q("SELECT id, amount, status FROM invoices WHERE customer_id = ?", (customer_id,))
@mcp.tool
async def issue_refund(invoice_id: int, ctx: Context) -> str | InputRequiredResult:
"""Refund an invoice. A human must approve the amount and give a reason."""
inv = q("SELECT id, customer_id, amount, status FROM invoices WHERE id = ?", (invoice_id,))
if not inv:
return f"No invoice {invoice_id}."
inv = inv[0]
responses = ctx.input_responses
if responses is None:
params = ElicitRequestFormParams(
message=(
f"Refund invoice {inv['id']} for customer {inv['customer_id']}? "
f"It is {inv['status']}, amount {inv['amount']:.2f}. "
"Confirm the amount to refund and give a reason."
),
requested_schema={
"type": "object",
"properties": {
"amount": {"type": "number", "description": "Amount to refund"},
"reason": {"type": "string", "description": "Why this refund is being issued"},
},
"required": ["amount", "reason"],
},
)
return InputRequiredResult(
result_type="input_required",
input_requests={"approve_refund": ElicitRequest(method="elicitation/create", params=params)},
)
answer = responses["approve_refund"]
if answer.action != "accept":
return f"Refund declined - invoice {inv['id']} unchanged."
w("UPDATE invoices SET status = 'refunded' WHERE id = ?", (inv["id"],))
return f"Refunded {answer.content['amount']:.2f} on invoice {inv['id']}. Reason: {answer.content['reason']}"
if __name__ == "__main__":
mcp.run(transport="http", host="127.0.0.1", port=8770, stateless_http=True)
The docstrings are not comments. They become the tool descriptions the model reads to choose, so write them as instructions to someone who can see nothing but that one line.
book_slot refuses a taken slot in SQL rather than checking first — WHERE customer_id IS NULL
means two agents racing for the same slot can't both win.
cd ~/mcp-tools && ./.venv/bin/python office_tools.py
Step 4 — The session that shouldn't exist
Leave it running. In a second session, ask what tools it has:
curl -s -X POST http://127.0.0.1:8770/mcp -H 'Content-Type: application/json' -H 'Accept: application/json, text/event-stream' -d '{"jsonrpc":"2.0","id":1,"method":"tools/list"}'
The first time we ran this, without stateless_http=True:
{"jsonrpc":"2.0","id":null,"error":{"code":-32600,"message":"Bad Request: Missing session ID"}}

The 2026-07-28 revision removed sessions from the protocol — no handshake, no session ID, nothing pinning a client to one instance. Every write-up of the change describes the client side accurately: there is nothing left to pin.
FastMCP 4.0.2 still defaults to the stateful transport. You opt in with
stateless_http=True, and until you do, the first request against a brand-new server fails
citing a concept the spec deleted five weeks ago.
run_http_async accepts both stateless_http and stateless; the docstring says the second
is "Alias for stateless_http for CLI consistency." One switch, two names, both defaulting to
None.
The startup line is where you check it:


With the flag, the same request returns the catalog — no handshake, no session:

issue_refund's schema contains only invoice_id — the ctx: Context parameter is absent,
because FastMCP injects it and the model can't set it. That's what makes a mandatory
confirmation mandatory.
Step 5 — An agent that discovers its tools
import asyncio, os, sys
from langchain.agents import create_agent
from langchain.mcp import MCPAdapter
from langchain_openai import ChatOpenAI
model = ChatOpenAI(
base_url="http://127.0.0.1:4000/v1",
api_key=os.environ["GATEWAY_KEY"],
model="qwen-mid",
temperature=0,
)
async def main():
async with MCPAdapter("http://127.0.0.1:8770/mcp") as adapter:
tools = await adapter.list_tools()
print("discovered:", [t.name for t in tools])
agent = create_agent(model, tools)
out = await agent.ainvoke({"messages": [{"role": "user", "content": sys.argv[1]}]})
print("\n--- answer ---")
print(out["messages"][-1].content)
asyncio.run(main())
Every published MCP example points the model at Anthropic, OpenAI or Gemini. This one points at your own LiteLLM gateway, which forwards to WEC Inference — the tools are local, the model is yours, nothing leaves the box.
Keep the key in a file rather than on the command line, where it would end up in your shell history and every screenshot:
cd ~/mcp-tools && printf 'GATEWAY_KEY=%s\n' "$KEY" > .env && chmod 600 .env
cd ~/mcp-tools && set -a && source .env && set +a && ./.venv/bin/python agent.py "Who is Maria and what invoices does she have?"

That answer took two tool calls in sequence: find_customer("Maria") to get her id, then
open_invoices(1) using it. Nothing routed that — the model read five descriptions and worked
out the order.
Now one that writes:
cd ~/mcp-tools && ./.venv/bin/python agent.py "Book Maria into a free slot on 2026-09-04"

Three calls this time. And because a booking is a real change, check it somewhere the agent can't reach:
cd ~/mcp-tools && ./.venv/bin/python -c "
import sqlite3
for r in sqlite3.connect('office.db').execute('select s.id,s.day,s.time,coalesce(c.name,\"free\") from slots s left join customers c on c.id=s.customer_id order by s.day,s.time'): print(r)"

You can stand up an MCP server over a real datastore, point an agent at it, and have it chain discovered tools into a change you can verify in the database — with the model on your own gateway rather than a vendor's.
Step 6 — The tool that refuses to finish alone
Part 1's approval pause was something you put in the graph. This one comes from the tool, and the agent has no way to skip it.
import asyncio, json, os
from langchain.agents import create_agent
from langchain.mcp import MCPAdapter
from langchain_openai import ChatOpenAI
from langgraph.checkpoint.memory import InMemorySaver
from langgraph.types import Command
model = ChatOpenAI(base_url="http://127.0.0.1:4000/v1",
api_key=os.environ["GATEWAY_KEY"], model="qwen-mid", temperature=0)
def ask_human(req):
print(f"\n>>> {req['message']}")
props = (req.get("requested_schema") or {}).get("properties", {})
content = {}
for name, spec in props.items():
raw = input(f" {name} ({spec.get('type','string')}): ").strip()
content[name] = float(raw) if spec.get("type") == "number" else raw
return {"action": "accept", "content": content}
async def main():
async with MCPAdapter("http://127.0.0.1:8770/mcp") as adapter:
tools = await adapter.list_tools()
agent = create_agent(model, tools, checkpointer=InMemorySaver())
cfg = {"configurable": {"thread_id": "refund-1"}}
out = await agent.ainvoke(
{"messages": [{"role": "user", "content": "Refund Maria's open invoice."}]}, cfg)
if "__interrupt__" not in out:
print("NO INTERRUPT:", out["messages"][-1].content); return
payload = out["__interrupt__"][0].value
print("=== INTERRUPT ===")
print(json.dumps(payload, indent=2))
responses = {r["key"]: ask_human(r) for r in payload["requests"]}
print("\n=== RESUMING ===")
out = await agent.ainvoke(Command(resume={"responses": responses}), cfg)
print(out["messages"][-1].content)
asyncio.run(main())
The checkpointer and thread_id are not optional — an interrupted run needs somewhere to
wait. Same requirement as Part 1's approval pause.

Read the order of events. The model resolved "Maria's open invoice" to invoice 2 by itself — she has two, and it picked the open one. Then it stopped. The amount and the reason were typed by a human, and only then did the refund happen.
The payload is {"type": "mcp_elicitation", "tool_name": "issue_refund", "requests": [...]}.
type is the discriminator, so a graph with several kinds of interrupt can tell which came
from MCP. Each request carries a key, a message for a human, a mode of form or url,
and for a form the requested_schema the answer must satisfy — here two required fields, one
of them a number. You resume with one answer per key.
And verify, again outside the agent:

You can make a tool refuse to finish without a human — returning a typed form the model cannot fill in itself, pausing the run as a LangGraph interrupt, and resuming it with an answer a person actually typed.
ctx.elicit is the old API, and its error lands in the model's contextEvery FastMCP elicitation example you will find calls await ctx.elicit(message, response_type)
inside the tool. That is the handshake-era mechanism: it blocks mid-execution and speaks
over the session's back-channel — the back-channel the stateless transport doesn't have.
Our first attempt used it. The agent didn't pause — but not because nothing was raised. FastMCP
raised its era error exactly as documented, LangChain converted it into a ToolMessage with
status="error" (deliberately, "instead of ending the run"), and the model read that error and
narrated it as progress:

No interrupt, exit code 0, and the invoice still open. Nothing was refunded and nothing is
waiting for a human — but the user was told an approval is pending.
On the modern protocol a tool returns an InputRequiredResult describing what it needs, and
exits. The client answers and retries the whole tool call with the answer attached. Each
round is a complete request, which is why it survives load balancing and redeploys — and why
the tool body must be written to be re-entered rather than resumed. Ours does that with
if ctx.input_responses is None.
FastMCP's docs are explicit that ctx.elicit works only on connections ≤ 2025-11-25, and that
a multi-era server should branch on ctx.request_context.protocol_version. They also promise the
mismatch "raises a clear era error rather than failing obscurely" — and it does. The error just
lands in the model's context rather than yours.
LangChain's post reads the interrupt as paused["__interrupt__"][0].value.requests[0] —
attribute access. MCPElicitationInterrupt is a TypedDict (see langchain/mcp/elicitation.py
in 1.4.0), so at runtime it is a dict and .requests doesn't resolve. Use value["requests"] —
as the next line of their own snippet does, reading question["key"].
Step 7 — Caching, and why it does nothing here
Every run starts by discovering tools, which is a round trip before the model sees anything. The new spec makes that cacheable:
import asyncio, time
from fastmcp import Client
from langchain.mcp import MCPAdapter
async def main():
client = Client("http://127.0.0.1:8770/mcp", cache=True)
async with MCPAdapter(client) as adapter:
for i in (1, 2):
t = time.perf_counter()
tools = await adapter.list_tools(cache_mode="use")
print(f"discovery {i}: {len(tools)} tools in {(time.perf_counter()-t)*1000:.1f} ms")
asyncio.run(main())
discovery 1: 5 tools in 14.6 ms
discovery 2: 5 tools in 12.5 ms

No effect — and that is correct behaviour, not a broken cache.
cache=True is necessary, not sufficientThe client cache "respects the ttlMs and cacheScope hints the server attaches to each
response" and works only against modern-era servers that advertise them. Our server advertises
none, so both calls go to the network.
Note also that cache=True goes on the fastmcp.Client, not on MCPAdapter — the
adapter takes exactly one argument. And the cache belongs to the client, so one client per
caller keeps catalogs from crossing between tenants.
Over loopback there was nothing to save anyway: 12–15 ms is the round trip. Caching is for remote servers with real latency and a TTL to honour.
Troubleshooting — the errors this run actually produced
ResolutionImpossible installing mcp[cli]
The cli extra was removed in mcp 2.x. Install "langchain[mcp]" instead; it pulls a
compatible mcp and fastmcp.
Bad Request: Missing session ID
Your server is running the stateful transport. Add stateless_http=True to mcp.run(...).
KeyError: 'GATEWAY_KEY'
source .env sets shell variables, not environment variables — the Python child never sees
them. Use set -a && source .env && set +a. And export does not cross SSH sessions, so a
second terminal needs its own source.
address already in use after restarting the server
The previous instance is still holding the port — a crashed restart leaves the old process
alive, so the fix you just made appears not to work. pkill -f office_tools.py, then check
with ss -tlnp | grep 8770.
The agent never pauses on a destructive tool
The tool is calling ctx.elicit rather than returning InputRequiredResult. On a stateless
connection there is no back-channel. The era error is raised, but it arrives as a failed
ToolMessage the model reads and talks past, so the run ends normally with nothing pending.
Check the tool messages, not just the final answer.
What this did and didn't buy you
Done: an MCP server over a real datastore, an agent that discovers tools rather than being wired to them and chains them without routing, two changes you verified in the database rather than taking the model's word for, and a refund that cannot happen without a human supplying the number.
Not done:
- No auth on the MCP server. It is bound to
127.0.0.1and anyone on the box can call every tool, includingissue_refund. FastMCP supports bearer tokens and OAuth 2.1; we configured neither. - The tools trust the caller completely.
book_slotwill book for anycustomer_idit is given. There is no notion of who is asking — the same gap Part 4 of the hardening series found between a control plane and a data plane. - One server.
ClientGroupconnects several at once and prefixes tool names by server; we used a single target. - SQLite, single process. Fine for one agent; the row-level guard in
book_slotis doing more work than it looks under any real concurrency. langchain.mcpis beta and warns on every import.
What's next
Scope the tools per agent. Right now one agent holds all five, so nothing stops the model
choosing issue_refund when it was asked to book a slot — the human pause is the only guard.
Part 2's supervisor already routes between a scheduler and a billing agent; hand each of them
only its own slice of adapter.list_tools() and the scheduler cannot issue a refund because it
has never been told the tool exists. That is a stronger guarantee than asking a model nicely.
Then authentication, so a tool call carries an identity rather than trusting whoever reached the port — the same question Part 4 of the hardening series asked about the gateway's own API, one layer up.
Further reading
- MCP in LangChain — quickstart, connections, auth, elicitation
- Migrating from
langchain-mcp-adapters - FastMCP client — transports, auth, response caching
- MCP
2026-07-28specification - Our write-up of the spec revision
