Skip to main content
intermediatePart 3

Book appointments and refund invoices from a LangGraph agent over MCP

· 20 min read
Rafael Fernandes
NLP Engineer & Tech Writer at WiLine
Share:
Model Context Protocol++LiteLLM
0/4
🎯 Skill path0/4 earned
Agent orchestration with LangGraph

Part 1 gave an agent state that survives a restart, and a pause where a human approves before it continues. Part 2 put a supervisor in front of a scheduler and a billing agent and routed work between them.

Neither of those agents scheduled or billed anything. They discussed it. Every tool they had was a function you wrote into the same process, and the interesting ones — look up a customer, take a slot, move money — didn't exist.

This part builds the toolbox those two roles need, over MCP: a calendar, a customer list, invoices, and a refund. To keep the moving parts down we point one agent at all five tools rather than rebuilding Part 2's supervisor — the agent discovers them at runtime instead of being wired to them, chains them to answer a question, and writes rows you can go and check in the database afterwards. Then the refund stops it dead and makes a human type the amount.

Splitting these tools back across the scheduler and the billing agent — so the scheduler cannot see issue_refund at all — is the natural next step, and the last section says how.

In July the MCP specification dropped sessions entirely. On the day we ran this, LangChain shipped support for that revision in the main package. So this is a first run against a five-week-old protocol and a same-day client — which is why three things in here aren't in anyone's documentation yet.

Reproducibility

One WEC Instance: 8 vCPU AMD EPYC 7601, 15 GB RAM, Ubuntu 22.04, Python 3.10. langchain 1.4.0, langchain-openai 1.6.0, langgraph 1.2.11, mcp 2.1.1, fastmcp 4.0.2 — all installed on the day of writing. Models served through the LiteLLM gateway from Part 1 of the gateway series, which forwards to WEC Inference. langchain.mcp is in beta and warns on every import; the API may move.

What MCP actually is​

A server publishes tools. A client asks it what tools exist. The model picks one and the client calls it. That's the whole protocol — the value is that "asks what tools exist" is standardised, so any client can use any server without glue written for the pair.

Three words you need:

Tool. A function the server exposes, with a description and a JSON schema for its arguments. The description is what the model reads to decide whether to call it; the schema is what it must fill in. Nothing else about your code reaches the model.

Discovery. The client asks tools/list and gets that catalog back. Your agent learns what it can do at runtime instead of at the time you wrote it.

Elicitation. A tool that can't finish without asking a human something — confirming a delete, supplying a parameter the model left out. Under the new spec this is an ordinary request the client retries with an answer attached.

The round trip that replaced the session​

Under the old spec a tool that needed input held the connection open and asked down a back-channel. Under 2026-07-28 it returns, and the client calls it again with the answer attached. Two complete requests instead of one long-lived one:

The tool runs twice. That's why its body starts with if ctx.input_responses is None — it has to be written to be re-entered, not resumed. Nothing is held open in between, so the answer can arrive after a redeploy, or at a different replica behind a load balancer.

Prerequisites​

  • The virtualenv and WEC Inference access from Part 1
  • A LiteLLM gateway with one virtual key, or any OpenAI-compatible endpoint you can point a model at
  • Python 3.10+
  • Two terminal sessions on the box — the server blocks one of them

Step 1 — Install​

MCP support now lives in the main langchain package. If you have langchain-mcp-adapters installed from before, it is the thing this replaces.

mkdir -p ~/mcp-tools && cd ~/mcp-tools && python3 -m venv .venv && ./.venv/bin/pip install -q "langchain[mcp]" langchain-openai httpx && ./.venv/bin/pip list | grep -iE 'langchain|langgraph|^mcp |fastmcp'

pip list showing langchain 1.4.0, langgraph 1.2.11, mcp 2.1.1 and fastmcp 4.0.2

langchain[mcp] needs 1.4.0 or newer — below that, MCPAdapter doesn't exist.

mcp[cli] is gone

Older guides tell you to pip install "mcp[cli]". On mcp 2.x that extra no longer exists, and pip fails with a wall of does not provide the extra 'cli' across sixty versions before giving up with ResolutionImpossible. It reads like a dependency conflict; it's a removed extra.

Step 2 — Something for the tools to act on​

The agents need data that exists. SQLite, seeded once — three customers, six slots (one already taken), four invoices.

~/mcp-tools/seed.py
import sqlite3
db = sqlite3.connect("office.db")
db.executescript("""
DROP TABLE IF EXISTS customers; DROP TABLE IF EXISTS slots; DROP TABLE IF EXISTS invoices;
CREATE TABLE customers(id INTEGER PRIMARY KEY, name TEXT, email TEXT);
CREATE TABLE slots(id INTEGER PRIMARY KEY, day TEXT, time TEXT, customer_id INTEGER);
CREATE TABLE invoices(id INTEGER PRIMARY KEY, customer_id INTEGER, amount REAL, status TEXT);
INSERT INTO customers(name,email) VALUES
('Maria Alvarez','maria@example.com'),
('Tomas Reis','tomas@example.com'),
('Priya Nair','priya@example.com');
INSERT INTO slots(day,time,customer_id) VALUES
('2026-09-04','09:00',NULL),('2026-09-04','11:00',NULL),('2026-09-04','14:00',2),
('2026-09-05','09:00',NULL),('2026-09-05','11:00',NULL),('2026-09-05','16:00',NULL);
INSERT INTO invoices(customer_id,amount,status) VALUES
(1,240.00,'paid'),(1,80.00,'open'),(2,150.00,'paid'),(3,410.00,'open');
""")
db.commit()

Two details are load-bearing. Maria has both a paid and an open invoice, so "refund Maria's invoice" is ambiguous and the agent has to resolve it. Tomas already holds Friday at 14:00, so availability is a real question rather than "everything".

Step 3 — The toolbox​

Five tools: three reads, one write, one that cannot finish without a human.

~/mcp-tools/office_tools.py
import os, sqlite3
from fastmcp import FastMCP, Context
from mcp.types import InputRequiredResult, ElicitRequest, ElicitRequestFormParams

DB = os.path.join(os.path.dirname(os.path.abspath(__file__)), "office.db")
mcp = FastMCP("office-tools")


def q(sql, args=()):
con = sqlite3.connect(DB); con.row_factory = sqlite3.Row
try:
rows = con.execute(sql, args).fetchall(); con.commit()
return [dict(r) for r in rows]
finally:
con.close()


def w(sql, args=()):
con = sqlite3.connect(DB)
try:
cur = con.execute(sql, args); con.commit(); return cur.rowcount
finally:
con.close()


@mcp.tool
async def find_customer(query: str) -> list[dict]:
"""Find customers whose name or email matches the query."""
like = f"%{query}%"
return q("SELECT id, name, email FROM customers WHERE name LIKE ? OR email LIKE ?", (like, like))


@mcp.tool
async def list_availability(day: str) -> list[dict]:
"""List free appointment slots on a given day, formatted YYYY-MM-DD."""
return q("SELECT id, day, time FROM slots WHERE day = ? AND customer_id IS NULL ORDER BY time", (day,))


@mcp.tool
async def book_slot(slot_id: int, customer_id: int) -> str:
"""Book a free slot for a customer. Fails if the slot is already taken."""
n = w("UPDATE slots SET customer_id = ? WHERE id = ? AND customer_id IS NULL", (customer_id, slot_id))
if n == 0:
return f"Slot {slot_id} is already taken or does not exist."
row = q("SELECT day, time FROM slots WHERE id = ?", (slot_id,))[0]
return f"Booked slot {slot_id} ({row['day']} {row['time']}) for customer {customer_id}."


@mcp.tool
async def open_invoices(customer_id: int) -> list[dict]:
"""List a customer's invoices with their amounts and status."""
return q("SELECT id, amount, status FROM invoices WHERE customer_id = ?", (customer_id,))


@mcp.tool
async def issue_refund(invoice_id: int, ctx: Context) -> str | InputRequiredResult:
"""Refund an invoice. A human must approve the amount and give a reason."""
inv = q("SELECT id, customer_id, amount, status FROM invoices WHERE id = ?", (invoice_id,))
if not inv:
return f"No invoice {invoice_id}."
inv = inv[0]

responses = ctx.input_responses
if responses is None:
params = ElicitRequestFormParams(
message=(
f"Refund invoice {inv['id']} for customer {inv['customer_id']}? "
f"It is {inv['status']}, amount {inv['amount']:.2f}. "
"Confirm the amount to refund and give a reason."
),
requested_schema={
"type": "object",
"properties": {
"amount": {"type": "number", "description": "Amount to refund"},
"reason": {"type": "string", "description": "Why this refund is being issued"},
},
"required": ["amount", "reason"],
},
)
return InputRequiredResult(
result_type="input_required",
input_requests={"approve_refund": ElicitRequest(method="elicitation/create", params=params)},
)

answer = responses["approve_refund"]
if answer.action != "accept":
return f"Refund declined - invoice {inv['id']} unchanged."

w("UPDATE invoices SET status = 'refunded' WHERE id = ?", (inv["id"],))
return f"Refunded {answer.content['amount']:.2f} on invoice {inv['id']}. Reason: {answer.content['reason']}"


if __name__ == "__main__":
mcp.run(transport="http", host="127.0.0.1", port=8770, stateless_http=True)

The docstrings are not comments. They become the tool descriptions the model reads to choose, so write them as instructions to someone who can see nothing but that one line.

book_slot refuses a taken slot in SQL rather than checking first — WHERE customer_id IS NULL means two agents racing for the same slot can't both win.

cd ~/mcp-tools && ./.venv/bin/python office_tools.py

Step 4 — The session that shouldn't exist​

Leave it running. In a second session, ask what tools it has:

curl -s -X POST http://127.0.0.1:8770/mcp -H 'Content-Type: application/json' -H 'Accept: application/json, text/event-stream' -d '{"jsonrpc":"2.0","id":1,"method":"tools/list"}'

The first time we ran this, without stateless_http=True:

{"jsonrpc":"2.0","id":null,"error":{"code":-32600,"message":"Bad Request: Missing session ID"}}

The MCP endpoint refusing a request with Bad Request: Missing session ID

The spec is stateless. Your server isn't, by default.

The 2026-07-28 revision removed sessions from the protocol — no handshake, no session ID, nothing pinning a client to one instance. Every write-up of the change describes the client side accurately: there is nothing left to pin.

FastMCP 4.0.2 still defaults to the stateful transport. You opt in with stateless_http=True, and until you do, the first request against a brand-new server fails citing a concept the spec deleted five weeks ago.

run_http_async accepts both stateless_http and stateless; the docstring says the second is "Alias for stateless_http for CLI consistency." One switch, two names, both defaulting to None.

The startup line is where you check it:

FastMCP starting without the stateless marker

FastMCP starting with transport http (stateless)

With the flag, the same request returns the catalog — no handshake, no session:

The five office tools listed with their descriptions

issue_refund's schema contains only invoice_id — the ctx: Context parameter is absent, because FastMCP injects it and the model can't set it. That's what makes a mandatory confirmation mandatory.

Step 5 — An agent that discovers its tools​

~/mcp-tools/agent.py
import asyncio, os, sys
from langchain.agents import create_agent
from langchain.mcp import MCPAdapter
from langchain_openai import ChatOpenAI

model = ChatOpenAI(
base_url="http://127.0.0.1:4000/v1",
api_key=os.environ["GATEWAY_KEY"],
model="qwen-mid",
temperature=0,
)

async def main():
async with MCPAdapter("http://127.0.0.1:8770/mcp") as adapter:
tools = await adapter.list_tools()
print("discovered:", [t.name for t in tools])
agent = create_agent(model, tools)
out = await agent.ainvoke({"messages": [{"role": "user", "content": sys.argv[1]}]})
print("\n--- answer ---")
print(out["messages"][-1].content)

asyncio.run(main())

Every published MCP example points the model at Anthropic, OpenAI or Gemini. This one points at your own LiteLLM gateway, which forwards to WEC Inference — the tools are local, the model is yours, nothing leaves the box.

Keep the key in a file rather than on the command line, where it would end up in your shell history and every screenshot:

cd ~/mcp-tools && printf 'GATEWAY_KEY=%s\n' "$KEY" > .env && chmod 600 .env
cd ~/mcp-tools && set -a && source .env && set +a && ./.venv/bin/python agent.py "Who is Maria and what invoices does she have?"

The agent chaining find_customer into open_invoices and reporting both of Maria's invoices

That answer took two tool calls in sequence: find_customer("Maria") to get her id, then open_invoices(1) using it. Nothing routed that — the model read five descriptions and worked out the order.

Now one that writes:

cd ~/mcp-tools && ./.venv/bin/python agent.py "Book Maria into a free slot on 2026-09-04"

The agent finding Maria, listing Friday's free slots, and booking 09:00

Three calls this time. And because a booking is a real change, check it somewhere the agent can't reach:

cd ~/mcp-tools && ./.venv/bin/python -c "
import sqlite3
for r in sqlite3.connect('office.db').execute('select s.id,s.day,s.time,coalesce(c.name,\"free\") from slots s left join customers c on c.id=s.customer_id order by s.day,s.time'): print(r)"

The slots table showing Maria booked into 2026-09-04 09:00

Skill unlocked 🏅

You can stand up an MCP server over a real datastore, point an agent at it, and have it chain discovered tools into a change you can verify in the database — with the model on your own gateway rather than a vendor's.

Step 6 — The tool that refuses to finish alone​

Part 1's approval pause was something you put in the graph. This one comes from the tool, and the agent has no way to skip it.

~/mcp-tools/elicit.py
import asyncio, json, os
from langchain.agents import create_agent
from langchain.mcp import MCPAdapter
from langchain_openai import ChatOpenAI
from langgraph.checkpoint.memory import InMemorySaver
from langgraph.types import Command

model = ChatOpenAI(base_url="http://127.0.0.1:4000/v1",
api_key=os.environ["GATEWAY_KEY"], model="qwen-mid", temperature=0)

def ask_human(req):
print(f"\n>>> {req['message']}")
props = (req.get("requested_schema") or {}).get("properties", {})
content = {}
for name, spec in props.items():
raw = input(f" {name} ({spec.get('type','string')}): ").strip()
content[name] = float(raw) if spec.get("type") == "number" else raw
return {"action": "accept", "content": content}

async def main():
async with MCPAdapter("http://127.0.0.1:8770/mcp") as adapter:
tools = await adapter.list_tools()
agent = create_agent(model, tools, checkpointer=InMemorySaver())
cfg = {"configurable": {"thread_id": "refund-1"}}

out = await agent.ainvoke(
{"messages": [{"role": "user", "content": "Refund Maria's open invoice."}]}, cfg)

if "__interrupt__" not in out:
print("NO INTERRUPT:", out["messages"][-1].content); return

payload = out["__interrupt__"][0].value
print("=== INTERRUPT ===")
print(json.dumps(payload, indent=2))

responses = {r["key"]: ask_human(r) for r in payload["requests"]}

print("\n=== RESUMING ===")
out = await agent.ainvoke(Command(resume={"responses": responses}), cfg)
print(out["messages"][-1].content)

asyncio.run(main())

The checkpointer and thread_id are not optional — an interrupted run needs somewhere to wait. Same requirement as Part 1's approval pause.

The interrupt payload, the typed amount and reason, and the resumed result

Read the order of events. The model resolved "Maria's open invoice" to invoice 2 by itself — she has two, and it picked the open one. Then it stopped. The amount and the reason were typed by a human, and only then did the refund happen.

The payload is {"type": "mcp_elicitation", "tool_name": "issue_refund", "requests": [...]}. type is the discriminator, so a graph with several kinds of interrupt can tell which came from MCP. Each request carries a key, a message for a human, a mode of form or url, and for a form the requested_schema the answer must satisfy — here two required fields, one of them a number. You resume with one answer per key.

And verify, again outside the agent:

The invoices table showing invoice 2 as refunded

Skill unlocked 🏅

You can make a tool refuse to finish without a human — returning a typed form the model cannot fill in itself, pausing the run as a LangGraph interrupt, and resuming it with an answer a person actually typed.

ctx.elicit is the old API, and its error lands in the model's context

Every FastMCP elicitation example you will find calls await ctx.elicit(message, response_type) inside the tool. That is the handshake-era mechanism: it blocks mid-execution and speaks over the session's back-channel — the back-channel the stateless transport doesn't have.

Our first attempt used it. The agent didn't pause — but not because nothing was raised. FastMCP raised its era error exactly as documented, LangChain converted it into a ToolMessage with status="error" (deliberately, "instead of ending the run"), and the model read that error and narrated it as progress:

The agent run printing interrupt raised? False, issue_refund with status=error carrying the era error, the model replying that it has initiated the refund and a human must approve, and sqlite3 showing invoice 2 still open

No interrupt, exit code 0, and the invoice still open. Nothing was refunded and nothing is waiting for a human — but the user was told an approval is pending.

On the modern protocol a tool returns an InputRequiredResult describing what it needs, and exits. The client answers and retries the whole tool call with the answer attached. Each round is a complete request, which is why it survives load balancing and redeploys — and why the tool body must be written to be re-entered rather than resumed. Ours does that with if ctx.input_responses is None.

FastMCP's docs are explicit that ctx.elicit works only on connections ≤ 2025-11-25, and that a multi-era server should branch on ctx.request_context.protocol_version. They also promise the mismatch "raises a clear era error rather than failing obscurely" — and it does. The error just lands in the model's context rather than yours.

The announcement's snippet doesn't run

LangChain's post reads the interrupt as paused["__interrupt__"][0].value.requests[0] — attribute access. MCPElicitationInterrupt is a TypedDict (see langchain/mcp/elicitation.py in 1.4.0), so at runtime it is a dict and .requests doesn't resolve. Use value["requests"] — as the next line of their own snippet does, reading question["key"].

Step 7 — Caching, and why it does nothing here​

Every run starts by discovering tools, which is a round trip before the model sees anything. The new spec makes that cacheable:

~/mcp-tools/cached.py
import asyncio, time
from fastmcp import Client
from langchain.mcp import MCPAdapter

async def main():
client = Client("http://127.0.0.1:8770/mcp", cache=True)
async with MCPAdapter(client) as adapter:
for i in (1, 2):
t = time.perf_counter()
tools = await adapter.list_tools(cache_mode="use")
print(f"discovery {i}: {len(tools)} tools in {(time.perf_counter()-t)*1000:.1f} ms")

asyncio.run(main())
discovery 1: 5 tools in 14.6 ms
discovery 2: 5 tools in 12.5 ms

Two discoveries of five tools timed at 14.6 ms and 12.5 ms, showing no cache effect

No effect — and that is correct behaviour, not a broken cache.

cache=True is necessary, not sufficient

The client cache "respects the ttlMs and cacheScope hints the server attaches to each response" and works only against modern-era servers that advertise them. Our server advertises none, so both calls go to the network.

Note also that cache=True goes on the fastmcp.Client, not on MCPAdapter — the adapter takes exactly one argument. And the cache belongs to the client, so one client per caller keeps catalogs from crossing between tenants.

Over loopback there was nothing to save anyway: 12–15 ms is the round trip. Caching is for remote servers with real latency and a TTL to honour.

Troubleshooting — the errors this run actually produced​

ResolutionImpossible installing mcp[cli]​

The cli extra was removed in mcp 2.x. Install "langchain[mcp]" instead; it pulls a compatible mcp and fastmcp.

Bad Request: Missing session ID​

Your server is running the stateful transport. Add stateless_http=True to mcp.run(...).

KeyError: 'GATEWAY_KEY'​

source .env sets shell variables, not environment variables — the Python child never sees them. Use set -a && source .env && set +a. And export does not cross SSH sessions, so a second terminal needs its own source.

address already in use after restarting the server​

The previous instance is still holding the port — a crashed restart leaves the old process alive, so the fix you just made appears not to work. pkill -f office_tools.py, then check with ss -tlnp | grep 8770.

The agent never pauses on a destructive tool​

The tool is calling ctx.elicit rather than returning InputRequiredResult. On a stateless connection there is no back-channel. The era error is raised, but it arrives as a failed ToolMessage the model reads and talks past, so the run ends normally with nothing pending. Check the tool messages, not just the final answer.

What this did and didn't buy you​

Done: an MCP server over a real datastore, an agent that discovers tools rather than being wired to them and chains them without routing, two changes you verified in the database rather than taking the model's word for, and a refund that cannot happen without a human supplying the number.

Not done:

  • No auth on the MCP server. It is bound to 127.0.0.1 and anyone on the box can call every tool, including issue_refund. FastMCP supports bearer tokens and OAuth 2.1; we configured neither.
  • The tools trust the caller completely. book_slot will book for any customer_id it is given. There is no notion of who is asking — the same gap Part 4 of the hardening series found between a control plane and a data plane.
  • One server. ClientGroup connects several at once and prefixes tool names by server; we used a single target.
  • SQLite, single process. Fine for one agent; the row-level guard in book_slot is doing more work than it looks under any real concurrency.
  • langchain.mcp is beta and warns on every import.
Finished this tutorial?
Mark it complete to earn Governed tools an agent can call on your skill path.

What's next​

Scope the tools per agent. Right now one agent holds all five, so nothing stops the model choosing issue_refund when it was asked to book a slot — the human pause is the only guard. Part 2's supervisor already routes between a scheduler and a billing agent; hand each of them only its own slice of adapter.list_tools() and the scheduler cannot issue a refund because it has never been told the tool exists. That is a stronger guarantee than asking a model nicely.

Then authentication, so a tool call carries an identity rather than trusting whoever reached the port — the same question Part 4 of the hardening series asked about the gateway's own API, one layer up.

Further reading​

Comments & questions

Hit an error, spotted a typo, or have a question? Leave a note below.