Split one MCP toolbox between two agents so the scheduler cannot issue refunds

- 1State that survives a restart
- 2Hand work between agents
- 3Governed tools an agent can call
- 🏆Engineer the context, not the prompt
Part 3 gave an agent five tools over MCP —
a calendar, a customer list, invoices, and a refund — and put a human in front of the refund.
It ended by admitting the obvious: one agent held all five. Nothing stopped the model
reaching for issue_refund when it had been asked to book an appointment. The human pause was
the only thing in the way, and a pause only fires if the model calls the tool at all.
This part removes the tool instead. The scheduler gets three tools and the billing agent gets
three, sharing one. Ask the scheduler to refund an invoice and it cannot — not because it was
told not to, not because a policy blocked it, but because issue_refund was never in the list
it was given.
Then Part 2's supervisor goes back on top, and the interesting property falls out: routing decides who works, scoping decides what is possible. A misrouted refund still cannot refund.
The same box, server and database as Part 3:
one WEC Instance, Ubuntu 22.04, Python 3.10, langchain 1.4.0, langgraph 1.2.11, mcp 2.1.1,
fastmcp 4.0.2. office_tools.py unchanged and still serving five tools on
127.0.0.1:8770. Models through the LiteLLM gateway to WEC Inference. Reset the database with
./.venv/bin/python seed.py before following along, so invoice 2 is open again.
What "scoping" actually means here
There is no new mechanism in this post. create_agent takes a list of tools; we hand it a
shorter list. That is the entire technique, and it is worth being precise about why it is
stronger than the alternatives:
- A system prompt telling the model not to issue refunds is a request. It survives exactly as long as the model cooperates.
- A policy check inside the tool is real, but it runs after the model decided to call it, and it has to be written into every tool you ever add.
- Not passing the tool removes the option from the model's context. There is nothing to refuse, because there is nothing to call.
The last one is the only one that doesn't depend on inference going well.
One server, two toolboxes
The server offers all five to anyone who asks. What differs is which ones each agent is
handed. find_customer goes to both, because looking a customer up is harmless.
The blocked arrow is the whole post. The scheduler cannot call issue_refund — not because a
rule forbids it, but because the tool is not in the list it was given, so it never reaches the
model's context at all.
Prerequisites
- The MCP server, virtualenv and seeded database from Part 3
- The gateway key in
~/mcp-tools/.env, as Part 3 set up - Python 3.10+
Part 3 ran the server in the foreground, which cost a terminal and died with the SSH session. Start it detached instead — one session is enough for everything below:
cd ~/mcp-tools && nohup ./.venv/bin/python office_tools.py > server.log 2>&1 & sleep 3; ss -tlnp | grep 8770
Step 1 — Two toolsets from one server
The server still offers five tools. The roles name subsets of them:
import asyncio, os, sys
from langchain.agents import create_agent
from langchain.mcp import MCPAdapter
from langchain_openai import ChatOpenAI
from langgraph.checkpoint.memory import InMemorySaver
MCP_URL = "http://127.0.0.1:8770/mcp"
ROLES = {
"scheduler": ("find_customer", "list_availability", "book_slot"),
"billing": ("find_customer", "open_invoices", "issue_refund"),
}
model = ChatOpenAI(
base_url="http://127.0.0.1:4000/v1",
api_key=os.environ["GATEWAY_KEY"],
model="qwen-mid",
temperature=0,
)
async def main():
role, request = sys.argv[1], sys.argv[2]
async with MCPAdapter(MCP_URL) as adapter:
catalog = {t.name: t for t in await adapter.list_tools()}
print(f"server offers : {sorted(catalog)}")
allowed = [catalog[n] for n in ROLES[role]]
print(f"{role} can see : {[t.name for t in allowed]}\n")
agent = create_agent(model, allowed, checkpointer=InMemorySaver())
cfg = {"configurable": {"thread_id": f"{role}-1"}}
out = await agent.ainvoke({"messages": [{"role": "user", "content": request}]}, cfg)
if "__interrupt__" in out:
req = out["__interrupt__"][0].value["requests"][0]
print(f"PAUSED — the server is asking a human:\n {req['message']}")
return
print("--- answer ---")
print(out["messages"][-1].content)
asyncio.run(main())
The checkpointer is carried over from
Part 3: issue_refund pauses the run with an
interrupt, and an interrupted run needs somewhere to wait. The scheduler never triggers one, but
both roles run the same code.
The whole of the technique is one line:
allowed = [catalog[n] for n in ROLES[role]]
find_customer appears in both roles — looking a customer up is harmless and both jobs need
it. book_slot and issue_refund appear in exactly one each.
The script prints both lists on every run for a reason. server offers is the full catalog the
MCP server returned; can see is what we passed to create_agent. Keeping them side by side
makes it obvious that the server hid nothing — the narrowing happened in the client, on the
line above.
Step 2 — The same sentence, to both specialists
Ask the billing agent to do its job:
cd ~/mcp-tools && set -a && source .env && set +a && ./.venv/bin/python scoped.py billing "Refund Maria's open invoice."

It found Maria, resolved "open invoice" to invoice 2, called issue_refund, and stopped —
because the tool demands a human. That is Part 3's guard, working.
Now the identical sentence to the wrong specialist:
cd ~/mcp-tools && ./.venv/bin/python scoped.py scheduler "Refund Maria's open invoice."

I don't have access to tools that can process refunds or handle invoices. The available tools I have are for finding customers, listing appointment availability, and booking appointment slots.
That answer is accurate and well-mannered, and it would be a mistake to read it as the model
behaving well. The model could not have called issue_refund. It was never in the list, so
it was never in the context, so there was no call to decline.
The distinction matters because a model that does hold a dangerous tool can also produce a polite sentence — including one describing an action it never took. We documented exactly that in the news piece: an era error handed to the model, and the model narrating an approval nobody would ever be asked for. A refusal you can trust is one the model had no way around.
And the scheduler doing the job it does have tools for, so the restriction is targeted rather than crippling:
cd ~/mcp-tools && ./.venv/bin/python scoped.py scheduler "Book Maria into a free slot on 2026-09-05"

Step 3 — Check that nothing moved
A refusal is a sentence. The database is the evidence.
cd ~/mcp-tools && ./.venv/bin/python -c "
import sqlite3
for r in sqlite3.connect('office.db').execute('select id, customer_id, amount, status from invoices'): print(r)"

Invoice 2 is still open, for two different reasons: the scheduler had no tool to change it,
and billing was stopped by the human gate before it could. Neither depended on the model
choosing well.
It is worth seeing the same agent report a boring truth, too. Before we reset the database, invoice 2 had already been refunded in Part 3 — and asked to refund it again, billing looked, found nothing open, and said so:

No invented invoice, no cheerful confirmation. That is the behaviour that makes its other answers worth anything, and it is not something to take on trust — it is something to check, which is why every claim in this post has a query behind it.
You can give two agents different slices of the same MCP server, so that a capability one of them must never have is absent from its context rather than forbidden by instruction — and verify in the data that the absence held.
Step 4 — The supervisor, back on top
Part 2 routed work between a scheduler and a billing agent that had no tools. Now they do.
supervisor.py is scoped.py with Part 2's graph around it: the same ROLES, MCPAdapter and
ChatOpenAI setup, plus a State TypedDict, a StateGraph, and one agent built per role
inside the async with block — agents = {role: create_agent(model, [catalog[n] for n in names]) for role, names in ROLES.items()}. Only the parts that are new to this post are shown here.
BILLING_WORDS = ("invoice", "refund", "bill", "charge", "payment")
def supervisor(state: State) -> Command[Literal["scheduler", "billing"]]:
nxt = "billing" if any(w in state["request"].lower() for w in BILLING_WORDS) else "scheduler"
print(f"supervisor routes to {nxt}")
return Command(goto=nxt, update={"completed": [f"supervisor->{nxt}"]})
async def run(role: str, state: State):
out = await agents[role].ainvoke(
{"messages": [{"role": "user", "content": state["request"]}]})
if "__interrupt__" in out:
q = out["__interrupt__"][0].value["requests"][0]["message"]
return {"answer": f"PAUSED — {q}", "completed": [role]}
return {"answer": out["messages"][-1].content, "completed": [role]}
async def scheduler_node(state: State):
return await run("scheduler", state)
async def billing_node(state: State):
return await run("billing", state)
The Command[Literal["scheduler", "billing"]] annotation is doing the same work it did in Part
2 — it is the only declaration of where the supervisor may send work, and dropping it makes
every graph diagram lie.
cd ~/mcp-tools && set -a && source .env && set +a && ./.venv/bin/python supervisor.py "Refund Maria's open invoice."

cd ~/mcp-tools && ./.venv/bin/python supervisor.py "Book Maria into a free slot on 2026-09-04"

That supervisor is a keyword match. It is trivially foolable — "settle Maria's outstanding
amount" contains none of BILLING_WORDS and lands on the scheduler.
Which is the point worth taking away. A misrouted refund still cannot refund. The router decides who gets the work; the toolsets decide what that agent can do with it. If routing were the only control, every bug in the supervisor would be a security bug. With scoping underneath, a routing mistake is just a bad answer.
You can put a router in front of scoped specialists and reason clearly about which failures are dangerous — because the blast radius of a routing bug is bounded by what the receiving agent was given, not by what it was asked to do.
Troubleshooting — the errors this run actually produced
InvalidUpdateError: Expected dict, got <coroutine object>
The node is an async function wrapped in a lambda:
builder.add_node("scheduler", lambda s: run("scheduler", s)) # returns a coroutine
LangGraph calls it, gets a coroutine back instead of a state update, and raises. The real clue
is the last line of the output — RuntimeWarning: coroutine 'run' was never awaited — which is
the part nobody reads. Define proper async def nodes instead.
RuntimeError: Client failed to connect: All connection attempts failed
Nothing is listening on 8770. If you started the MCP server in a foreground shell, it died with your SSH session. The error names the client, which sends you looking in the wrong place.
ss -tlnp | grep 8770
Start it with nohup ... & so it survives, and check the port before blaming the code.
A backgrounded run prints nothing for minutes
Python buffers stdout when it is redirected to a file, so a long run looks identical to a hang.
Add -u:
nohup ./.venv/bin/python -u supervisor.py "..." > run.log 2>&1 &
tail -f run.log
KeyError: 'GATEWAY_KEY'
source .env sets shell variables, not environment variables, and export does not cross SSH
sessions. Every command that starts the agent needs its own
set -a && source .env && set +a — including the backgrounded ones.
What this did and didn't buy you
Done: two agents drawing different subsets of one MCP server, a capability that is absent rather than forbidden, and a router whose mistakes cost correctness instead of money.
Not done:
- The MCP server still trusts everyone. Scoping happens in the client. Anything that can
reach
127.0.0.1:8770can callissue_refunddirectly, agent or not. This is a control on what your agents can do, not on what your server will accept. - No identity on the tool call. The server cannot tell the scheduler from the billing agent, because nothing in the protocol says which is calling.
- The subsets are hardcoded. A real deployment reads them from the same place it reads team membership — which is the hardening series' territory, one layer down.
- The supervisor is a keyword match, kept deliberately dumb so the scoping argument stands on its own.
What's next
Push the boundary from the client to the server: authenticate the tool call, so the server
knows which agent is asking and refuses issue_refund to anything that isn't billing. That is
the same question Part 4 of the hardening series
asked about the gateway's own API — a control plane and a data plane, one layer apart — and
FastMCP already supports bearer tokens and OAuth 2.1 for it.
Further reading
- MCP in LangChain — connections, auth, elicitation
- FastMCP — client authentication
- LangGraph —
Commandand conditional edges - MCP
2026-07-28specification
