Skip to main content
intermediatePart 4

Split one MCP toolbox between two agents so the scheduler cannot issue refunds

· 13 min read
Rafael Fernandes
NLP Engineer & Tech Writer at WiLine
Share:
Model Context Protocol++LiteLLM
0/4
🎯 Skill path0/4 earned
Agent orchestration with LangGraph

Part 3 gave an agent five tools over MCP — a calendar, a customer list, invoices, and a refund — and put a human in front of the refund. It ended by admitting the obvious: one agent held all five. Nothing stopped the model reaching for issue_refund when it had been asked to book an appointment. The human pause was the only thing in the way, and a pause only fires if the model calls the tool at all.

This part removes the tool instead. The scheduler gets three tools and the billing agent gets three, sharing one. Ask the scheduler to refund an invoice and it cannot — not because it was told not to, not because a policy blocked it, but because issue_refund was never in the list it was given.

Then Part 2's supervisor goes back on top, and the interesting property falls out: routing decides who works, scoping decides what is possible. A misrouted refund still cannot refund.

Reproducibility

The same box, server and database as Part 3: one WEC Instance, Ubuntu 22.04, Python 3.10, langchain 1.4.0, langgraph 1.2.11, mcp 2.1.1, fastmcp 4.0.2. office_tools.py unchanged and still serving five tools on 127.0.0.1:8770. Models through the LiteLLM gateway to WEC Inference. Reset the database with ./.venv/bin/python seed.py before following along, so invoice 2 is open again.

What "scoping" actually means here​

There is no new mechanism in this post. create_agent takes a list of tools; we hand it a shorter list. That is the entire technique, and it is worth being precise about why it is stronger than the alternatives:

  • A system prompt telling the model not to issue refunds is a request. It survives exactly as long as the model cooperates.
  • A policy check inside the tool is real, but it runs after the model decided to call it, and it has to be written into every tool you ever add.
  • Not passing the tool removes the option from the model's context. There is nothing to refuse, because there is nothing to call.

The last one is the only one that doesn't depend on inference going well.

One server, two toolboxes​

The server offers all five to anyone who asks. What differs is which ones each agent is handed. find_customer goes to both, because looking a customer up is harmless.

The blocked arrow is the whole post. The scheduler cannot call issue_refund — not because a rule forbids it, but because the tool is not in the list it was given, so it never reaches the model's context at all.

Prerequisites​

  • The MCP server, virtualenv and seeded database from Part 3
  • The gateway key in ~/mcp-tools/.env, as Part 3 set up
  • Python 3.10+

Part 3 ran the server in the foreground, which cost a terminal and died with the SSH session. Start it detached instead — one session is enough for everything below:

cd ~/mcp-tools && nohup ./.venv/bin/python office_tools.py > server.log 2>&1 & sleep 3; ss -tlnp | grep 8770

Step 1 — Two toolsets from one server​

The server still offers five tools. The roles name subsets of them:

~/mcp-tools/scoped.py
import asyncio, os, sys
from langchain.agents import create_agent
from langchain.mcp import MCPAdapter
from langchain_openai import ChatOpenAI
from langgraph.checkpoint.memory import InMemorySaver

MCP_URL = "http://127.0.0.1:8770/mcp"

ROLES = {
"scheduler": ("find_customer", "list_availability", "book_slot"),
"billing": ("find_customer", "open_invoices", "issue_refund"),
}

model = ChatOpenAI(
base_url="http://127.0.0.1:4000/v1",
api_key=os.environ["GATEWAY_KEY"],
model="qwen-mid",
temperature=0,
)

async def main():
role, request = sys.argv[1], sys.argv[2]
async with MCPAdapter(MCP_URL) as adapter:
catalog = {t.name: t for t in await adapter.list_tools()}
print(f"server offers : {sorted(catalog)}")

allowed = [catalog[n] for n in ROLES[role]]
print(f"{role} can see : {[t.name for t in allowed]}\n")

agent = create_agent(model, allowed, checkpointer=InMemorySaver())
cfg = {"configurable": {"thread_id": f"{role}-1"}}
out = await agent.ainvoke({"messages": [{"role": "user", "content": request}]}, cfg)

if "__interrupt__" in out:
req = out["__interrupt__"][0].value["requests"][0]
print(f"PAUSED — the server is asking a human:\n {req['message']}")
return

print("--- answer ---")
print(out["messages"][-1].content)

asyncio.run(main())

The checkpointer is carried over from Part 3: issue_refund pauses the run with an interrupt, and an interrupted run needs somewhere to wait. The scheduler never triggers one, but both roles run the same code.

The whole of the technique is one line:

allowed = [catalog[n] for n in ROLES[role]]

find_customer appears in both roles — looking a customer up is harmless and both jobs need it. book_slot and issue_refund appear in exactly one each.

The script prints both lists on every run for a reason. server offers is the full catalog the MCP server returned; can see is what we passed to create_agent. Keeping them side by side makes it obvious that the server hid nothing — the narrowing happened in the client, on the line above.

Step 2 — The same sentence, to both specialists​

Ask the billing agent to do its job:

cd ~/mcp-tools && set -a && source .env && set +a && ./.venv/bin/python scoped.py billing "Refund Maria's open invoice."

The billing agent holding find_customer, open_invoices and issue_refund, pausing with the server's confirmation question about invoice 2

It found Maria, resolved "open invoice" to invoice 2, called issue_refund, and stopped — because the tool demands a human. That is Part 3's guard, working.

Now the identical sentence to the wrong specialist:

cd ~/mcp-tools && ./.venv/bin/python scoped.py scheduler "Refund Maria's open invoice."

The scheduler agent holding find_customer, list_availability and book_slot, replying that it has no tools for refunds or invoices

I don't have access to tools that can process refunds or handle invoices. The available tools I have are for finding customers, listing appointment availability, and booking appointment slots.

The politeness is not the guarantee

That answer is accurate and well-mannered, and it would be a mistake to read it as the model behaving well. The model could not have called issue_refund. It was never in the list, so it was never in the context, so there was no call to decline.

The distinction matters because a model that does hold a dangerous tool can also produce a polite sentence — including one describing an action it never took. We documented exactly that in the news piece: an era error handed to the model, and the model narrating an approval nobody would ever be asked for. A refusal you can trust is one the model had no way around.

And the scheduler doing the job it does have tools for, so the restriction is targeted rather than crippling:

cd ~/mcp-tools && ./.venv/bin/python scoped.py scheduler "Book Maria into a free slot on 2026-09-05"

The scheduler agent booking Maria into the 09:00 slot on 2026-09-05

Step 3 — Check that nothing moved​

A refusal is a sentence. The database is the evidence.

cd ~/mcp-tools && ./.venv/bin/python -c "
import sqlite3
for r in sqlite3.connect('office.db').execute('select id, customer_id, amount, status from invoices'): print(r)"

The invoices table showing invoice 2 still open after both the refusal and the paused run

Invoice 2 is still open, for two different reasons: the scheduler had no tool to change it, and billing was stopped by the human gate before it could. Neither depended on the model choosing well.

It is worth seeing the same agent report a boring truth, too. Before we reset the database, invoice 2 had already been refunded in Part 3 — and asked to refund it again, billing looked, found nothing open, and said so:

The billing agent reporting that Maria has no open invoices, her invoices being either paid or already refunded

No invented invoice, no cheerful confirmation. That is the behaviour that makes its other answers worth anything, and it is not something to take on trust — it is something to check, which is why every claim in this post has a query behind it.

Skill unlocked 🏅

You can give two agents different slices of the same MCP server, so that a capability one of them must never have is absent from its context rather than forbidden by instruction — and verify in the data that the absence held.

Step 4 — The supervisor, back on top​

Part 2 routed work between a scheduler and a billing agent that had no tools. Now they do.

supervisor.py is scoped.py with Part 2's graph around it: the same ROLES, MCPAdapter and ChatOpenAI setup, plus a State TypedDict, a StateGraph, and one agent built per role inside the async with block — agents = {role: create_agent(model, [catalog[n] for n in names]) for role, names in ROLES.items()}. Only the parts that are new to this post are shown here.

~/mcp-tools/supervisor.py (the parts that matter)
BILLING_WORDS = ("invoice", "refund", "bill", "charge", "payment")

def supervisor(state: State) -> Command[Literal["scheduler", "billing"]]:
nxt = "billing" if any(w in state["request"].lower() for w in BILLING_WORDS) else "scheduler"
print(f"supervisor routes to {nxt}")
return Command(goto=nxt, update={"completed": [f"supervisor->{nxt}"]})

async def run(role: str, state: State):
out = await agents[role].ainvoke(
{"messages": [{"role": "user", "content": state["request"]}]})
if "__interrupt__" in out:
q = out["__interrupt__"][0].value["requests"][0]["message"]
return {"answer": f"PAUSED — {q}", "completed": [role]}
return {"answer": out["messages"][-1].content, "completed": [role]}

async def scheduler_node(state: State):
return await run("scheduler", state)

async def billing_node(state: State):
return await run("billing", state)

The Command[Literal["scheduler", "billing"]] annotation is doing the same work it did in Part 2 — it is the only declaration of where the supervisor may send work, and dropping it makes every graph diagram lie.

cd ~/mcp-tools && set -a && source .env && set +a && ./.venv/bin/python supervisor.py "Refund Maria's open invoice."

The supervisor routing a refund request to the billing agent, which pauses on the human gate

cd ~/mcp-tools && ./.venv/bin/python supervisor.py "Book Maria into a free slot on 2026-09-04"

The supervisor routing a booking request to the scheduler agent, which books the slot

Routing is not the guarantee

That supervisor is a keyword match. It is trivially foolable — "settle Maria's outstanding amount" contains none of BILLING_WORDS and lands on the scheduler.

Which is the point worth taking away. A misrouted refund still cannot refund. The router decides who gets the work; the toolsets decide what that agent can do with it. If routing were the only control, every bug in the supervisor would be a security bug. With scoping underneath, a routing mistake is just a bad answer.

Skill unlocked 🏅

You can put a router in front of scoped specialists and reason clearly about which failures are dangerous — because the blast radius of a routing bug is bounded by what the receiving agent was given, not by what it was asked to do.

Troubleshooting — the errors this run actually produced​

InvalidUpdateError: Expected dict, got <coroutine object>​

The node is an async function wrapped in a lambda:

builder.add_node("scheduler", lambda s: run("scheduler", s)) # returns a coroutine

LangGraph calls it, gets a coroutine back instead of a state update, and raises. The real clue is the last line of the output — RuntimeWarning: coroutine 'run' was never awaited — which is the part nobody reads. Define proper async def nodes instead.

RuntimeError: Client failed to connect: All connection attempts failed​

Nothing is listening on 8770. If you started the MCP server in a foreground shell, it died with your SSH session. The error names the client, which sends you looking in the wrong place.

ss -tlnp | grep 8770

Start it with nohup ... & so it survives, and check the port before blaming the code.

A backgrounded run prints nothing for minutes​

Python buffers stdout when it is redirected to a file, so a long run looks identical to a hang. Add -u:

nohup ./.venv/bin/python -u supervisor.py "..." > run.log 2>&1 &
tail -f run.log

KeyError: 'GATEWAY_KEY'​

source .env sets shell variables, not environment variables, and export does not cross SSH sessions. Every command that starts the agent needs its own set -a && source .env && set +a — including the backgrounded ones.

What this did and didn't buy you​

Done: two agents drawing different subsets of one MCP server, a capability that is absent rather than forbidden, and a router whose mistakes cost correctness instead of money.

Not done:

  • The MCP server still trusts everyone. Scoping happens in the client. Anything that can reach 127.0.0.1:8770 can call issue_refund directly, agent or not. This is a control on what your agents can do, not on what your server will accept.
  • No identity on the tool call. The server cannot tell the scheduler from the billing agent, because nothing in the protocol says which is calling.
  • The subsets are hardcoded. A real deployment reads them from the same place it reads team membership — which is the hardening series' territory, one layer down.
  • The supervisor is a keyword match, kept deliberately dumb so the scoping argument stands on its own.
Finished this tutorial?
Mark it complete to earn Engineer the context, not the prompt on your skill path.

What's next​

Push the boundary from the client to the server: authenticate the tool call, so the server knows which agent is asking and refuses issue_refund to anything that isn't billing. That is the same question Part 4 of the hardening series asked about the gateway's own API — a control plane and a data plane, one layer apart — and FastMCP already supports bearer tokens and OAuth 2.1 for it.

Further reading​

Comments & questions

Hit an error, spotted a typo, or have a question? Leave a note below.