Skip to main content

14 posts tagged with "agents"

View all tags
intermediatePart 10

Give each MCP tool its own scope, and return a refusal the client can act on

· 12 min read
Rafael Fernandes
NLP Engineer & Tech Writer at WiLine
Share:
authentik+Model Context Protocol

Part 6 put required_scopes=["mcp:invoke"] on the JWT verifier and called the server authorised. It is, in the sense that an unauthenticated caller gets nothing. It is not, in the sense that matters: mcp:invoke opens every tool on the server. The same token that lets an agent run find_customer lets it run issue_refund.

That is the whole gap. A read-only agent and a refund-issuing agent hold identical credentials, and the only thing standing between "look up Maria" and "refund invoice 2" is that nobody asked.

This post closes it, badly first and then properly, because the badly is what most people ship and the difference only shows up in what the client can do about a refusal.

Both versions refuse the same call — a read-only token asking for issue_refund. They differ in where the scope is checked, and that one choice decides what the client gets back:

Prerequisites​

Parts 5 and 6, specifically:

  • Authentik issuing client-credentials tokens to an agent-tools provider (part 5)
  • A FastMCP server verifying those tokens against the JWKS endpoint (part 6)
  • ~/mcp-auth/token.sh, which takes a scope string and returns an access token

Part 6's server stays on :8770 throughout. The two versions below run on :8771 and :8772 so you can compare all three without stopping anything.

Step 1 — Two new scopes​

Scopes are Property Mappings in Authentik. Customization → Property Mappings → New Property Mapping → Scope Mapping, twice:

NameScope nameDescription
office-readoffice:readLook up customers and invoices
office-refundoffice:refundIssue refunds

Two new scope mappings alongside the existing gateway-invoke and mcp-invoke

Creating them is not enough. A provider will only issue a scope it has been given, so open Applications → Providers → agent-tools → Edit and move both into Selected Scopes.

The agent-tools provider with mcp-invoke, office-read and office-refund selected

Note the line under the picker: "Select which scopes can be used by the client. The client still has to specify the scope to access the data." Both halves matter. Selecting a scope here does not put it in every token — it permits the client to ask. A client that asks for nothing gets nothing, which is why every token.sh call below passes an explicit scope string.

Confirm the token actually carries what you asked for:

~/mcp-auth/token.sh "mcp:invoke office:read" \
| cut -d. -f2 | base64 -d 2>/dev/null | jq .scope
"mcp:invoke office:read"

If that comes back without office:read, the scope is not on the provider — fix that before writing any server code, or you will spend an hour debugging enforcement that is working correctly on a token that was never scoped.

Step 2 — The obvious implementation​

Keep required_scopes=["mcp:invoke"] on the verifier as the price of admission, then ask per tool whether the caller holds what that operation needs. FastMCP exposes the verified token through get_access_token(), so a decorator can read the claims the verifier already checked:

from fastmcp.exceptions import ToolError
from fastmcp.server.dependencies import get_access_token

def requires(scope: str):
def decorate(fn):
@functools.wraps(fn)
async def wrapper(*args, **kwargs):
token = get_access_token()
held = set(getattr(token, "scopes", None) or [])
if scope not in held:
raise ToolError(
f'insufficient_scope: this tool requires "{scope}"; '
f'token carries {sorted(held) or "nothing"}'
)
return await fn(*args, **kwargs)
return wrapper
return decorate

Then one line per tool:

@mcp.tool
@requires("office:read")
async def find_customer(query: str) -> list[dict]: ...

@mcp.tool
@requires("office:refund")
async def issue_refund(invoice_id: int) -> str: ...

Run it on :8771 and drive both tools with both tokens:

Two tokens, two tools: each token allows one and refuses the other

--- token: mcp:invoke + office:read ---
find_customer : ALLOWED -> [{'id': 1, 'name': 'Maria Alvarez', ...}]
issue_refund : REFUSED -> insufficient_scope: this tool requires "office:refund"

--- token: mcp:invoke + office:refund ---
find_customer : REFUSED -> insufficient_scope: this tool requires "office:read"
issue_refund : ALLOWED -> Refunded 80.00 on invoice 2.

That is real enforcement. The refund token cannot read, the read token cannot refund, and the refusal names the missing scope. For a lot of deployments this is where you stop.

Step 3 — Why that refusal is not good enough​

Watch the HTTP layer rather than the client library.

TOK=$(~/mcp-auth/token.sh "mcp:invoke office:read")
curl -sS -i -X POST http://127.0.0.1:8771/mcp \
-H "Authorization: Bearer $TOK" \
-H "Content-Type: application/json" \
-H "Accept: application/json, text/event-stream" \
-d '{"jsonrpc":"2.0","id":1,"method":"tools/call","params":{"name":"issue_refund","arguments":{"invoice_id":2}}}'

The tool-body check returns HTTP 200 with the refusal inside the body

HTTP/1.1 200 OK
content-type: text/event-stream

data: {"jsonrpc":"2.0","id":1,"result":{"content":[{"text":"insufficient_scope: this tool
requires \"office:refund\"; token carries ['mcp:invoke', 'office:read']","type":"text"},
"isError":true}}

HTTP 200. The request succeeded; the tool declined. That is correct JSON-RPC semantics and it is useless to an OAuth client.

An OAuth client that wants to step up — go back to the authorization server and ask for office:refund — is watching for a 401 or 403 with a WWW-Authenticate header telling it what to request. It does not parse English out of a tool result. So the refusal is legible to a human reading logs and invisible to the machinery designed to handle exactly this case.

It also arrives late. The tool body runs after routing, after session setup, after the server has committed to a successful response. By then there is no status line left to change.

Step 4 — The challenge the spec actually wants​

Getting a 403 means checking scopes before the JSON-RPC layer answers, which means ASGI middleware wrapping the FastMCP app:

class ScopeChallengeMiddleware:
def __init__(self, app):
self.app = app

async def __call__(self, scope, receive, send):
if scope["type"] != "http" or scope["method"] != "POST":
return await self.app(scope, receive, send)

# Buffer the body so we can inspect it and still pass it on.
chunks, more = [], True
while more:
msg = await receive()
chunks.append(msg.get("body", b""))
more = msg.get("more_body", False)
body = b"".join(chunks)

needed = None
try:
rpc = json.loads(body)
if rpc.get("method") == "tools/call":
needed = TOOL_SCOPES.get((rpc.get("params") or {}).get("name"))
except Exception:
pass

if needed and needed not in token_scopes(dict(scope.get("headers") or [])):
challenge = (
f'Bearer error="insufficient_scope", scope="{needed}", '
f'resource_metadata="{RESOURCE_METADATA}", '
f'error_description="This operation requires the {needed} scope"'
)
await send({"type": "http.response.start", "status": 403,
"headers": [(b"content-type", b"application/json"),
(b"www-authenticate", challenge.encode())]})
await send({"type": "http.response.body",
"body": json.dumps({"error": "insufficient_scope",
"scope": needed}).encode()})
return

# Replay the buffered body downstream.
replayed = False
async def replay():
nonlocal replayed
if not replayed:
replayed = True
return {"type": "http.request", "body": body, "more_body": False}
return await receive()

await self.app(scope, replay, send)

app = ScopeChallengeMiddleware(mcp.http_app(stateless_http=True))

Same request, against :8772:

The middleware returns 403 with a WWW-Authenticate challenge naming the missing scope

HTTP/1.1 403 Forbidden
www-authenticate: Bearer error="insufficient_scope", scope="office:refund",
resource_metadata="http://127.0.0.1:8772/.well-known/oauth-protected-resource",
error_description="This operation requires the office:refund scope"

Now a client has something to act on: the status says refused, scope= says what to ask for, and resource_metadata says where to look up how. That is step-up authorization as a protocol rather than as a log message.

Step 5 — What it cost​

The middleware works. It is also worse code than the decorator, in three specific ways, and pretending otherwise would be dishonest.

It does not know what a tool is. ASGI middleware sees bytes and headers. To find out which tool is being called it parses the JSON-RPC envelope itself and looks the name up in a table it has to maintain:

TOOL_SCOPES = {
"find_customer": "office:read",
"open_invoices": "office:read",
"issue_refund": "office:refund",
}

That table is a second source of truth. Add a tool and forget the entry and it is unprotected — silently, because the middleware just passes through anything it does not recognise. The decorator could not have that bug: the requirement sat on the function.

It decodes the token unverified. The middleware runs upstream of the verifier, so the verified claims do not exist yet. It splits the JWT and base64-decodes the payload without checking the signature:

payload = auth.split(None, 1)[1].split(".")[1]
claims = json.loads(base64.urlsafe_b64decode(payload + "=" * (-len(payload) % 4)))

This is not the hole it looks like — the verifier still runs downstream and still rejects a forged token, so nothing reaches a tool on a bad signature. But the scope decision is made on unauthenticated bytes, and the only reason that is survivable is the second check behind it. It is a wart, not a vulnerability, and it is the kind of thing worth writing down before someone later removes the "redundant" verifier.

It buffers every request body. To read the envelope and still pass it downstream, the middleware drains receive() into memory and replays it. Fine for JSON-RPC calls; think harder before putting this in front of large uploads.

Troubleshooting — the errors this run actually produced​

insufficient_scope on a tool you did grant​

Check the token, not the server:

~/mcp-auth/token.sh "mcp:invoke office:refund" | cut -d. -f2 | base64 -d 2>/dev/null | jq .scope

If office:refund is missing, the scope exists as a Property Mapping but was never moved into the provider's Selected Scopes. Authentik silently drops scopes a client is not permitted to request rather than erroring, so the token comes back valid and short.

The middleware never fires​

It only inspects POST. MCP clients open a GET for the event stream first, and that request carries no JSON-RPC envelope — if you are watching the wrong request you will conclude the middleware is dead. Confirm with the curl above, which is a single POST.

Every tool suddenly returns 403​

TOOL_SCOPES is keyed by the tool's registered name, which is the function name, not the decorated label. Rename a function and the table stops matching. The pass-through case is the dangerous direction (unprotected), but a typo in the table produces this one.

The 403 body is empty in some clients​

The challenge lives in the WWW-Authenticate header. Clients that only log response bodies will show you {"error":"insufficient_scope"} and nothing about which scope. Use curl -i.

What this did and didn't buy you​

It bought genuine per-tool authorisation: two tokens that differ by one scope, each able to run exactly one of two tools, proven at the HTTP layer rather than asserted. And with the middleware, a refusal a client can programmatically recover from.

It did not buy a clean design. Both implementations are compromises pointing opposite ways:

tool decorator (:8771)ASGI middleware (:8772)
Scope requirement liveson the functionin a separate table
Reads verified claimsyesno — decodes unverified, verifier runs after
RefusalHTTP 200, isError: trueHTTP 403 + WWW-Authenticate
Client can step upnoyes
New tool unprotected by defaultnoyes

The honest summary is that the correct protocol behaviour requires leaving the abstraction the framework gives you, and the ergonomic version cannot produce it. If your clients do not implement step-up — and most agent clients today do not — the decorator is the better trade. If they do, you pay for it with a table you must remember to update.

It also did not buy authorisation that survives the tool doing something else. issue_refund is gated on office:refund; nothing stops a future find_customer from being edited to write. Scopes gate entry, not behaviour — which is the same boundary part 4 drew around policy-in-the-tool.

What's next​

The obvious remaining gap is that both versions trust the token's scope list and nothing else. Neither asks who the caller is or what they are acting on — a token with office:refund refunds any invoice, for any customer, for any amount. That is object-level authorisation, and it does not live in OAuth scopes at all.

Finished this tutorial?
Mark it complete to earn A scope per tool, and a refusal clients can act on on your skill path.

Further reading​

intermediatePart 9

Sandbox the code your agent writes, and prove every limit actually applied

· 11 min read
Rafael Fernandes
NLP Engineer & Tech Writer at WiLine
Share:
+

Every agent tutorial in this series so far has given a model tools — functions you wrote, with arguments you defined. This post is about the other thing agents do, which is write code and then run it.

That is a different risk, and the difference is worth being precise about. A tool call is the model choosing from a menu you control. Executing generated code is the model handing you something nobody has ever read, which you then run on your machine.

intermediatePart 8

Move the Docker daemon off root so a container escape lands on an ordinary user

· 14 min read
Rafael Fernandes
NLP Engineer & Tech Writer at WiLine
Share:
+

Here is the part nobody says out loud when they tell you to add yourself to the docker group so you can stop typing sudo.

That group is root. Not "close to root", not "root for Docker things". If you can run a container, you can read, change or delete any file on the machine — including the password file, including other people's home directories, including the files the administrator deliberately kept away from you.

No exploit required. It is one command, and it takes about four seconds.

intermediatePart 7

Firewall an agent container so it can reach one API and nothing else

· 21 min read
Rafael Fernandes
NLP Engineer & Tech Writer at WiLine
Share:
+

Agent orchestration part 4 split a toolbox so the scheduler could not issue refunds. Part 5 gave the agent an identity at the gateway. Part 6 put a JWT verifier in front of the tools server so it refuses anonymous callers.

Every one of those controls what the agent may call. Not one of them controls where the agent may go.

That distinction is the whole of this post. Exfiltration does not need a tool. It needs a socket.

intermediatePart 6

The wrong valid token: authenticating an MCP tools server with authentik

· 16 min read
Rafael Fernandes
NLP Engineer & Tech Writer at WiLine
Share:
authentik+Model Context Protocol

Part 5 gave an agent its own identity at the gateway: a token issued by authentik over client credentials, carrying a scope, expiring in five minutes. It ended by naming what it had not covered — the MCP tools server from agent orchestration part 4 still trusts anything that can reach its port, issue_refund included.

That post was explicit about the limit of what it had built:

The MCP server still trusts everyone. Scoping happens in the client. Anything that can reach 127.0.0.1:8770 can call issue_refund directly, agent or not.

Splitting the toolbox per role stopped an agent from reaching a tool it shouldn't. It did nothing about a curl. This part closes that.

intermediatePart 5

The key that expires: giving an agent its own identity at the gateway

· 19 min read
Rafael Fernandes
NLP Engineer & Tech Writer at WiLine
Share:
authentik+LiteLLM

Part 4 put group access control on the gateway's admin UI and then, at the end, called the API from a shell with no account, no session and no group. It answered normally. The conclusion was that SSO protects a control plane and the data plane authenticates machine callers with keys — which it has to, because an agent running at 3am cannot complete a browser login.

That was true and it was also a stopping point rather than an answer. "Machine callers use keys" leaves you with a credential that never expires, that no identity provider knows about, and that survives the person who created it. Part 4 said so plainly: removing someone from a group does not revoke their keys, a leaked key is unaffected by identity entirely, and keys outlive people.

This part gives the machine an identity instead of a key.

authentik issues the agent a token over the client credentials grant — no browser, no consent screen, no human. The token is signed, carries a scope, and expires in five minutes. The gateway verifies it locally against authentik's public keys and refuses anything without the right scope. At the end, three status codes show the boundary holding.

intermediatePart 4

Split one MCP toolbox between two agents so the scheduler cannot issue refunds

· 13 min read
Rafael Fernandes
NLP Engineer & Tech Writer at WiLine
Share:
Model Context Protocol++LiteLLM
0/4
🎯 Skill path0/4 earned
Agent orchestration with LangGraph

Part 3 gave an agent five tools over MCP — a calendar, a customer list, invoices, and a refund — and put a human in front of the refund. It ended by admitting the obvious: one agent held all five. Nothing stopped the model reaching for issue_refund when it had been asked to book an appointment. The human pause was the only thing in the way, and a pause only fires if the model calls the tool at all.

This part removes the tool instead. The scheduler gets three tools and the billing agent gets three, sharing one. Ask the scheduler to refund an invoice and it cannot — not because it was told not to, not because a policy blocked it, but because issue_refund was never in the list it was given.

Then Part 2's supervisor goes back on top, and the interesting property falls out: routing decides who works, scoping decides what is possible. A misrouted refund still cannot refund.

intermediatePart 3

Book appointments and refund invoices from a LangGraph agent over MCP

· 20 min read
Rafael Fernandes
NLP Engineer & Tech Writer at WiLine
Share:
Model Context Protocol++LiteLLM
0/4
🎯 Skill path0/4 earned
Agent orchestration with LangGraph

Part 1 gave an agent state that survives a restart, and a pause where a human approves before it continues. Part 2 put a supervisor in front of a scheduler and a billing agent and routed work between them.

Neither of those agents scheduled or billed anything. They discussed it. Every tool they had was a function you wrote into the same process, and the interesting ones — look up a customer, take a slot, move money — didn't exist.

This part builds the toolbox those two roles need, over MCP: a calendar, a customer list, invoices, and a refund. To keep the moving parts down we point one agent at all five tools rather than rebuilding Part 2's supervisor — the agent discovers them at runtime instead of being wired to them, chains them to answer a question, and writes rows you can go and check in the database afterwards. Then the refund stops it dead and makes a human type the amount.

Splitting these tools back across the scheduler and the billing agent — so the scheduler cannot see issue_refund at all — is the natural next step, and the last section says how.

In July the MCP specification dropped sessions entirely. On the day we ran this, LangChain shipped support for that revision in the main package. So this is a first run against a five-week-old protocol and a same-day client — which is why three things in here aren't in anyone's documentation yet.

intermediatePart 2

Hand work between LangGraph agents without corrupting shared state

· 16 min read
Rafael Fernandes
NLP Engineer & Tech Writer at WiLine
Share:
+PostgreSQL+
0/4
🎯 Skill path0/4 earned
Agent orchestration with LangGraph

You have two agents. One books appointments. One handles billing.

A ticket arrives that needs both: reschedule my install, and my bill looks wrong. So you send it to both.

The first one books Tuesday. The second one sees the account is past due and freezes it. Both finish at almost the same moment, and both save what they decided.

Only one of them is saved. Which one? Whichever finished first — which comes down to how slow an API call was that day. So you book an appointment on a frozen account, or freeze an account you just promised an engineer to. Afterwards it looks like one clean decision was made.

Part 1 built one agent that runs its steps in a fixed order. This post has several: a supervisor that picks who works on what, an agent that passes the job to another one halfway through, and two agents put in each other's way on purpose — to find out what LangGraph does when they disagree.

intermediatePart 1

Checkpoint a LangGraph agent on a WEC Instance so crashes cost you nothing

· 20 min read
Rafael Fernandes
NLP Engineer & Tech Writer at WiLine
Share:
+PostgreSQL+
0/4
🎯 Skill path0/4 earned
Agent orchestration with LangGraph

Your agent has been working on a customer's request for forty seconds. It has read the ticket, pulled the account, and checked three engineers' calendars. It is about to book the appointment.

Then the process dies. A deploy going out, the host running out of memory, someone restarting the container — it does not matter which.

The agent does not pick up where it left off, because there is nothing to pick up from. Everything it learned lived in variables inside a process that no longer exists. The customer is still waiting. Run it again and you pay for all that work a second time. And the part that should worry you most: nobody can say whether the appointment was booked in the last second before it died.

An agent is a model in a loop. As an ordinary Python script, that loop is exactly as fragile as the process holding it.

Here we build one that saves its state to Postgres after every step, so a crash costs nothing.

intermediatePart 2

Root by default: hardening container privilege on a self-hosted AI stack

· 8 min read
Rafael Fernandes
NLP Engineer & Tech Writer at WiLine
Share:
+

Part 1 of this series audited a live multi-agent box and found every network-exposure question worth asking. One line from that audit didn't get followed up: "the Postgres, ClickHouse and Redis behind Langfuse were published to 127.0.0.1 instead of the world — someone made a good decision there. Hold that thought."

Here's the other half of that thought: getting the network right says nothing about what happens after someone's already inside a container. If the process running in there is root, a compromise starts with the keys to the whole filesystem. So we checked — on the same box, the same Langfuse stack Part 1 already praised — whether "network correct" also meant "privilege correct." It didn't, for two of the six containers. Here's what fixing that actually looked like, including the part that broke.

intermediatePart 1

Your firewall is lying to you: hardening Docker networks for multi-agent systems

· 15 min read
Rafael Fernandes
NLP Engineer & Tech Writer at WiLine
Share:
+

You start with one agent. Then it needs a database. Then you add a second agent, a messaging bridge, an observability stack. Six months later a single WEC Instance is running five compose projects, twenty-something containers, and nobody remembers which ports are open to the world.

That's not a hypothetical — that's the box this tutorial was written on. So instead of theorizing, we probed it: can containers reach each other across stacks? Can they reach the databases? Is the firewall actually protecting anything?

Three of the answers surprised me. One of them was a database sitting there with no password. And the firewall — the firewall was lying.

intermediatePart 7

Component-level tracing: debugging agent tool calls

· 12 min read
Rafael Fernandes
NLP Engineer & Tech Writer at WiLine
Share:
+
0/8
🎯 Skill path0/8 earned
AI evals & observability

An agent that calls tools does two very different things: it reasons ("I should search the docs") and it acts (actually calls the tool). When something goes wrong, the first question is always which layer failed — did it think wrong, or did the doing break? A flat log can't answer that. A trace can.

In this tutorial you build a small tool-calling agent on WEC Inference, instrument it with Langfuse so every reasoning step and every tool call becomes an inspectable node, then debug two real failures from the trace tree — including the worst kind: a confident, wrong answer that never throws an error. Every command, error, and screenshot below is from a real run.

intermediatePart 1

Self-host the Hermes Agent with persistent memory

· 14 min read
Rafael Fernandes
NLP Engineer & Tech Writer at WiLine
Share:
Hermes Agent+SQLite
0/2
🎯 Skill path0/2 earned
Self-hosting Hermes

Hermes Agent is Nous Research's open-source (MIT) AI agent — "the agent that grows with you." Its standout feature is persistent memory: it learns your projects and doesn't forget across restarts. This guide deploys it on the same WEC Instance you already use for OpenClaw, points it at a model, and proves the memory survives a full reboot.