Skip to main content

10 posts tagged with "hardening"

View all tags
intermediatePart 10

Give each MCP tool its own scope, and return a refusal the client can act on

· 12 min read
Rafael Fernandes
NLP Engineer & Tech Writer at WiLine
Share:
authentik+Model Context Protocol

Part 6 put required_scopes=["mcp:invoke"] on the JWT verifier and called the server authorised. It is, in the sense that an unauthenticated caller gets nothing. It is not, in the sense that matters: mcp:invoke opens every tool on the server. The same token that lets an agent run find_customer lets it run issue_refund.

That is the whole gap. A read-only agent and a refund-issuing agent hold identical credentials, and the only thing standing between "look up Maria" and "refund invoice 2" is that nobody asked.

This post closes it, badly first and then properly, because the badly is what most people ship and the difference only shows up in what the client can do about a refusal.

Both versions refuse the same call — a read-only token asking for issue_refund. They differ in where the scope is checked, and that one choice decides what the client gets back:

Prerequisites​

Parts 5 and 6, specifically:

  • Authentik issuing client-credentials tokens to an agent-tools provider (part 5)
  • A FastMCP server verifying those tokens against the JWKS endpoint (part 6)
  • ~/mcp-auth/token.sh, which takes a scope string and returns an access token

Part 6's server stays on :8770 throughout. The two versions below run on :8771 and :8772 so you can compare all three without stopping anything.

Step 1 — Two new scopes​

Scopes are Property Mappings in Authentik. Customization → Property Mappings → New Property Mapping → Scope Mapping, twice:

NameScope nameDescription
office-readoffice:readLook up customers and invoices
office-refundoffice:refundIssue refunds

Two new scope mappings alongside the existing gateway-invoke and mcp-invoke

Creating them is not enough. A provider will only issue a scope it has been given, so open Applications → Providers → agent-tools → Edit and move both into Selected Scopes.

The agent-tools provider with mcp-invoke, office-read and office-refund selected

Note the line under the picker: "Select which scopes can be used by the client. The client still has to specify the scope to access the data." Both halves matter. Selecting a scope here does not put it in every token — it permits the client to ask. A client that asks for nothing gets nothing, which is why every token.sh call below passes an explicit scope string.

Confirm the token actually carries what you asked for:

~/mcp-auth/token.sh "mcp:invoke office:read" \
| cut -d. -f2 | base64 -d 2>/dev/null | jq .scope
"mcp:invoke office:read"

If that comes back without office:read, the scope is not on the provider — fix that before writing any server code, or you will spend an hour debugging enforcement that is working correctly on a token that was never scoped.

Step 2 — The obvious implementation​

Keep required_scopes=["mcp:invoke"] on the verifier as the price of admission, then ask per tool whether the caller holds what that operation needs. FastMCP exposes the verified token through get_access_token(), so a decorator can read the claims the verifier already checked:

from fastmcp.exceptions import ToolError
from fastmcp.server.dependencies import get_access_token

def requires(scope: str):
def decorate(fn):
@functools.wraps(fn)
async def wrapper(*args, **kwargs):
token = get_access_token()
held = set(getattr(token, "scopes", None) or [])
if scope not in held:
raise ToolError(
f'insufficient_scope: this tool requires "{scope}"; '
f'token carries {sorted(held) or "nothing"}'
)
return await fn(*args, **kwargs)
return wrapper
return decorate

Then one line per tool:

@mcp.tool
@requires("office:read")
async def find_customer(query: str) -> list[dict]: ...

@mcp.tool
@requires("office:refund")
async def issue_refund(invoice_id: int) -> str: ...

Run it on :8771 and drive both tools with both tokens:

Two tokens, two tools: each token allows one and refuses the other

--- token: mcp:invoke + office:read ---
find_customer : ALLOWED -> [{'id': 1, 'name': 'Maria Alvarez', ...}]
issue_refund : REFUSED -> insufficient_scope: this tool requires "office:refund"

--- token: mcp:invoke + office:refund ---
find_customer : REFUSED -> insufficient_scope: this tool requires "office:read"
issue_refund : ALLOWED -> Refunded 80.00 on invoice 2.

That is real enforcement. The refund token cannot read, the read token cannot refund, and the refusal names the missing scope. For a lot of deployments this is where you stop.

Step 3 — Why that refusal is not good enough​

Watch the HTTP layer rather than the client library.

TOK=$(~/mcp-auth/token.sh "mcp:invoke office:read")
curl -sS -i -X POST http://127.0.0.1:8771/mcp \
-H "Authorization: Bearer $TOK" \
-H "Content-Type: application/json" \
-H "Accept: application/json, text/event-stream" \
-d '{"jsonrpc":"2.0","id":1,"method":"tools/call","params":{"name":"issue_refund","arguments":{"invoice_id":2}}}'

The tool-body check returns HTTP 200 with the refusal inside the body

HTTP/1.1 200 OK
content-type: text/event-stream

data: {"jsonrpc":"2.0","id":1,"result":{"content":[{"text":"insufficient_scope: this tool
requires \"office:refund\"; token carries ['mcp:invoke', 'office:read']","type":"text"},
"isError":true}}

HTTP 200. The request succeeded; the tool declined. That is correct JSON-RPC semantics and it is useless to an OAuth client.

An OAuth client that wants to step up — go back to the authorization server and ask for office:refund — is watching for a 401 or 403 with a WWW-Authenticate header telling it what to request. It does not parse English out of a tool result. So the refusal is legible to a human reading logs and invisible to the machinery designed to handle exactly this case.

It also arrives late. The tool body runs after routing, after session setup, after the server has committed to a successful response. By then there is no status line left to change.

Step 4 — The challenge the spec actually wants​

Getting a 403 means checking scopes before the JSON-RPC layer answers, which means ASGI middleware wrapping the FastMCP app:

class ScopeChallengeMiddleware:
def __init__(self, app):
self.app = app

async def __call__(self, scope, receive, send):
if scope["type"] != "http" or scope["method"] != "POST":
return await self.app(scope, receive, send)

# Buffer the body so we can inspect it and still pass it on.
chunks, more = [], True
while more:
msg = await receive()
chunks.append(msg.get("body", b""))
more = msg.get("more_body", False)
body = b"".join(chunks)

needed = None
try:
rpc = json.loads(body)
if rpc.get("method") == "tools/call":
needed = TOOL_SCOPES.get((rpc.get("params") or {}).get("name"))
except Exception:
pass

if needed and needed not in token_scopes(dict(scope.get("headers") or [])):
challenge = (
f'Bearer error="insufficient_scope", scope="{needed}", '
f'resource_metadata="{RESOURCE_METADATA}", '
f'error_description="This operation requires the {needed} scope"'
)
await send({"type": "http.response.start", "status": 403,
"headers": [(b"content-type", b"application/json"),
(b"www-authenticate", challenge.encode())]})
await send({"type": "http.response.body",
"body": json.dumps({"error": "insufficient_scope",
"scope": needed}).encode()})
return

# Replay the buffered body downstream.
replayed = False
async def replay():
nonlocal replayed
if not replayed:
replayed = True
return {"type": "http.request", "body": body, "more_body": False}
return await receive()

await self.app(scope, replay, send)

app = ScopeChallengeMiddleware(mcp.http_app(stateless_http=True))

Same request, against :8772:

The middleware returns 403 with a WWW-Authenticate challenge naming the missing scope

HTTP/1.1 403 Forbidden
www-authenticate: Bearer error="insufficient_scope", scope="office:refund",
resource_metadata="http://127.0.0.1:8772/.well-known/oauth-protected-resource",
error_description="This operation requires the office:refund scope"

Now a client has something to act on: the status says refused, scope= says what to ask for, and resource_metadata says where to look up how. That is step-up authorization as a protocol rather than as a log message.

Step 5 — What it cost​

The middleware works. It is also worse code than the decorator, in three specific ways, and pretending otherwise would be dishonest.

It does not know what a tool is. ASGI middleware sees bytes and headers. To find out which tool is being called it parses the JSON-RPC envelope itself and looks the name up in a table it has to maintain:

TOOL_SCOPES = {
"find_customer": "office:read",
"open_invoices": "office:read",
"issue_refund": "office:refund",
}

That table is a second source of truth. Add a tool and forget the entry and it is unprotected — silently, because the middleware just passes through anything it does not recognise. The decorator could not have that bug: the requirement sat on the function.

It decodes the token unverified. The middleware runs upstream of the verifier, so the verified claims do not exist yet. It splits the JWT and base64-decodes the payload without checking the signature:

payload = auth.split(None, 1)[1].split(".")[1]
claims = json.loads(base64.urlsafe_b64decode(payload + "=" * (-len(payload) % 4)))

This is not the hole it looks like — the verifier still runs downstream and still rejects a forged token, so nothing reaches a tool on a bad signature. But the scope decision is made on unauthenticated bytes, and the only reason that is survivable is the second check behind it. It is a wart, not a vulnerability, and it is the kind of thing worth writing down before someone later removes the "redundant" verifier.

It buffers every request body. To read the envelope and still pass it downstream, the middleware drains receive() into memory and replays it. Fine for JSON-RPC calls; think harder before putting this in front of large uploads.

Troubleshooting — the errors this run actually produced​

insufficient_scope on a tool you did grant​

Check the token, not the server:

~/mcp-auth/token.sh "mcp:invoke office:refund" | cut -d. -f2 | base64 -d 2>/dev/null | jq .scope

If office:refund is missing, the scope exists as a Property Mapping but was never moved into the provider's Selected Scopes. Authentik silently drops scopes a client is not permitted to request rather than erroring, so the token comes back valid and short.

The middleware never fires​

It only inspects POST. MCP clients open a GET for the event stream first, and that request carries no JSON-RPC envelope — if you are watching the wrong request you will conclude the middleware is dead. Confirm with the curl above, which is a single POST.

Every tool suddenly returns 403​

TOOL_SCOPES is keyed by the tool's registered name, which is the function name, not the decorated label. Rename a function and the table stops matching. The pass-through case is the dangerous direction (unprotected), but a typo in the table produces this one.

The 403 body is empty in some clients​

The challenge lives in the WWW-Authenticate header. Clients that only log response bodies will show you {"error":"insufficient_scope"} and nothing about which scope. Use curl -i.

What this did and didn't buy you​

It bought genuine per-tool authorisation: two tokens that differ by one scope, each able to run exactly one of two tools, proven at the HTTP layer rather than asserted. And with the middleware, a refusal a client can programmatically recover from.

It did not buy a clean design. Both implementations are compromises pointing opposite ways:

tool decorator (:8771)ASGI middleware (:8772)
Scope requirement liveson the functionin a separate table
Reads verified claimsyesno — decodes unverified, verifier runs after
RefusalHTTP 200, isError: trueHTTP 403 + WWW-Authenticate
Client can step upnoyes
New tool unprotected by defaultnoyes

The honest summary is that the correct protocol behaviour requires leaving the abstraction the framework gives you, and the ergonomic version cannot produce it. If your clients do not implement step-up — and most agent clients today do not — the decorator is the better trade. If they do, you pay for it with a table you must remember to update.

It also did not buy authorisation that survives the tool doing something else. issue_refund is gated on office:refund; nothing stops a future find_customer from being edited to write. Scopes gate entry, not behaviour — which is the same boundary part 4 drew around policy-in-the-tool.

What's next​

The obvious remaining gap is that both versions trust the token's scope list and nothing else. Neither asks who the caller is or what they are acting on — a token with office:refund refunds any invoice, for any customer, for any amount. That is object-level authorisation, and it does not live in OAuth scopes at all.

Finished this tutorial?
Mark it complete to earn A scope per tool, and a refusal clients can act on on your skill path.

Further reading​

intermediatePart 9

Sandbox the code your agent writes, and prove every limit actually applied

· 11 min read
Rafael Fernandes
NLP Engineer & Tech Writer at WiLine
Share:
+

Every agent tutorial in this series so far has given a model tools — functions you wrote, with arguments you defined. This post is about the other thing agents do, which is write code and then run it.

That is a different risk, and the difference is worth being precise about. A tool call is the model choosing from a menu you control. Executing generated code is the model handing you something nobody has ever read, which you then run on your machine.

intermediatePart 8

Move the Docker daemon off root so a container escape lands on an ordinary user

· 14 min read
Rafael Fernandes
NLP Engineer & Tech Writer at WiLine
Share:
+

Here is the part nobody says out loud when they tell you to add yourself to the docker group so you can stop typing sudo.

That group is root. Not "close to root", not "root for Docker things". If you can run a container, you can read, change or delete any file on the machine — including the password file, including other people's home directories, including the files the administrator deliberately kept away from you.

No exploit required. It is one command, and it takes about four seconds.

intermediatePart 7

Firewall an agent container so it can reach one API and nothing else

· 21 min read
Rafael Fernandes
NLP Engineer & Tech Writer at WiLine
Share:
+

Agent orchestration part 4 split a toolbox so the scheduler could not issue refunds. Part 5 gave the agent an identity at the gateway. Part 6 put a JWT verifier in front of the tools server so it refuses anonymous callers.

Every one of those controls what the agent may call. Not one of them controls where the agent may go.

That distinction is the whole of this post. Exfiltration does not need a tool. It needs a socket.

intermediatePart 6

The wrong valid token: authenticating an MCP tools server with authentik

· 16 min read
Rafael Fernandes
NLP Engineer & Tech Writer at WiLine
Share:
authentik+Model Context Protocol

Part 5 gave an agent its own identity at the gateway: a token issued by authentik over client credentials, carrying a scope, expiring in five minutes. It ended by naming what it had not covered — the MCP tools server from agent orchestration part 4 still trusts anything that can reach its port, issue_refund included.

That post was explicit about the limit of what it had built:

The MCP server still trusts everyone. Scoping happens in the client. Anything that can reach 127.0.0.1:8770 can call issue_refund directly, agent or not.

Splitting the toolbox per role stopped an agent from reaching a tool it shouldn't. It did nothing about a curl. This part closes that.

intermediatePart 5

The key that expires: giving an agent its own identity at the gateway

· 19 min read
Rafael Fernandes
NLP Engineer & Tech Writer at WiLine
Share:
authentik+LiteLLM

Part 4 put group access control on the gateway's admin UI and then, at the end, called the API from a shell with no account, no session and no group. It answered normally. The conclusion was that SSO protects a control plane and the data plane authenticates machine callers with keys — which it has to, because an agent running at 3am cannot complete a browser login.

That was true and it was also a stopping point rather than an answer. "Machine callers use keys" leaves you with a credential that never expires, that no identity provider knows about, and that survives the person who created it. Part 4 said so plainly: removing someone from a group does not revoke their keys, a leaked key is unaffected by identity entirely, and keys outlive people.

This part gives the machine an identity instead of a key.

authentik issues the agent a token over the client credentials grant — no browser, no consent screen, no human. The token is signed, carries a scope, and expires in five minutes. The gateway verifies it locally against authentik's public keys and refuses anything without the right scope. At the end, three status codes show the boundary holding.

intermediatePart 4

The binding that wasn't there: group access control, and what SSO doesn't protect

· 16 min read
Rafael Fernandes
NLP Engineer & Tech Writer at WiLine
Share:
authentik+LiteLLM

Part 3 put authentik in front of Langfuse and ended with an admission: the bindings step was left empty, so every account in authentik could reach Langfuse and nothing would warn you.

This part closes that gap on the LiteLLM gateway — and the closing went wrong in a way worth more than the procedure. We configured the bindings in authentik's wizard, watched them appear in the wizard's own table, submitted, and ended up with zero bindings saved. The application was open to everyone, the UI said nothing, and the only reason we found out is that we tested with an account that should have been refused and wasn't.

Then, once access control genuinely worked, we called the gateway's API from a shell with no account, no session and no group. It answered normally. That one isn't a bug — it's the difference between a control plane and a data plane, and assuming SSO covers both is how people conclude they've secured something they haven't.

intermediatePart 3

One login for everything: putting authentik in front of a self-hosted app

· 20 min read
Rafael Fernandes
NLP Engineer & Tech Writer at WiLine
Share:
authentik+

Part 1 closed the ports nobody meant to open. Part 2 stopped two containers from running as root. Both were about the machine. Neither touched the question a self-hosted AI stack answers worst: who is allowed to log in, and where is that decided?

Same box as before. The Langfuse we've been hardening since Part 1 — the one whose Postgres, ClickHouse and Redis Part 1 found correctly bound to localhost, and whose worker Part 2 dropped off root — was deployed back in Catch what your tests miss. Two parts have now secured the machine underneath it without once touching the application's own front door. This is the first part that changes Langfuse itself.

Right now that decision is made in each app, separately. Langfuse has its own email-and- password table. So does every other tool on the box. Each one is a place where an account can outlive the person who owned it, where a password can be reused, and where "remove this person's access" means remembering that the app exists at all.

This post moves that decision to one place. That place is authentik — an open-source identity provider you run yourself, the self-hosted counterpart to Okta or Auth0. It holds the accounts, runs the login screen, and vouches for who someone is to any app that asks. Apps stop storing passwords and start asking authentik.

It's the same job Keycloak does, and Keycloak is the better-known name. authentik earns the pick here on setup cost: a compose file and a wizard against Keycloak's realms, clients and JVM tuning. For one box and one app, that difference is the whole decision.

We deploy it, connect Langfuse to it over OIDC, and end with a login that goes through authentik and back.

intermediatePart 2

Root by default: hardening container privilege on a self-hosted AI stack

· 8 min read
Rafael Fernandes
NLP Engineer & Tech Writer at WiLine
Share:
+

Part 1 of this series audited a live multi-agent box and found every network-exposure question worth asking. One line from that audit didn't get followed up: "the Postgres, ClickHouse and Redis behind Langfuse were published to 127.0.0.1 instead of the world — someone made a good decision there. Hold that thought."

Here's the other half of that thought: getting the network right says nothing about what happens after someone's already inside a container. If the process running in there is root, a compromise starts with the keys to the whole filesystem. So we checked — on the same box, the same Langfuse stack Part 1 already praised — whether "network correct" also meant "privilege correct." It didn't, for two of the six containers. Here's what fixing that actually looked like, including the part that broke.

intermediatePart 1

Your firewall is lying to you: hardening Docker networks for multi-agent systems

· 15 min read
Rafael Fernandes
NLP Engineer & Tech Writer at WiLine
Share:
+

You start with one agent. Then it needs a database. Then you add a second agent, a messaging bridge, an observability stack. Six months later a single WEC Instance is running five compose projects, twenty-something containers, and nobody remembers which ports are open to the world.

That's not a hypothetical — that's the box this tutorial was written on. So instead of theorizing, we probed it: can containers reach each other across stacks? Can they reach the databases? Is the firewall actually protecting anything?

Three of the answers surprised me. One of them was a database sitting there with no password. And the firewall — the firewall was lying.