Sandbox the code your agent writes, and prove every limit actually applied
- 1Lock down Docker networks
- 2Run containers as non-root
- 3Identity in front of every port
- 4Membership, not just an account
- 5An identity for the agent, not a key
- 6A tool server that checks who is asking
- 7Where the agent can go, not just what it can call
- 8Move the daemon off root
- 9Run the model's own code without trusting it
- 🏆A scope per tool, and a refusal clients can act on
Every agent tutorial in this series so far has given a model tools — functions you wrote, with arguments you defined. This post is about the other thing agents do, which is write code and then run it.
That is a different risk, and the difference is worth being precise about. A tool call is the model choosing from a menu you control. Executing generated code is the model handing you something nobody has ever read, which you then run on your machine.




