Skip to main content

2 posts tagged with "guardrails"

View all tags

An agent can narrow its own web search, but never widen it

· 4 min read
Rafael Fernandes
NLP Engineer & Tech Writer at WiLine
Share:
Agents · AI News

Narrow only, never wider

Please use trusted sourcesTrusted sources are all there are

Your agent can search the web. You write in the prompt: only use sec.gov and the big financial wires. It usually listens. The times it does not are the times you learn that a line in a prompt is a request, not a rule.

On 19 August AWS added site filters to the web search tool in Bedrock AgentCore. Small feature. But the way it handles a disagreement is worth knowing, because that is what makes it a guardrail instead of a note.

Mask Your Logs, Not Your Prompts

· 7 min read
Rafael Fernandes
NLP Engineer & Tech Writer at WiLine
Share:
Privacy · AI News

Mask your logs, not your prompts

Redact before the modelRedact before the logs

Almost every "secure your LLM app" guide gives the same advice: before a prompt reaches the model, strip the personal data out of it — swap names, emails, and account numbers for [REDACTED] or <PERSON>, then call the model.

It sounds obviously right. I assumed it was, too. Then I went and read the research on what masking actually does to a model, and the papers point the other way. The short version: mask your logs, not your prompts.