Watch an AI agent get hijacked — then stopped
A support copilot retrieves a poisoned document and tries to email your customer data to a stranger. Flip one switch and run the same attack again — this time the send is refused and waits for a person. Deterministic replay, no live model, nothing to install.
One agent. One poisoned document. One switch.
The same attack runs in both modes. In observe, enforcement is off: the collector records every call and lets it through, so you get the record after the data is gone. In enforce, your workspace policy says sending email needs a person's approval, so the send is refused before it runs.
A support copilot is about to retrieve a poisoned document. Press Run the attack to watch it unfold — then flip to Enforce and run it again.
A policy on the tool, decided before the call runs
The collector did not have to recognise the injection. The policy names the tool — any send_email call needs a person's approval — and the collector's MCP proxy decides that before the call reaches the mail server. Trefur records each call the proxy sees, and the refused one carries its reason code.
- callkb/search_knowledge_base
- callfiles/read_file
- callmail/send_email[refused · policy.approval_required]
- agent
- support-copilot
- server
- outcome
- denied
- reason
- policy.approval_required
What was refused, for which agent, and the reason code — not a score, not a rating.
A refusal at the MCP proxy is reported, not just enforced
When the collector refuses a call from an identified agent at the MCP proxy, it reports the refusal to your workspace: it is listed in the enforcement history with the server and the reason code, and recorded as a step. Refusals at a coding agent's hook are enforced on the machine but are not reported this way.
Every agent you deploy, governed like this
Start by watching — every agent step on the record. Turn on enforcement when you're ready, and keep the record either way.