Most agent tools ask you to trust them. The people who buy back-office AI are the people who get audited.
Tell an agent “never publish at org scope” and it will usually comply. Usually is not a control. In Annex the permitted scopes live in the template and the check runs inside the tool that writes — so the platform refuses regardless of what the model decided.
Retrieved memory gets the same treatment: it’s returned as reference data, never as instructions. A memory shaped like a command doesn’t become one.
Every run opens a row before the first model call and closes with its outcome, tokens, and duration. Every tool call is its own row — inputs, outputs, timing, and whether it succeeded, errored, or was blocked. Memory writes record their scope and path.
None of this is a report you generate afterwards. It’s how the system runs, so it’s complete by construction.
Reads run under row-level security as the signed-in member, and org membership is re-checked on every route — not just at the edge.
Scoped to your org in the database — not filtered in app code a bug could bypass.
The check runs in the tool handler. A jailbroken agent still can't publish above its station.
A blocked write isn't a silent no-op. It lands in the log, attributed to the run.
"Who told the agent that?" has an answer you can look up.