Blog
Securing AI agents, in depth
How the attacks on agents actually work, what stops them, and what does not. Every number links to the benchmark it came from, losses included.

MCP rug pulls: why approving an MCP server once is not enough
An MCP server can pass review and change its tools a week later. How tool poisoning and rug pulls work, and how to catch drift at the moment of the call.

Prompt injection detection is a speed bump. Contain the tool call instead.
You will not detect every prompt injection. Capability grants, impact tiers and argument provenance stop the damaging tool call even when detection misses.

Which Claude Code hooks can actually stop a tool call
Claude Code hooks are the only place a coding agent's tool call can be refused before it runs. Which hooks can block, which cannot, and a working baseline.

We attack our own AI agent every hour and publish what gets through
A red-team report is out of date the day after it ships. How scheduled, opt-in probes against a deployed agent work, and why we publish our own misses.

Indirect prompt injection that hides: three techniques a keyword filter misses
Letter-spaced overrides, instructions hidden in markup, and persona jailbreaks. What they look like in a retrieved document, and how a detector catches them.

An AI agent audit trail your auditor can check without your vendor
A log the vendor can rewrite is not evidence. How a hash-chained audit trail, signed checkpoints and a standalone verifier turn agent decisions into proof.
Subscribe with the RSS feed, or watch our own agent get attacked every hour.
Red team a call in the browser
No account. The playground uses the same enforcement code as the product.
Offline: no API key, no downloaded weights, no network egress. Install from the docs.