Anthropic reports zero-percent prompt injection for its browser agent
Anthropic says its Opus 5 model is nearly immune to prompt injection, the attack where hidden instructions, such as text buried in a webpage, hijack an AI agent. According to the company's system card, Opus 5's browser agent recorded a zero-percent attack success rate across 129 test scenarios. On the separate Gray Swan benchmark, attacker success after 15 attempts fell from 5.5 percent for Opus 4.8 to 2.0 percent for Opus 5, which now leads that benchmark ahead of two rival models.
The zero-percent figure comes with a firm condition: it holds only when Auto Mode is enabled in products like Claude Cowork. Auto Mode adds two defenses, one scanning incoming data for hidden instructions and one blocking dangerous actions before they run. Without those layers, Opus 5 sits at 3.7 percent. So the result reflects a model-plus-software combination, not the model alone.
Prompt injection has been the central obstacle to trusting AI agents with real tasks like browsing, buying, or handling accounts. If it can be driven close to zero in practice, delegating routine digital work becomes safer, which is the precondition for agents doing genuinely useful things cheaply and at scale.
The caveats matter. These figures come from Anthropic's own system card and a third-party benchmark, not independent audits, and OpenAI stated in December that prompt injection may never be fully solved. Worth watching is whether outside testers reproduce the numbers and whether the defenses hold against novel attacks.
Source: The Decoder
MANY MINDED