In this post, we explore how a simple website summary request hijacksClaude Code Opus 5inAuto Modeand achieves code execution with 60-80% attack success rate using a small sample size.
This is interesting because a third-party evaluation commissioned by Anthropic showed a0.00%prompt injection attack success rate for Opus 5 in Auto Mode.Auto Mode Is Now the Default in Claude Code
Auto Mode replaces human approval prompts with a safety classifier. Since mid-August it is the default starting mode for Claude Code.
To make my key point right away: If you care about what’s happening and are worried about misalignment, hallucinations and prompt injection, then Auto Mode IS NOT a substitute for running your agent in an isolated environment and monitoring what it is up to.
Boris Cherny from Anthropic recently posted that layered defenses could reduce indirect prompt injection on unseen attacks to approximately zero. The layers were model training, input probes and an intent classifier. They hired a vendor (Trajectory Labs) to test 72 indirect prompt injection scenarios ten times each. The evaluation seems to not have a published benchmark name, and the shared chart shows 0.00% attack success for Opus 5 in Auto Mode.
I wanted to see how that result holds up against a targeted attack chain.
...read more at embracethered.com
pull down to refresh
related posts
This is in case you didn't realize yet that all these graphs they pitch are bulshytt
Boris Cherny charts?
All the charts! (But yeah, the ones made/sponsored by the AI labs)
I've seen a lot of bs charts over the years but these guys just make 'em up as they go. And then defend it by challenging definitions.
Like a dog takes after its owner and the owner takes after its dog, LLMs take after their creators and their creators take after their LLMs.