GitHub - rudratoshs/buried-injections: 🛡️ Regex catches 0%, Meta's Prompt Guard 2 catches 1% of 629 realistic AgentDojo injection attacks when they're buried in tool output. Reproducible benchmark.
Can open-source prompt-injection detectors catch realistic AI agent attacks?
🎯 TL;DR
I ran 10 open-source detectors against 629 real AgentDojo
injection attacks, each buried inside ordinary tool output — the way an agent
firewall actually sees them. None catches most attacks without also blocking
normal traffic.
🥇 Best trade-off: 51% caught at 2% false positives
🔴 Meta's Prompt Guard 2: 1% caught
🚫 Two detectors flag 98% of safe tool outputs too
And they fail in three different ways 👇
📊 Le...
Read more at github.com