About this Paper
Standard prompt-injection benchmarks reward detectors that flag nearly everything, so models that score well often break in production. This paper explains why and introduces Hydra, a multi-head classifier built for real deployment: one shared encoder with four separately calibrated heads for prompt injection, SQL injection, PII and toxicity. At a 1% false-positive rate, Hydra catches 57.8% of attacks.
What's Covered
- The over-defense failure mode in standard benchmarks
- A contamination audit of every benchmark
- Hydra's architecture and per-head operating points
- Recommended reporting standards for guard classifiers
