<style>

/* ---------- Primary button ---------- */

.button_bg {
  background-color: #a83018;
  transition: background-color 0.3s cubic-bezier(0.4, 0, 0.2, 1);
}

.button:hover .button_bg,
.button:focus-visible .button_bg,
.submit-wrapper:hover .button_bg,
.submit-wrapper:focus-within .button_bg {
  background-color: #d24a2f;
}

/* ---------- Secondary button ---------- */

.button_bg.secondary {
  background-color: var(--neutral--900);
  border: 1px solid var(--neutral--500);
  transition:
    background-color 0.3s cubic-bezier(0.4, 0, 0.2, 1),
    border-color 0.3s cubic-bezier(0.4, 0, 0.2, 1);
}

.button:hover .button_bg.secondary,
.button:focus-visible .button_bg.secondary,
.submit-wrapper:hover .button_bg.secondary,
.submit-wrapper:focus-within .button_bg.secondary {
  background-color: var(--neutral--700);
  border-color: var(--neutral--300);
}

</style>
[fs-list-field]:has(> *),
.tag-component {
  transition: background-color 0.25s ease, border-color 0.25s ease;
}

.fs-list-active .tag-component,
.tag-component.fs-list-active {
  background-color: #d24a2f;
  border-color: #d24a2f;
}
Research

Hydra: Multi-Head Prompt Security Classification and the Over-Defense Failure Mode of Prompt Injection Benchmarks

Why standard prompt-injection benchmarks reward over-flagging, and a multi-head classifier built for deployment

About this Paper

Standard prompt-injection benchmarks reward detectors that flag nearly everything, so models that score well often break in production. This paper explains why and introduces Hydra, a multi-head classifier built for real deployment: one shared encoder with four separately calibrated heads for prompt injection, SQL injection, PII and toxicity. At a 1% false-positive rate, Hydra catches 57.8% of attacks.

What's Covered

  • The over-defense failure mode in standard benchmarks
  • A contamination audit of every benchmark
  • Hydra's architecture and per-head operating points
  • Recommended reporting standards for guard classifiers

Lower runtime risk.
Govern without friction.

Run at agent speed.

See across every cloud, vendor,
and team in days, not months.