# Why route Decepticon at ExploitBench?

Strategic motivation for the `benchmark/exploitbench-adapt` branch.

## TL;DR

ExploitBench measures something XBOW cannot: how far up the
exploitation ladder a frontier model walks, not whether it grabs a
flag. Wiring Decepticon at it converts the existing harness from a
binary CTF runner into a **capability-graded evaluation surface** —
the same surface the public leaderboard publishes against
`anthropic/claude-mythos-preview`, `openai/gpt-5.5`, and
`anthropic/claude-opus-4-7`. Every gain below is downstream of that
shift.

## What XBOW gives us today

```
provider XBOW  →  binary pass/fail (FLAG{<sha256>} captured or not)
                  filtered by level (1-3) and tag (sqli/xss/…)
                  driven by Decepticon's main agent + sub-agent middleware
```

Strengths: cheap, hermetic, easy to score, ~100 challenges available
already.

Limits:

- **One bit of evidence per run.** Solve-rate compresses every
  partial result into "fail". A run that landed `addrof` but missed
  `ace` looks identical to one that never reached the patched line.
- **All targets at the web-app layer.** No memory-corruption surface,
  no JIT compiler bugs, no V8 sandbox, no exploit-primitive
  composition.
- **No vendor-grade frontier comparison.** XBOW results don't map onto
  any external leaderboard, so we can't say where Decepticon places
  against Mythos / GPT-5.5 / Opus on the same axis.

## What ExploitBench gives us

```
provider ExploitBench → 16-capability bitmap per run
                       collapsed into 5 tiers (T1 ace … T5 coverage)
                       byte-identical with exploitbench.ai numbers
                       driven by the same Decepticon agent loop
```

The provider is a thin adapter (≈600 lines including capability
helpers) — it does not change the agent loop, does not change the
harness, does not change the sub-agent topology. It changes what the
benchmark *grades*.

### Concrete gains, in priority order

1. **Capability-graded signal.** Every result carries a 16-bit
   bitmap, a tier label, and a continuous score. A failing run now
   tells us *how it failed* (cov_func only ≠ caged_write ≠ infoleak):
   the bench-v8 capability ladder is the right level for diagnosing
   sub-agent regressions because each transition (T5→T4→T3→T2→T1)
   exercises a different skill — coverage discovery, crash crafting,
   JS-engine primitive construction, sandbox bypass, control-flow
   hijack.

2. **External calibration.** ExploitBench publishes a leaderboard
   keyed on the same 16 capabilities. Decepticon results computed by
   this provider drop straight onto that axis: we can say
   "Decepticon + claude-opus-4-7 nudged hit T1 on X/41 bugs against
   the public 27/41 Opus baseline" without inventing our own metric.
   That number is the cheapest external proof that our agent loop is
   competitive with vendor scaffolding.

3. **Pressure on the right components.** The bench-v8 tasks stress
   exactly the parts of Decepticon that are hard to test
   internally:

   - `setup()` → sub-agent recon over a C++ source tree.
   - `exec()` → harness sandbox stability under long-running native
     subprocesses (gdb, lldb, autoninja).
   - `grade()` → multi-turn capability accumulation; punishes
     amnesiac agents.
   - end-to-end → ladder climbing tests OPPLAN discipline (you
     cannot accidentally exploit your way to T1; the agent must plan
     across hours of compute).

4. **N-day forecasting workload.** Each bug-image is a 1-day exploit
   scenario: vulnerable commit + patch commit + working environment.
   Running Decepticon on the 41-bug v8 corpus produces a per-bug
   tier curve we can correlate with Decepticon's prompt iterations,
   sub-agent versions, and model knobs — i.e. it gives us a
   *regression test* for the agent loop instead of a one-shot demo.

5. **Defensive metric.** Mythos hit T1 on 21/41 bugs; the human
   observations post argues every defender now has to assume the
   patch-gap is "how long does it take an LLM to climb the ladder".
   For Decepticon-as-defender (triage, repro, severity scoring), the
   same ladder is the right measurement axis: we can publish
   "Decepticon reaches T3 on Y/41 patched bugs in 30 minutes" as a
   triage-readiness number. That's the kind of claim corporate
   defensive customers actually care about.

6. **Honest failure mode for security tasks.** XBOW lets a degraded
   sub-agent look like a small regression in pass-rate. ExploitBench
   surfaces *what the agent stopped being able to do*: losing T2
   primitives is a different bug from losing T5 coverage. Future
   regressions land in PRs with capability deltas attached, which is
   far more useful than a flag-rate trend line.

7. **Cheap baseline for OSS / cheap-tier models.** The bench-v8
   environments scale down to small-context models (Haiku, GLM, Kimi,
   MiniMax) without code changes. We can publish a Decepticon
   side-by-side on cheap models the same way the upstream leaderboard
   already does, and immediately know whether our agent loop adds
   capability *over the raw model* — the multi-agent thesis of
   Decepticon. That number is currently unmeasured.

8. **Re-runs without flag pollution.** ExploitBench's grader signs
   per-round ACE flags inside the V8 d8 patch, so re-running a bug
   cannot cause an agent to reuse a leaked flag. Decepticon's XBOW
   provider uses static `FLAG{sha256(ID)}` strings which leak across
   runs once a sub-agent caches them. Migrating future high-stakes
   bugs to this grader pattern is a process win that the
   ExploitBench port unlocks.

### Non-goals (things this branch is NOT trying to do)

- It does not change Decepticon's agent loop. Same `decepticon.md`
  prompt, same `EngagementContextMiddleware`, same sub-agents.
- It does not run reinforcement learning against the benchmark —
  upstream explicitly asks consumers not to, and our agent loop is
  not RL-shaped anyway.
- It does not replace XBOW. Web-app CTFs are still useful for the
  recon-→-exploit pipeline and stay on the XBOW provider. The two
  providers coexist behind `--provider`.

## Cost shape

The 14-bug small cohort against one model and one seed is the cheapest
parity check we can run. ExploitBench publishes per-bug ~$80-200 on
Opus 4.7 and ~$20 on GPT-5.5 with 300-turn budgets; the same matrix
under Decepticon will cost roughly the same since the agent loop is
the dominant token producer. The smoke config (1 bug, 1 seed) is the
go/no-go before scaling, matching upstream's "verification ladder"
discipline (`docs/RUNBOOK.md`).

## What success looks like at PR merge

- One smoke run lands a complete `ChallengeResult` with non-empty
  `capabilities`, `tier_reached`, and `capability_score`.
- The 14-bug baseline run reproduces a tier distribution within ±1
  tier per bug against the public Opus 4.7 nudged row.
- A second-run regression report uses the capability bitmap delta as
  the primary signal, not the binary pass/fail column.

When those three boxes are checked, the branch is providing more
information per dollar of compute than XBOW does today, and our agent
loop has its first external calibration point against frontier model
scaffolding.
