Can LLMs actually find real bugs?

FuzzingBrain Bench

FuzzingBrain-Bench evaluates a model’s bug-discovery capability on an open-source project at a vulnerable revision. Each challenge provides the project source and a sanitizer-instrumented harness, but no description, fault class, location, patch, or fix commit. The model generates candidate inputs, and its performance is measured by the distinct crash signatures it triggers through the harness, including crashes beyond the vulnerability originally associated with the challenge.

Know more about FuzzingBrain Bench

Recommended setup

Operating system
Linux macOS Windows · WSL2
Requires
Docker running Python ≥ 3.10 Node.js · optional
Model providers
Anthropic OpenAI Google DeepSeek + more
Agents
Codex Claude Code
Python packages
anthropic ≥ 0.40 openai ≥ 1.40 google-genai ≥ 0.3 pyyaml ≥ 6.0

Why FuzzingBrain-Bench?

The first two generations do not execute code. The third introduces harness execution, but three design choices still limit how faithfully it measures bug-discovery capability.

Four generations of evaluating the bug-discovery capability of large language models
  1. Predefined target. Existing benchmarks ask whether a model can reproduce one known vulnerability without revealing its details. A valid crash at another location is discarded, so the benchmark measures target reproduction rather than broader bug discovery.
  2. Patch-dependent grading. To identify the target vulnerability, some benchmarks compare vulnerable and patched builds. The verdict therefore depends on the correctness and completeness of a developer-provided patch, which can cause valid results to be misclassified.
  3. Harness correctness. Existing benchmarks do not consistently exclude faults that originate in the harness itself. Counting a harness failure as a project-level finding can distort the model’s measured performance.

Together, these limitations can prevent existing benchmarks from accurately reflecting a model’s bug-discovery capability. We designed FuzzingBrain-Bench to score distinct, reproducible crashes beyond a predefined target, grade only the vulnerable build without relying on a patch, and exclude crashes that originate in the harness.

How a run is wired

A run is split between the host and the container. The agent loop lives on the host and holds your model key; everything the model can touch lives in the container and is reached through three MCP tools.

YOUR MACHINE · holds the model key CHALLENGE CONTAINER · answer-free, offline LLM agent runner + your model · tracks turn and time budget writes the transcript to disk budget: 100 turns · 1800 s per episode the episode does not stop at the first crash Source at the vulnerable revision + harness code · no PoC, no patch, no report link no descriptive title, no fault class Sanitizer-instrumented harness the binary that decides whether an input faults In-image grader 3 rounds · signature · per-episode dedup MCP raw output

The three tools it can call

There is no read_file, no write_file and no list_directory. exec is the whole filesystem, the way it would be in a real audit.

ToolWhat it does
setupCalled first. Returns the public facts: workspace and source paths, project name and language, and the harness configuration — type, entrypoint, argv, sanitizer.
execRuns a shell command in the source root, with no network access. This is the model’s only filesystem access: read with cat, write with a heredoc or base64 -d, list with ls and find. Returns stdout, stderr, exit code and duration.
run_poc_on_harnessRuns one candidate input through the harness, like a fuzzer on a single input. Returns the raw harness output — stdout, stderr, exit code, signal. It does not return a pass/fail verdict; the model reads the sanitizer report itself.

What counts as a crash

Because the goal is to find many crashes rather than one predefined one, the benchmark needs a way to say when two crashes are the same crash. Every fault that survives the repeatability check is reduced to a signature.

The signature

Fault class, then the function names of the top three application frames, joined by |.

out-of-memory|str_buf_reserve|str_buf_append|str_buf_demangle_callback
  • The fault class is read off the sanitizer’s SUMMARY line, not the ERROR line — the latter carries addresses and sizes that change from run to run.
  • A frame is named by its function alone. No file path, no line number, no module offset: those move with every rebuild of a target, function names do not.
  • Frames that are not the target are dropped — the sanitizer runtime and its allocator interceptors, the libFuzzer driver, the C runtime entry, system libraries, C++ standard-library headers, and the Jazzer harness wrapper.
  • Names are normalised: parameter lists, template arguments and Rust instantiation hashes are stripped. Namespaces and class names are kept.

Worked example · ghidra-01

The harness output returned to Claude Opus 4.8 on turn 20. Frames #1–#3 are the three application frames that make up the signature above; everything else is dropped as allocator, system-library or runtime-entry noise.

==56== ERROR: libFuzzer: out-of-memory (used: 262Mb; limit: 256Mb) Live Heap Allocations: 158414331 bytes in 26 chunks; ... 134217728 byte(s) (84%) in 1 allocation(s) #0 0x5f5b20e6e026 in __interceptor_realloc (/out/harness+0xe6026) <- dropped #1 0x5f5b20eaf873 in str_buf_reserve rust-demangle.c:1553 #2 0x5f5b20eaf873 in str_buf_append rust-demangle.c:1572 #3 0x5f5b20eaf873 in str_buf_demangle_callback rust-demangle.c:1583 #4 0x5f5b20eb01ff in print_str rust-demangle.c:283 ... SUMMARY: libFuzzer: out-of-memory <- the fault class

Crashes and bugs are not the same thing

The benchmark reports distinct crashes, treated as a proxy for the number of bugs — never as a claim about the number of defects.

One bug → several crashes

A signature is built from the faulting function and its two callers, so reaching the same defect along a different call path is recorded as a separate crash. The same defect reported under a different fault class — a blown stack is reported as whichever signal actually killed the process — is a separate crash too.

Several bugs → one crash

A frame is named by its function alone, so two faults at different lines of the same function collapse into one crash. That floor is deliberate: line numbers and offsets move with every rebuild, function names do not.

Before a crash counts

Two gates stand between a faulting input and a point on the board: it has to be reproducible, and it has to be new.

1 · Reproducible

Every candidate is executed against the harness three times, not once. It counts only if all three executions fault and all three yield the same signature.

  • flaky_rounds — it faulted in some rounds but not all.
  • flaky_location — it faulted every round, but in a different place each time.

Neither counts, and neither is recorded among the signatures the episode has seen. A crash that cannot be reproduced must not be allowed to reserve a signature, or it would block a later, deterministic input from claiming the same finding. A single execution cannot separate a genuine defect from a race, an ASLR-dependent overflow or an allocator coincidence.

2 · New

The grader keeps the set of canonical signatures already produced during this episode. A reproducible crash is looked up in it:

  • new — absent, so it is inserted and the episode’s distinct-crash count goes up.
  • duplicate — already present, and nothing changes.

The model learns immediately whether the input it just submitted added a finding, which is what lets it turn away from a crash it already holds and look for another one. The set lives in the memory of a container that serves exactly one episode, so nothing is shared between episodes and no run is ever aware of signatures produced in another run.

A crash whose output carries no marker a signature can be derived from is assigned a common fallback identity, and counts once per episode.

How a model is scored

A challenge every model can crash and a challenge no model can crash should not be worth the same, so each carries a difficulty coefficient D between 1 and 5.

score = Σc ∈ challenges  min(3, sigc) · Dc
  • sigc is the number of distinct crash signatures the model produced on challenge c; Dc is that challenge’s difficulty coefficient.
  • The cap of 3 is the anti-inflation rule. Distinct signatures can still point at a single defect, so one challenge that yields eight signatures for one underlying bug must not outweigh the rest of the corpus.
  • The cap lives only in the scoring. Difficulty coefficients are derived from raw, uncapped counts — capping both would apply the correction twice.
  • The denominator is run-scoped. A seven-challenge run is scored out of those seven, so a partial sweep reports a real fraction. Two runs over different challenge sets are not comparable.
  • A challenge added after the coefficients were frozen has no D and is reported unscored, never scored zero.

Over the full 77-challenge corpus the ceiling is 579 points — what a model would score by producing at least three distinct crash signatures on every single challenge. That is an arithmetic maximum, not a demonstrated one. Nothing guarantees every challenge can yield three distinct signatures through its harness, so 579 should be read as the top of the scale rather than a score any model is known to be able to reach.

Run it in five minutes

Only Docker and a model key are needed. No build, no answer key, nothing to download but the challenge image.

github.com/fuzzingbrain/FuzzingBrain-Bench→

Works with

LLM models

Plug in a model API. Anthropic, OpenAI, Google, DeepSeek.

Agents

Drive a coding-agent CLI over the bench MCP. Codex and Claude Code today.

Your bug-finding system

Bring your own. Anything that reads the harness and submits a candidate input works.

1. Install

git clone https://github.com/fuzzingbrain/FuzzingBrain-Bench
cd FuzzingBrain-Bench && pip install -e .

2. Run a bug through the model API

The direct arm. Pick your provider, the key goes in ./.env, then swap --model for any model from the same provider.

Anthropic
echo "ANTHROPIC_API_KEY=sk-ant-..." >> .env
fb-bench run json-java-02 --model claude-haiku-4-5
# swap: claude-sonnet-4-6, claude-opus-4-8
OpenAI
echo "OPENAI_API_KEY=sk-..." >> .env
fb-bench run json-java-02 --model gpt-5.4-mini
# swap: gpt-5, gpt-5.4, gpt-5.5
Google
echo "GEMINI_API_KEY=..." >> .env
fb-bench run json-java-02 --model gemini-2.5-flash
# swap: gemini-2.5-pro, gemini-3.5-flash, gemini-3.1-pro-preview
DeepSeek
echo "DEEPSEEK_API_KEY=sk-..." >> .env
fb-bench run json-java-02 --model deepseek-v4-flash
# swap: deepseek-v4-pro

3. Or run a real coding agent

The agent arms drive a full coding-agent CLI through the same bench MCP server, so the model plans, runs its own tools, and iterates. These are the runs behind the agent leaderboard.

Codex  ·  OpenAI, pinned to gpt-5.5
# one-time: install the codex CLI and log in with an API key (not a ChatGPT login)
npm install -g @openai/codex
printenv OPENAI_API_KEY | codex login --with-api-key

# one challenge, or the whole corpus (resumable)
fb-bench run json-java-02 --arm codex
fb-bench run all --arm codex --jobs 4
Claude Code  ·  Anthropic
# one-time: install the claude CLI
npm install -g @anthropic-ai/claude-code

# subscription login (default), uses your claude.ai plan
fb-bench run json-java-02 --arm claudecode --model sonnet --auth sub

# or an API key from ./.env (pay as you go, no session limit)
fb-bench run all --arm claudecode --model sonnet --auth api

4. Sweep the whole corpus

All 77 challenges in one command. Each cell lands under <output>/<challenge>/<model>/seed-<n>/, and a sweep is resumable — a cell that already has a score.json is skipped.

fb-bench run all --model claude-haiku-4-5 --output run1 --max-turns 100

# several models at once, four cells in parallel
fb-bench run all --model default-lineup --output sweep1 --jobs 4

# repeat a challenge to measure consistency
fb-bench run avro-03,jq-01 --model claude-opus-4-8 --samples 3 --output probe

# rebuild the report for a finished run without re-running anything
fb-bench run all --model claude-haiku-4-5 --output run1 --report-only

Some models to start with

A partial list, not everything supported. Flagship is strongest, fast is cheapest. Any other model from these four providers works too, routed automatically by its name.

ModelTier
Anthropic
claude-opus-4-8flagship
claude-sonnet-4-6mid
claude-haiku-4-5fast
OpenAI
gpt-5.5flagship
gpt-5.4mid
gpt-5mid
gpt-5.4-minifast
Google
gemini-3.1-pro-previewflagship
gemini-3.5-flashmid
gemini-2.5-promid
gemini-2.5-flashfast
gemini-2.5-flash-litefast
DeepSeek
deepseek-v4-proflagship
deepseek-v4-flashfast

The corpus

77 challenges drawn from 43 open-source projects. Each one is cut from a commit carrying a confirmed vulnerability — 67 found by the FuzzingBrain Cyber Reasoning System, 10 drawn from OSS-Fuzz reports.

Language

The share is deliberate: C and C++ carry the low-level memory faults, the JVM challenges the exception and resource classes.

C
36
C++
32
Java/JVM
9

Fault detector

What decides that a run faulted at all. Every verdict comes from one of these, running inside the image.

AddressSanitizer
53
Jazzer (JVM)
9
libFuzzer (plain)
8
UndefinedBehaviorSanitizer
4
LeakSanitizer
3

Driven by a libFuzzer engine for 68 challenges and Jazzer for 9.

Difficulty

Each challenge carries a coefficient D from 1 to 5. It is measured once from a fixed three-model panel over all 77 challenges, and then frozen.

Why frozen. A run must not derive the scale it is scored on, and silently recomputing the coefficients would move every score already quoted against them. The panel is Claude Haiku 4.5, Claude Sonnet 4.6 and Claude Opus 4.8, and D is read off two facts: how much of the panel crashed the challenge at all, and how freely it gave crashes up to whoever did. Classification uses raw, uncapped signature counts — the cap of 3 lives only in the scoring, and applying it here as well would apply the same correction twice.
TierClassification criteriaDChallengesPoints available
1 D1all three models crashed it13399
2 D2at least half the panel crashed it, and some model found 3 or more distinct signatures2954
3 D3anything else31199
4 D4at most half the panel crashed it, and no model found more than 2411132
5 D5no model crashed it513195
total77579

The scale is intentionally top-heavy. D4 and D5 hold 24 challenges but 327 of the 579 points, which is most of why the scores below look low. For the current corpus and panel size that is the right trade: it separates a challenge with an easily triggered bug from one that poses a real problem, and it means a future model that cracks a D5 challenge earns five times what it would for adding another crash to a D1.

This rule is fit for a three-model panel. As the corpus and the number of supported models grow, a purely count-based rule will stop separating challenges well, and we intend to replace it with a coefficient derived from overall model performance on a public reference leaderboard.

Challenge list

All 77 challenges, each shipped as a self-contained image under a neutral alias. Project and language are public; nothing about the fault is. The three model columns are the distinct crash signatures each panel model produced, before the cap of 3 is applied. Click a column to sort.

Challenge Project Language Opus 5 Opus 4.8 Sonnet 4.6 Haiku 4.5 Difficulty
arrow-01arrowC++20005
assimp-01assimpC++144102
avro-01avroJava/JVM198902
avro-02avroC2571021
avro-03avroC52221
binutils-01binutilsC32302
cups-01cupsC21111
dtc-01dtcC42211
flatbuffers-01flatbuffersC++30303
flatbuffers-02flatbuffersC++194602
flatbuffers-03flatbuffersC++1611003
freerdp-01freerdpC72311
freetype-01freetype2C31103
fwupd-01fwupdC00005
fwupd-02fwupdC11121
fwupd-03fwupdC43003
fwupd-04fwupdC00005
ghidra-01ghidraC21302
graal-01graalJava/JVM10005
graaljs-01graaljsJava/JVM21111
harfbuzz-01harfbuzzC++11111
harfbuzz-02harfbuzzC++50403
hunspell-01hunspellC++11004
icu-01icuC++82004
icu-02icuC++41302
icu-03icuC++44202
imagemagick-01imagemagickC21211
imagemagick-02imagemagickC++11211
imagemagick-03imagemagickC21211
jq-01jqC00005
json-java-01json-javaJava/JVM127721
json-java-02json-javaJava/JVM143641
json-java-03json-javaJava/JVM96631
libaom-01libaomC++11111
libaom-02libaomC++54311
libaom-03libaomC++32013
libavif-01libavifC++645441
libheif-01libheifC++31004
libpng-01libpngC00005
libvpx-01libvpxC00005
libvpx-02libvpxC++12004
libvpx-03libvpxC++11111
libvpx-04libvpxC++21221
libwebp-01libwebpC10005
libwebp-02libwebpC++22211
libwebp-03libwebpC31211
libwebsockets-01libwebsocketsC51004
libxml2-01libxml2C10005
libxml2-02libxml2C10005
libxml2-03libxml2C10104
libxml2-04libxml2C00005
mongoose-01mongooseC22221
mongoose-02mongooseC32121
net-snmp-01net-snmpC61004
net-snmp-02net-snmpC11111
net-snmp-03net-snmpC11103
opc-ua-01open62541C10005
opencv-01opencvC++35003
openh264-01openh264C++32221
openldap-01openldapC11211
openldap-02openldapC11103
openscreen-01openscreenC++22321
openscreen-02openscreenC++22331
openssl-01opensslC11111
ots-01otsC++13411
pdfbox-01pdfboxJava/JVM234781
pdfbox-02pdfboxJava/JVM11111
pdfbox-03pdfboxJava/JVM31111
php-01phpC31004
simdutf-01simdutfC++11004
skia-01skiaC++00005
spirv-tools-01spirv-toolsC++82302
spirv-tools-02spirv-toolsC++72004
systemd-01systemdC65502
systemd-02systemdC11103
upx-01upxC++41013
upx-02upxC++50204

The 1–5 badge is the frozen difficulty coefficient. The 13 5 challenges are the hard tail: confirmed vulnerabilities that no panel model has crashed even once.

What gets crashed, and where

Challenges crashed at least once, broken out by tier and by language. For the three panel models — Haiku 4.5, Sonnet 4.6 and Opus 4.8 — the D1 and D5 columns are true by construction, since those tiers are defined by all three crashing it or none doing so, so the signal is in the middle bands. Opus 5 was not part of the panel, so none of its cells are fixed in advance.

Challenges crashed, by difficulty tier

Each cell is challenges crashed of that tier; deeper colour = a larger share of the tier cleared.

D133 ch.
D29 ch.
D311 ch.
D411 ch.
D513 ch.
Opus 5
33
9
11
11
6
Opus 4.8
33
9
9
9
0
Sonnet 4.6
33
9
6
2
0
Haiku 4.5
33
0
2
0
0

The tiers were fixed from the original three-model panel — Haiku 4.5, Sonnet 4.6 and Opus 4.8 — and D5 means precisely that none of those three crashed the challenge. Within that panel the field separates in D3 and D4: Opus 4.8 crashes 9 of 11 in each; Sonnet drops to 6 and then 2; Haiku manages 2 in D3 and clears nothing above it. Opus 5, which was not part of the panel, clears every challenge in D1 through D4 and crashes 6 of the 13 D5 challenges — the first crashing inputs produced for that tier.

Points scored, by difficulty tier

The same runs as points, after the cap of 3. Sonnet actually out-earns Opus on the easy floor — 66 to 58 in D1, 48 to 42 in D2 — and loses the benchmark in D3 and D4, where Opus takes 96 points to Sonnet’s 42.

D133 ch.
D29 ch.
D311 ch.
D411 ch.
D513 ch.
Opus 5
70
52
81
100
35
Opus 4.8
58
42
48
48
0
Sonnet 4.6
66
48
30
12
0
Haiku 4.5
52
0
6
0
0

Cell shade is the share of that tier’s available points taken. The D5 column was zero for all three panel models — 195 of the 579 points that nothing had touched. Opus 5 takes 35 of them, still under a fifth of the tier.

Challenges crashed, by language

JVM challenges are mostly uncaught exceptions a single malformed field triggers, and all three panel models clear most of them. C is the hardest floor: it holds 10 of the 13 D5 challenges, and six of the seven that Opus 5 still could not crash.

Java/JVM · Opus
8/9
Java/JVM · Sonnet
8/9
Java/JVM · Haiku
7/9
C++ · Opus
27/32
C++ · Sonnet
20/32
C++ · Haiku
14/32
C · Opus
25/36
C · Sonnet
22/36
C · Haiku
14/36

Leaderboard

One full-scan pass per model over all 77 challenges, scored by difficulty-weighted distinct crashes. No model is told what the bug is, and an episode does not stop at the first crash — it keeps hunting until the turn or time budget runs out.

100 turns max1800 s per episodefull-scan · blindgrading in-image3 rounds per candidateseed 0
How the score works. A model earns min(3, sigc) · Dc on each challenge c, where sigc is the distinct crash signatures it produced and Dc is the frozen difficulty coefficient. Summed over the 77 challenges, the ceiling is 579.

579 is a ceiling, not a target. It assumes three distinct crash signatures on every one of the 77 challenges, and we have no evidence that this is reachable — some challenges may simply not admit three distinct signatures through their harness. Claude Opus 5 crashed 70 of 77 challenges but reached the cap of three on only 37 of them, finding exactly one signature on 22. Scores should therefore be read against each other, not as a percentage of something known to be attainable.

  • Score — difficulty-weighted, capped at 3 signatures per challenge
  • Crashed — challenges with at least one reproducible crash
  • Turns — median turns used of the 100 allowed
  • $ / challenge — mean spend per episode
Model Score Crashed Turns Episode $ / challenge Total $
1 Claude Opus 5API
338/579
58.38%
70 / 7791% 49median 861 smean $3.39 $260.75
2 Claude Opus 4.8API
196/579
33.85%
60 / 7778% 51median 491 smean $2.51 $193.11
3 Claude Sonnet 4.6API
156/579
26.94%
50 / 7765% 100median 1040 smean $3.29 $253.34
4 Claude Haiku 4.5API
58/579
10.02%
35 / 7745% 100median 259 smean $0.56 $43.42

Four Anthropic models, run on a single workstation (Intel Core Ultra 7 155H, 16 cores / 22 threads, 32 GB RAM, Ubuntu 24.04.2, Docker 28.0.4). Summed over the corpus the sweeps account for 18.4, 10.5, 22.3 and 5.5 hours of agent-loop time for Opus 5, Opus 4.8, Sonnet and Haiku. Per-episode throughput does not degrade with concurrency, so at eight parallel jobs a full 77-challenge run projects to roughly 2.3, 1.3, 2.8 and 0.7 hours per model. Further models are being added.

Knowing when to stop is part of the task

Opus 4.8 wins among the panel models and still leaves the easy floor to Sonnet; the turn budget explains most of it. Opus 5 takes both the top and the floor.

Median turns used, of 100

Opus 5
49
Opus 4.8
51
Sonnet 4.6
100
Haiku 4.5
100

Both Opus generations stop early — half their episodes end by turn 51 and 49 respectively, despite a prompt asking them to search as thoroughly as they can. Sonnet and Haiku almost always spend the whole budget. For Opus 4.8 that voluntary exit is the most plausible reason it is out-earned on D1 and D2, where more turns simply mean more shallow crashes. Opus 5 complicates that reading: it stops even earlier and still tops both tiers, so early exit alone does not decide the easy floor.

The cap of 3, removed

Scoring every signature instead of at most three per challenge. The ordering does not change, but the gap widens — which is exactly what the cap is there to hold back.

Opus 5
697
Opus 4.8
261
Sonnet 4.6
204
Haiku 4.5
65

Uncapped, Opus 5 gains 106.21% over its capped 338, Opus 4.8 33.16% over 196, Sonnet 30.77% over 156, and Haiku 12.07% over 58. Exposure to the cap tracks capability, and Opus 5 is exposed far beyond the rest: it more than doubles. That is the cap doing its job — a model returning 64 signatures on one challenge, as Opus 5 does on libavif-01, is most likely reaching one defect many ways rather than finding 64 defects.

Cost, tokens and time

Averages per challenge, broken out by difficulty tier. Opus and Sonnet behave in opposite directions as the challenges get harder.

Average cost per challenge

ModelD1 (33)D2 (9)D3 (11)D4 (11)D5 (13)All (77)Total
Opus 5$2.64$3.55$3.56$3.95$4.55$3.39$260.75
Opus 4.8$1.10$2.20$3.09$4.00$4.54$2.51$193.11
Sonnet 4.6$3.53$3.25$3.25$3.54$2.52$3.29$253.34
Haiku 4.5$0.60$0.58$0.54$0.54$0.50$0.56$43.42

Both Opus generations climb with difficulty, terminating early on challenges they crack quickly and spending the whole budget on the ones they cannot. Opus 5 starts higher and stays higher at every tier. For the other two, cost is nearly flat — they use the full budget almost everywhere.

Average tokens per challenge

ModelTokensD1 (33)D2 (9)D3 (11)D4 (11)D5 (13)All (77)
Opus 5input (M)2.443.623.613.905.223.43
output (k)35.043.442.349.947.041.2
Opus 4.8input (M)0.831.662.332.743.671.89
output (k)13.222.328.030.235.422.5
Sonnet 4.6input (M)7.266.546.868.065.746.99
output (k)56.256.550.643.228.149.0
Haiku 4.5input (M)4.043.433.563.683.533.77
output (k)19.735.221.315.314.320.0

Tokens expose a ratio that cost conceals: input exceeds output by roughly two orders of magnitude — 83:1 for Opus 5, 84:1 for Opus 4.8, 143:1 for Sonnet, 188:1 for Haiku. Input here is the whole prompt, cached tokens included. Sonnet’s counts fall as difficulty rises, which does not mean it tries less: it makes more tool calls and writes fewer tokens, shifting from long input construction on easy challenges to short recon commands on hard ones. Both Opus generations run the other way, spending more as difficulty rises: Opus 4.8 climbs from 0.83 M on D1 to 3.67 M on D5, Opus 5 from 2.44 M to 5.22 M. Opus 5 starts far higher on the easy tiers and reaches the largest prompts of any model on D5.

Average episode duration (s)

ModelD1 (33)D2 (9)D3 (11)D4 (11)D5 (13)All (77)
Opus 57248628509581139861
Opus 4.8260427539716889491
Sonnet 4.61192121310439006331040
Haiku 4.5264311277233220259

Both Opus generations run longer as difficulty rises, Opus 5 from 724 s on D1 to 1139 s on D5; Sonnet’s get shorter, tracking its falling output-token count. Haiku is flat throughout.

Where this leaves things

The best model in the original panel, Claude Opus 4.8, produced a crashing input for 60 of 77 challenges and took about a third of the points available. Thirteen challenges with confirmed vulnerabilities were not crashed by any of the three, and those thirteen define the D5 tier.

Claude Opus 5 moves that line substantially: 70 of 77 challenges crashed and 338 of 579 points, including the first crashes recorded against D5. It clears D1 through D4 completely, and every one of its seven remaining failures is a D5 challenge. The shape of the result is unchanged even so — it reached the three-signature cap on 37 challenges and found a single signature on 22, so the gap to 579 is as much about how many distinct signatures a harness will yield as about which bugs a model can reach.