Skip to main content
BENCHMARKS

8 of 8 on OWASP Juice Shop.

Graded by the target's own scoreboard, not by reading our report. We withdrew our previous numbers in August because the runs behind them had been cut short and the write-up was done by hand. These come from untruncated runs, scored by the harness, with the run count and the exclusions stated.

OWASP JUICE SHOP · 8 CHALLENGES · 2026-09-02 & 2026-09-28
2026-09-28 run8 / 8recorded

reproduced the full sweep three and a half weeks later on the same engine; $0.38 in tokens, 17 findings filed alongside

2026-09-02 run8 / 8recorded

every challenge solved, including Repetitive Registration, which had never been solved in any prior run

Full-stack run7 / 8recorded

DOM XSS and Forged Review both solved; only Repetitive Registration missed

Two earlier runs7 / 8, 7 / 8observed

ended early by hand to free the queue; a killed scan writes no harness score

One run excludedvoidexcluded

47 connection failures in 219 events; it measured a dead container, not the engine

Model: the Ryvx internal engine, the same one hosted scans run on. The score on a frontier model is unmeasured, not assumed, and on this corpus we have no reason to expect it would be higher: a run on a model roughly fifteen times more expensive per token scored lower on the same eight challenges the same night. On 2026-09-28 two further models, one at our own price point and one below it, were run on the same eight challenges and scored 1 of 8 each: the recall tracks the engine and its tooling, not the price per token.

The 8 of 8 has now scored twice, on 2026-09-02 and again on 2026-09-28, three and a half weeks apart on the same engine, at $0.37 and $0.38 in tokens. Two scored runs is still not a rate, so we quote them as measurements and not as "we score 8/8"; we will replace them with an averaged figure once we have several consecutive runs. Repetitive Registration, unsolved in every run before 2026-09-02, was solved in both.

WHAT IT FILED, AND WHAT WE CAN ACTUALLY CLAIM

On a separate Juice Shop run, Ryvx filed 19 findings the evidence gate accepted and rejected 10 for incomplete evidence. Several of those were completed and re-filed by the agent, which is the retry loop working. The CWE classification was correct on all 19, and the one proof-of-concept we executed by hand reproduced first try against the live target: 15 records, 6 users, no auth, exactly as the finding claimed.

What that is not: an independently verified precision figure. Ryvx did not run those PoCs itself. A person did. We are not going to quote a false-positive rate off a set we checked one member of. These numbers are also ours, on a public CTF target we triaged ourselves; treat them as our own measurement rather than an independent evaluation.

WHY THE OLD NUMBERS CAME DOWN

In the run we had been quoting, the coordinating agent finished while five subagents were still working. They were given sixty seconds and then cancelled mid-request. One had just had a finding rejected over a formatting error and never got to re-file it. The score was real, but it measured an investigation that stopped early.

Agents now get as long as their own turn and cost budgets allow, every run records whether anything was cut off, and the harness refuses to score a run it detects as truncated or aimed at an unreachable target, which is why one run above is void rather than a 3 of 8. A tool that refuses to report an unproven finding should hold itself to the same rule.

STAY IN THE LOOP

Release notes and product updates, by email.

We'll send a confirmation email; you're not on the list until you click the link in it. See our privacy policy.