Guidance is not a mechanism
2026-09-02
We set out to answer a straightforward question: is Ryvx held back by the model it runs on? The internal engine is cheap, and cheap felt like it ought to be the constraint. So we priced up the alternatives, picked a model roughly fifteen times more expensive per token, and ran it against the same eight OWASP Juice Shop challenges we have been scoring against for months.
It scored lower. Not dramatically, but lower, on the same corpus, on the same night, for about fifteen times the money. That result is what the rest of this post is about, because the reason it lost turned out to be far more interesting than the score.
Three misses, one cause
One Juice Shop challenge requires reaching /ftp/acquisitions.md, a directory that is not linked from the application anywhere. The cheap engine found it by requesting /ftp twenty-four times until something stuck. The expensive engine never requested it once. Our prompts name /ftp explicitly, in a checklist called the unauthenticated exposure sweep. The difference was not that one model knew about the path and the other did not. The difference was that the checklist had been moved behind an on-demand tool call, and one model thought to make that call and the other did not. In a separate run, that tool was called zero times across 587 events.
A second challenge asks whether the registration endpoint enforces that a password and its confirmation match. We pulled every registration our agents had ever sent and counted: sixteen requests carrying both fields, and all sixteen sent them identical. The agent was not failing to exploit the bug. It had never attempted it. Registration was something it used to obtain a token so it could go and test other things, and it never turned around to ask whether registration itself was broken.
A third: a hunter posted product reviews to test for stored cross-site scripting, and set the author field on every one of them to its own generated address. The challenge is to set that field to somebody else's. The agent had the exact request in hand and never varied the one field that mattered.
Written out together, the pattern is hard to miss. Ryvx uses the fields it controls. It does not tamper with them. And in the first case our own prompt file contains a paragraph describing precisely the failure that occurred, including the phrase cleverness spent before cheapness was tried. The guidance was there. It was correct. It was ignored.
What we were actually telling our agents
Once we started checking prompts against reality rather than against our memory of them, the picture got worse. The tool catalogue injected into every reconnaissance and hunting agent advertised eight tools that are not installed in the sandbox and never have been. It also stayed silent about four that are. A skill file served to exactly the agents meant to do content discovery told them the shell was unavailable on URL-only targets, which stopped being true in August when sandbox support landed. And a gate designed to refuse premature completion ended its own refusal message by telling the model that finish would be honoured next time regardless, which a live run took as the permission it plainly was.
None of that is a model problem. Every one of those is us, telling an agent something false or something optional, and then being disappointed when it behaved accordingly.
The fixes are boring, which is the point
We stopped writing guidance and started writing mechanisms. The deterministic pre-scan pass, which runs before any agent makes a decision, went from probing eleven well-known paths to thirty, and every one of the new entries is a directory or backup artifact rather than a config leak. Finding /ftp no longer depends on a model guessing. A gate now reads the scan's own traffic log at the end of a run, and if a confirm-field pair was submitted repeatedly and never once submitted mismatched, it refuses to finish and says so, naming the endpoint and the count. Discovered URLs that nothing ever fetched now get a plain GET automatically, with the response handed to the coordinating agent, because fetching a URL you have already found needs no agent and no decision. And the catalogue now lists what is actually installed.
We also added the missing vulnerability class to the checklist, since it had no coverage anywhere in the codebase. That one is worth calling out honestly, because we predicted it would be worthless. The reasoning was sound: the checklist is only served on request, and we had measured that request happening zero times. On the next run the model requested it, read it, and went and tampered the registration. The prose worked. It worked because that particular model happens to ask, which is exactly the coin flip the mechanisms exist to remove.
Eight out of eight
The run after those changes scored 8 of 8 on Juice Shop. Harness scored, not read off a scoreboard by hand: run status scored, not tainted, root agent completed normally, every figure taken from the run's own artifacts. It filed 17 findings alongside the eight challenges and cost 37 cents in tokens. Repetitive Registration, which had never been solved in any run on any model we have tried, came in. So did Confidential Document and DOM XSS.
For comparison, the same corpus previously scored 6 of 8 on a frontier model at $4.47, and 7 of 8 on our internal engine. The expensive run that started all of this was still sitting at 6 of 8 when we stopped it.
What we are not claiming
This is one run. One run is not a rate, and we are publishing it as a single scored measurement rather than as a recall figure. Juice Shop has real variance between runs, and an earlier run the same evening reached 7 of 8 through a different combination of challenges. We will replace this with a repeated figure once we have several consecutive runs to average, and if that average comes in below 8 we will publish that instead.
We also cannot yet attribute each challenge to each fix. The gate we built specifically for the registration bug never fired on the winning run, so something else got there first. That trace is unfinished, and we would rather say so than tidy it into a cleaner story.
The part we are confident about is the lesson, because it repeated four times in one night without a single counterexample. Every miss we diagnosed was a place where the system asked a model to choose correctly instead of making the correct thing happen. In three of those four cases our own documentation already described the right behaviour in plain language, and that changed nothing at all. Guidance tells a model what good looks like. A mechanism decides what happens. When the two disagree the mechanism wins every time, and if you have not written one, there is no mechanism, there is only hope.
← Back to the blog