The bugs that only showed up when it ran for real
2026-08-28
The GitHub remediation-PR bot has been on the pricing page for weeks. Nothing had ever run it end to end against a real repository. Every test around it passed: the diff-application logic, the Markdown-comment rendering, the finding-eligibility predicate, all of it. Running it for real took four attempts. Three of the four failures were distinct bugs, and every one of them was in code that had never been exercised anywhere the bug was possible.
Four attempts
Attempt one cloned the wrong branch. A bare `git clone` checks out the remote's default branch (whatever HEAD points to), not the job's spec.base. The fix branch was then created from wherever that landed, while the PR itself targeted spec.base. On a repo whose default branch differs from the base the job was given, that produces a bloated diff or a patch applied against the wrong version of a file. The fix was one flag: `git clone --branch base`. It had survived until now because the only other caller of this code, the local `ryvx fix-pr` CLI path, never clones anything, it operates on a developer's own pre-existing checkout, which is already on the right commit by construction. The hosted path is the one caller that clones fresh, and it was also the one that had never run.
Attempt two failed to commit at all. git refuses to commit with no identity configured: "unable to auto-detect email address", exit 128. The hosted worker clones into a brand-new container with no global git config in it. A developer's own machine always has one, so the local path had never hit this either. The fix passes identity per-command, `-c user.name=Ryvx -c user.email=ryvx-bot@users.noreply.github.com`, rather than writing it with `git config`, nothing persists in a checkout that belongs to a customer.
Attempt three failed with no usable explanation. `subprocess.run(check=True, capture_output=True)` raises a `CalledProcessError` whose message is only "returned non-zero exit status 128", git's real explanation sits in `.stderr` and nothing ever printed it. The actual cause was mundane: the customer's fine-grained token was missing Contents: write, and GitHub says exactly that ("Resource not accessible by personal access token") if anything reads its response. Nobody did. That one cost two extra rounds by itself, chasing an exit code instead of a sentence, before the git wrapper was rewritten to raise git's own stderr in the exception message.
Attempt four worked, once the token had write access and the failures ahead of it had somewhere to be seen.
The database moved and the interface didn't
The remediation bot wasn't the only place this happened this week, and the pattern is the actual point. A migration removed the Pro-plan entitlement gate behind the bot's two usage caps, but the account page's connect-a-GitHub-token form was still hidden behind `planTier !== "pro"`, showing an upgrade prompt in its place. With every org now on the free plan by default, the form could not be opened by anyone, on any plan, at all. A second migration extended `--skip-verification-i-own-this` from local-only to hosted scans; `/security` and `/docs/authorization` still described it as "local/dev use only" after that shipped. And the one receipt a customer gets when a hosted scan finishes reads its severity counts from `public.findings`, a table the RE tiers and the fix-PR job never write a row to, by design; their output lands elsewhere (`re_runs/{run}/report.json`, findings already indexed by the scan that produced them). An empty count rendered as "Findings: none," so a customer who had just paid 25 credits for a Tier 2 reverse-engineering solve got an email telling them their scan found nothing, while the actual report sat in storage the whole time.
None of these three needed a new bug to exist. The schema and the entitlement model had already changed underneath them. Nothing about the UI or the email was wrong on the day it was written, it just stopped being true the moment the migration landed, silently, with no test positioned to notice, because nothing had told the tests the ground had moved.
A verification that couldn't fail
The sharpest version of this happened somewhere nothing was even broken yet. A change to the Content-Security-Policy dropped `'unsafe-inline'` from `script-src` in favour of per-page build-time hashes, the standard hardening move. It was verified two ways: the hashes were checked against the served bytes, and the page was loaded in a browser. Both passed. Fourteen inline scripts on `/pricing`, fourteen hashes in the header, none missing, none extra. The page rendered perfectly.
It was still completely broken. Next's App Router injects further inline scripts at runtime, during hydration, they don't exist at build time, so no build-time hash can cover them. `script-src 'self'` plus fourteen correct hashes blocked those, which meant React never hydrated, which meant every page on the site shipped with no JavaScript at all. And because the HTML is fully prerendered on a static export, a page with zero JavaScript looks exactly like a page with working JavaScript, right up until something on it needs to move. It was caught because a nav-bar scroll animation had stopped condensing, not because any check said no.
That's the case worth sitting with. The two checks that were run (do the hashes match? does the page render?) were both real, both passed honestly, and neither one was capable of detecting the failure they existed to catch. A verification that structurally cannot fail isn't a weaker verification. It isn't a verification.
What this says
Every bug above passed its tests, because its tests were written against the code as it was written, not against the thing it had to be true. A test confirms the code does what you told it to do. It says nothing about whether what you told it to do was the right thing to tell it, and it says nothing at all about a caller, a container, or a migration that the test suite doesn't know exists yet. The clone-branch bug, the missing identity, the swallowed stderr, all three lived in code path that unit tests exercised correctly and that had simply never been run anywhere the failure could occur. The gate, the docs, the email, all three were correct descriptions of a system that had since changed underneath them. And the CSP fix is the case that makes the whole thing hardest to wave away, because it wasn't skipped or undertested, it was checked by exactly the two methods that sounded sufficient, and neither could have caught it, by construction.
None of this means the tests were pointless, every one of them still caught what it was built to catch. It means passing them was never the question that mattered. The question was always whether the thing had been run for real, and for four bugs this week, in code that had shipped and sat unrun, the answer had quietly been no the whole time.
← Back to the blog