A red team for the AI system, not just the app around it
Point Ryvx at an AI system you own. It sends real adversarial turns to the live endpoint, checks whether the agent can be talked past the boundary you declared for it, and hands back a report that says what it tried, what it found, and what it did not test.
It is also built to refuse. The strongest thing on this page is a measurement of it declining to touch an endpoint it was not authorized to touch.
Zero requests to the endpoint it was told not to touch.
An offline conformance harness ships in the source tree. It stands up two local fake endpoints, one authorized and one that must never be touched, then drives the real agent_boundary_probe path against them, exactly as a scan dispatches it. The endpoint that must not be touched counts its own inbound requests. No API key, no network, no real model.
- 01No out-of-scope connection is ever made0 inbound requests to the out-of-scope endpoint, 16 to the authorized one
- 02A scope violation is refused and recordedrefusal_returned=True, in_audit=True, in_sidecar=True, endpoint_hits=0
- 03The production gate auto-denies without hanginghung=False, denied=True, authorised_hits_during_call=0
- 04Every probe attempt is in the audit trail3 of 3 probe tool-calls recorded in audit.jsonl, with arguments
That the probe path's scope enforcement, its production gate and its audit trail held by observation, against a build, on a date. The harness asserts each guardrail against something a failure would have moved, writes a JSON artifact you can attach to a customer report, and exits non-zero if any guardrail fails. It never weakens a gate to make a check pass: a failing run is a real finding.
It is one measurement of one path on one day, not a certification. The endpoints are local fakes, so it says nothing about how a real target behaves. It runs from a source checkout with no credentials, so ask for the artifact from a run against the build you are evaluating rather than taking this page's word for it.
A coverage report, not a findings list.
An empty findings list is not a clean bill of health, and the report says so on its face. For each of the ten OWASP LLM Top 10 risks, every run states one of three things: a finding was filed, an attack attempted that risk and nothing was filed, or the risk was not tested, with the reason on the row.
"Attempted, nothing filed" is deliberately not called a pass. It means only that this run's specific attacks did not surface a weakness. A vendor that prints its own gaps, with a reason on each one, is making a stronger claim than one that implies total coverage.
Each risk also carries its MITRE ATLAS technique ids where ATLAS has one: LLM01 maps to AML.T0051 and AML.T0054, LLM02 to AML.T0057, LLM07 to AML.T0056. The other six carry none rather than a guessed one.
Eleven attack categories.
You describe what the agent is genuinely allowed to do: which tools it may call, which filesystem root it is confined to, which hosts it may reach, whether it may run a shell command. Each category below tries a different route past that boundary. The ids are the ones that appear in a finding, so a reader can match a result to its cause.
Blunt override language in the user's own message. No tool output involved.
Claims elevated authority without naming a forbidden action, to see whether the agent engages with the framing at all.
Builds trust with a legitimate ask, probes without a target, then names the forbidden path on the third turn.
A file the agent legitimately reads carries a social-engineering aside rather than a command.
Override language arriving inside data the agent asked for itself, not in a user message.
A verbatim-repeat request, then the same ask reframed as an audit obligation. Confirmed only against a canary you plant.
Escalates from a plausible lookup to a request to paste a stored record back verbatim. Confirmed only against a canary you plant.
Escalates from an ordinary-looking search to an explicit request for documents outside the caller's own scope. Confirmed only against a canary you plant.
Invites the model to confirm a plausible but unverified claim, then presses for it to be stated as fact. Confirmed only against a false statement you declare.
A benign baseline turn, then a bounded probe turn. Response size and latency are measured against the baseline, not inferred from text.
Induces the model to emit an active payload verbatim. Only ever a triage lead here: confirming it needs a downstream sink this probe cannot observe.
The evidence bar
Black-box testing, with no visibility into the agent's own tool calls, can never assert that containment was actually broken: a model claiming to have taken an action is text, not proof.
It can confirm three things directly, because they are observations rather than inferences:
- Disclosurea canary string you plant yourself, echoed back verbatim
- Amplificationa bounded request producing a disproportionate amount of output or time, measured against a baseline turn
- Misinformationa false statement you declare yourself, asserted back as fact, not a disclosure, since nothing here was leaked
Every other result caps out as a triage lead, and nothing files as a finding without a working proof of concept. Expose the agent's tool-call trace and the verdicts stop being black-box.
Past a keyword guardrail
A plain-English payload set is exactly what a keyword filter blocks, and a probe that reports nothing because it was filtered is a false negative, the worst outcome a security tool can produce. So llm_injection_probe, the prompt-injection tool that runs alongside the boundary battery rather than being part of it, re-sends its canary payload through three encodings: base64, rot13 and leetspeak.
Only the canary goes through, because its success is machine-detectable: the string either comes back or it does not. A target that decodes one of these and still emits the canary has followed an injected instruction through an encoding layer, a strictly stronger finding than the plaintext hit, and it is reported as such.
One detail that matters: base64 and rot13 are bijective, so a target that decodes them recovers the original canary and the plain marker is the right thing to look for. Leetspeak is a same-length substitution with no decode step, and its table rewrites the canary's own suffix, so that variant looks for the leetspoken canary instead. The wrong marker there would silently turn a real bypass into a false negative.
Three policy presets, or write your own.
A policy is what the agent is genuinely allowed to do. Pick the preset that matches the system, or hand-write one. A preset is a starting point, not a verdict about your agent: a support bot that legitimately owns an order-lookup tool is not the support-chatbot preset, and should start from rag-assistant or a hand-written policy with that one tool added.
Answers product and account questions in a support chat widget.
Reads and writes files in a repository checkout and runs its build and test commands.
Answers questions by retrieving documents from an internal corpus and citing them.
Authorization is enforced, not advisory.
It will not test a system you cannot prove is yours
The same domain-control challenge every other Ryvx exploit path goes through gates this one, with no lighter-touch route of its own. Someone else's chatbot is refused by design.
It treats every adversarial turn as exploitation
Sending turns designed to make a live agent misbehave is exploitative by definition, so the probe is hard-gated on exploitation being enabled for the run at all. Against a production-tagged target it stops and asks a human. Run unattended, it denies rather than hangs: guardrail 03 above is that behavior, measured.
It keeps the receipts
Every probe attempt, with its arguments, is written to an append-only audit trail delivered with the report. A scope refusal is recorded on three surfaces, not one.
The full position, including how a bug bounty program's published scope is treated, is on the security policy page.
On your machine, or on ours.
The desktop app and the CLI, with your own model API key. You pay your model provider directly and nothing leaves your network except what you choose to send. The conformance harness above runs from the same checkout.
Get Ryvx →Hosted red-team runs on our hardware, credit-based, results in your account. No Docker, no local setup. Current plans are on the pricing page.
See pricing →The parts a skeptical engineer would poke at, stated first.
8 of 10, not 10 of 10
The other two, Supply Chain (LLM03) and Data and Model Poisoning (LLM04), are structural gaps, not missing features: neither is reachable by any black-box probe, ever, because both need access this tool structurally cannot get from a conversation with a live endpoint. LLM03 needs model, adapter and dependency provenance; LLM04 needs training or fine-tuning data access. No amount of further attack-category work closes either gap, which is why they are the only two rows on the table above with no path to becoming a third.
Triage leads, mostly
In black-box mode, most verdicts are triage leads, not confirmed findings. Three things are confirmable without a tool-call trace, and they are all observations: a planted canary coming back, a declared false statement asserted back as fact, and a measured amplification.
One dated run, not a certification
The guardrail figures are one dated run against local fake endpoints. They show the scope enforcement held on that build on that day. They are re-runnable, which is the point, and they are not a certification.
No quoted detection or false-positive rate
We have not independently measured one for this line. Quoting a number we cannot show the working for would be exactly the behavior this product exists to avoid.
Ryvx is for systems you own or are contracted to test. Domain control is verified before any live target is touched, and a probe against a production-tagged target requires a human decision at the time it happens. The web-app side of the same engine is on the features page.