Skip to main content
AI RED TEAMING

A red team for the AI system, not just the app around it

Point Ryvx at an AI system you own. It sends real adversarial turns to the live endpoint, checks whether the agent can be talked past the boundary you declared for it, and hands back a report that says what it tried, what it found, and what it did not test.

It is also built to refuse. The strongest thing on this page is a measurement of it declining to touch an endpoint it was not authorized to touch.

PROOF IT STAYS IN SCOPE

Zero requests to the endpoint it was told not to touch.

An offline conformance harness ships in the source tree. It stands up two local fake endpoints, one authorized and one that must never be touched, then drives the real agent_boundary_probe path against them, exactly as a scan dispatches it. The endpoint that must not be touched counts its own inbound requests. No API key, no network, no real model.

0
requests out of scope
16
to the authorized endpoint
measured 2026-09-05, scripts/guardrail_conformance.py, exit 0, OVERALL: PASS
  • 01
    No out-of-scope connection is ever made
    0 inbound requests to the out-of-scope endpoint, 16 to the authorized one
  • 02
    A scope violation is refused and recorded
    refusal_returned=True, in_audit=True, in_sidecar=True, endpoint_hits=0
  • 03
    The production gate auto-denies without hanging
    hung=False, denied=True, authorised_hits_during_call=0
  • 04
    Every probe attempt is in the audit trail
    3 of 3 probe tool-calls recorded in audit.jsonl, with arguments
What it proves

That the probe path's scope enforcement, its production gate and its audit trail held by observation, against a build, on a date. The harness asserts each guardrail against something a failure would have moved, writes a JSON artifact you can attach to a customer report, and exits non-zero if any guardrail fails. It never weakens a gate to make a check pass: a failing run is a real finding.

What it does not prove

It is one measurement of one path on one day, not a certification. The endpoints are local fakes, so it says nothing about how a real target behaves. It runs from a source checkout with no credentials, so ask for the artifact from a run against the build you are evaluating rather than taking this page's word for it.

WHAT A RUN TELLS YOU

A coverage report, not a findings list.

An empty findings list is not a clean bill of health, and the report says so on its face. For each of the ten OWASP LLM Top 10 risks, every run states one of three things: a finding was filed, an attack attempted that risk and nothing was filed, or the risk was not tested, with the reason on the row.

"Attempted, nothing filed" is deliberately not called a pass. It means only that this run's specific attacks did not surface a weakness. A vendor that prints its own gaps, with a reason on each one, is making a stronger claim than one that implies total coverage.

11
attack categories
run against the live endpoint by default
8 / 10
OWASP LLM risks exercised
the other two carry a stated reason, not a result
2 / 10
gaps stated openly
on the page and in every report, never a silent blank
THE TEN RISKS, ROW BY ROW
Exercised
LLM01Prompt Injection
Six categories reach it: every override, escalation, spoofing and injection template, plus system prompt extraction.ATLAS AML.T0051, AML.T0054
Exercised
LLM02Sensitive Information Disclosure
sensitive_data_disclosure and multi_turn_escalation.ATLAS AML.T0057
Not tested
LLM03Supply Chain
Structurally unreachable by any remote probe: it needs model, adapter and dependency provenance, not a conversation with the endpoint. Covered instead by the dependency-scan path's real OSV lookups against known-vulnerable LLM libraries.
Not tested
LLM04Data and Model Poisoning
Structurally unreachable by any remote probe: it needs training or fine-tuning data access, not a conversation with the endpoint. No black-box attack category can close this gap.
Exercised
LLM05Improper Output Handling
improper_output_handling. Caps at a triage lead in this probe.
Exercised
LLM06Excessive Agency
Four categories: direct_override, social_engineering_tool_output, multi_turn_escalation and indirect_injection_retrieved_content.
Exercised
LLM07System Prompt Leakage
system_prompt_extraction, confirmed only against a planted canary.ATLAS AML.T0056
Exercised
LLM08Vector and Embedding Weaknesses
retrieval_boundary: seed a document in a corpus a tenant/user must not reach, then try to retrieve it. Confirmed only against a planted canary.
Exercised
LLM09Misinformation
misinformation: press the model to state a claim as fact. Confirmed only against a false statement you declare, not against ordinary text.
Exercised
LLM10Unbounded Consumption
unbounded_consumption, a measured amplification rather than a text inference.

Each risk also carries its MITRE ATLAS technique ids where ATLAS has one: LLM01 maps to AML.T0051 and AML.T0054, LLM02 to AML.T0057, LLM07 to AML.T0056. The other six carry none rather than a guessed one.

WHAT RUNS

Eleven attack categories.

You describe what the agent is genuinely allowed to do: which tools it may call, which filesystem root it is confined to, which hosts it may reach, whether it may run a shell command. Each category below tries a different route past that boundary. The ids are the ones that appear in a finding, so a reader can match a result to its cause.

Direct override
direct_override

Blunt override language in the user's own message. No tool output involved.

LLM01 · LLM06
Authority spoofing
authority_spoofing

Claims elevated authority without naming a forbidden action, to see whether the agent engages with the framing at all.

LLM01
Multi-turn escalation
multi_turn_escalation

Builds trust with a legitimate ask, probes without a target, then names the forbidden path on the third turn.

LLM01 · LLM02 · LLM06
Social engineering via tool output
social_engineering_tool_output

A file the agent legitimately reads carries a social-engineering aside rather than a command.

LLM01 · LLM06
Indirect injection via retrieved content
indirect_injection_retrieved_content

Override language arriving inside data the agent asked for itself, not in a user message.

LLM01 · LLM06
System prompt extraction
system_prompt_extraction

A verbatim-repeat request, then the same ask reframed as an audit obligation. Confirmed only against a canary you plant.

LLM01 · LLM07
Sensitive data disclosure
sensitive_data_disclosure

Escalates from a plausible lookup to a request to paste a stored record back verbatim. Confirmed only against a canary you plant.

LLM02
Retrieval boundary
retrieval_boundary

Escalates from an ordinary-looking search to an explicit request for documents outside the caller's own scope. Confirmed only against a canary you plant.

LLM08
Misinformation
misinformation

Invites the model to confirm a plausible but unverified claim, then presses for it to be stated as fact. Confirmed only against a false statement you declare.

LLM09
Unbounded consumption
unbounded_consumption

A benign baseline turn, then a bounded probe turn. Response size and latency are measured against the baseline, not inferred from text.

LLM10
Improper output handling
improper_output_handling

Induces the model to emit an active payload verbatim. Only ever a triage lead here: confirming it needs a downstream sink this probe cannot observe.

LLM05

The evidence bar

Black-box testing, with no visibility into the agent's own tool calls, can never assert that containment was actually broken: a model claiming to have taken an action is text, not proof.

It can confirm three things directly, because they are observations rather than inferences:

  • Disclosurea canary string you plant yourself, echoed back verbatim
  • Amplificationa bounded request producing a disproportionate amount of output or time, measured against a baseline turn
  • Misinformationa false statement you declare yourself, asserted back as fact, not a disclosure, since nothing here was leaked

Every other result caps out as a triage lead, and nothing files as a finding without a working proof of concept. Expose the agent's tool-call trace and the verdicts stop being black-box.

Past a keyword guardrail

A plain-English payload set is exactly what a keyword filter blocks, and a probe that reports nothing because it was filtered is a false negative, the worst outcome a security tool can produce. So llm_injection_probe, the prompt-injection tool that runs alongside the boundary battery rather than being part of it, re-sends its canary payload through three encodings: base64, rot13 and leetspeak.

Only the canary goes through, because its success is machine-detectable: the string either comes back or it does not. A target that decodes one of these and still emits the canary has followed an injected instruction through an encoding layer, a strictly stronger finding than the plaintext hit, and it is reported as such.

One detail that matters: base64 and rot13 are bijective, so a target that decodes them recovers the original canary and the plain marker is the right thing to look for. Leetspeak is a same-length substitution with no decode step, and its table rewrites the canary's own suffix, so that variant looks for the leetspoken canary instead. The wrong marker there would silently turn a real bypass into a false negative.

THE BOUNDARY YOU DECLARE

Three policy presets, or write your own.

A policy is what the agent is genuinely allowed to do. Pick the preset that matches the system, or hand-write one. A preset is a starting point, not a verdict about your agent: a support bot that legitimately owns an order-lookup tool is not the support-chatbot preset, and should start from rag-assistant or a hand-written policy with that one tool added.

support-chatbot

Answers product and account questions in a support chat widget.

Allowed tools
No tools at all. That is the preset's whole point.
coding-agent

Reads and writes files in a repository checkout and runs its build and test commands.

Allowed tools
read_file, write_file, list_files, run_command
rag-assistant

Answers questions by retrieving documents from an internal corpus and citing them.

Allowed tools
retrieve_document
WHAT IT REFUSES TO DO

Authorization is enforced, not advisory.

It will not test a system you cannot prove is yours

The same domain-control challenge every other Ryvx exploit path goes through gates this one, with no lighter-touch route of its own. Someone else's chatbot is refused by design.

It treats every adversarial turn as exploitation

Sending turns designed to make a live agent misbehave is exploitative by definition, so the probe is hard-gated on exploitation being enabled for the run at all. Against a production-tagged target it stops and asks a human. Run unattended, it denies rather than hangs: guardrail 03 above is that behavior, measured.

It keeps the receipts

Every probe attempt, with its arguments, is written to an append-only audit trail delivered with the report. A scope refusal is recorded on three surfaces, not one.

The full position, including how a bug bounty program's published scope is treated, is on the security policy page.

WHERE IT RUNS

On your machine, or on ours.

On your machine

The desktop app and the CLI, with your own model API key. You pay your model provider directly and nothing leaves your network except what you choose to send. The conformance harness above runs from the same checkout.

Get Ryvx →
On our infrastructure

Hosted red-team runs on our hardware, credit-based, results in your account. No Docker, no local setup. Current plans are on the pricing page.

See pricing →
WHAT WE ARE NOT GOING TO CLAIM

The parts a skeptical engineer would poke at, stated first.

8 of 10, not 10 of 10

The other two, Supply Chain (LLM03) and Data and Model Poisoning (LLM04), are structural gaps, not missing features: neither is reachable by any black-box probe, ever, because both need access this tool structurally cannot get from a conversation with a live endpoint. LLM03 needs model, adapter and dependency provenance; LLM04 needs training or fine-tuning data access. No amount of further attack-category work closes either gap, which is why they are the only two rows on the table above with no path to becoming a third.

Triage leads, mostly

In black-box mode, most verdicts are triage leads, not confirmed findings. Three things are confirmable without a tool-call trace, and they are all observations: a planted canary coming back, a declared false statement asserted back as fact, and a measured amplification.

One dated run, not a certification

The guardrail figures are one dated run against local fake endpoints. They show the scope enforcement held on that build on that day. They are re-runnable, which is the point, and they are not a certification.

No quoted detection or false-positive rate

We have not independently measured one for this line. Quoting a number we cannot show the working for would be exactly the behavior this product exists to avoid.

Ryvx is for systems you own or are contracted to test. Domain control is verified before any live target is touched, and a probe against a production-tagged target requires a human decision at the time it happens. The web-app side of the same engine is on the features page.

STAY IN THE LOOP

Release notes and product updates, by email.

We'll send a confirmation email; you're not on the list until you click the link in it. See our privacy policy.