Triage a suspicious binary without it ever touching your disk
Hand Ryvx a sample and it runs the static battery inside a hardware-backed VM the file cannot escape, then writes a report an analyst can hand a client: identification, a capability read, IOCs, a starter YARA rule, and hash reputation.
The strongest thing on this page is not a capability. It is the machinery that made the tool withdraw its own conclusion the day it was wrong.
A false alarm costs an analyst an hour. False reassurance costs them the incident.
A malware tool has two ways to fail, and they are not equally expensive. Crying wolf wastes an hour. Telling an analyst a live sample is clean when it is not loses the incident. So this report is gated in both directions, and the gate that matters most exists because Ryvx made exactly that mistake and now catches it.
Sample 699ec052... was analyzed at Tier 3, full mode, twice. Both reports concluded it appeared to be a legitimate .NET application with no suspicious indicators. VirusTotal flags that same hash with 59 of 74 engines. The reputation lookup had already run, and its answer sat in report.json, unread by anything that formed a conclusion.
- 01Reputation contradictiongate_reputation_contradictionWhen an independent source says malicious and this analysis said nothing of the kind, it says so loudly, at the top, before anyone reads the reassuring part. The verdict is withdrawn to insufficient_evidence, never flipped to malicious on someone else's say-so.
- 02Unsupported verdict_gate_unsupported_verdictThe other direction, and older: a suspicious or malicious label the analysis cannot support is withdrawn. Both directions are gated, because a tool that only guards against crying wolf is guarding the cheaper failure.
- 03Coverage warningstool_coverage_warningsSurfaces tools that said, in their own output, that they could not really look at this file. An absence-of-findings conclusion drawn on top of a tool that opted out is not evidence of absence, and the report says so where a reader will see it.
- 04Withdrawal threshold_REPUTATION_WITHDRAW_MIN_ENGINES = 5A threshold, so no single engine flips a verdict: five independent engines before a reputation source is treated as contradicting. A MalwareBazaar hit withdraws regardless of count, because getting onto a malware repository is a categorically different signal from one engine's heuristic firing.
When an independent source contradicts a reassuring verdict, the gate withdraws that verdict to insufficient_evidence and states the contradiction at the top of the report, above the part a hurried reader trusts. It is the same safe direction the unsupported-verdict gate already chose: withdraw the claim without asserting its opposite.
It never asserts malicious on reputation alone. A VirusTotal score is other vendors' judgment, not this tool's analysis, and this product's whole stance is that a claim needs its own evidence. So a clean verdict is withdrawn, never flipped to malicious, and a verdict that already agrees with reputation is left untouched.
Three tiers, from static triage to decompilation.
Most samples are identified at Tier 1 without ever being run. The deeper tiers exist for the ones that hide behind a check or need their code read directly. Each tier is a real job type, and the id below is the one ryvx re-analyze runs internally.
The static tool battery, no execution. Raw strings, FLOSS-decoded strings, radare2 auto-analysis and capa capability detection, run against the sample inside the VM. Enough to fingerprint a sample and file its IOCs without ever executing it. No detection rate is claimed: the sample this page opens with was not caught by this battery, which is why the gates above exist.
A symbolic solve pass with angr, and a Wine path for Windows binaries, for samples where the answer is behind a check the static battery cannot see through on its own.
Ghidra decompilation for the samples that need the code read, not just its surface described. This is also where the managed-.NET limitation below bites, and the report is built to notice.
The Tier 1 battery
A fixed static battery, run identically every time so the output is comparable across samples. Four tools, none of which executes the file:
- stringsraw extracted strings, and FLOSS-decoded strings for the obfuscated ones
- radare2auto-analysis: architecture, sections, imports and the function graph
- capacapability detection: what the binary can do, mapped to named behaviors
The captured tool output, not the sample, is what an LLM ever reads. And because that text can carry C2 domains, embedded credentials and victim data, RE jobs default to a local model: sample-derived text does not leave the machine unless you pass a flag that says it may.
A VM boundary, not a container
The sample runs inside a real QEMU virtual machine, one fresh VM per job, always force-destroyed after. The rootfs is mounted readonly=on and the guest has -nic none: no writable disk to persist to, no network to reach. Hardware acceleration is KVM on Linux or WHPX on Windows.
The sample bytes never touch the host disk. This is a hardware-backed boundary, and it is not optional: the sandbox refuses to run without a hypervisor and has no container-only fallback to weaker isolation. If the box cannot give it a real VM, the job does not silently downgrade, it stops.
A container shares the host kernel. For a file whose entire purpose may be to break out of wherever you put it, that is the wrong boundary, so it is not the one used.
Artifacts you can act on
From the Tier 1 output, Ryvx extracts IOCs and generates a starter YARA rule keyed to the sample's hash and strings. Deterministic, so the same sample yields the same rule, and something you can drop into a detection pipeline rather than a paragraph you have to translate.
A starter rule, not a finished one, and the generator is honest about why that distinction matters: a rule that never matches is merely useless, a rule that matches everything is worse. It is a lead an analyst tightens, not a signature to ship unread.
Hash reputation, kept in its lane
The sample's hash is checked against MalwareBazaar and VirusTotal. This is where the contradiction gate above draws its independent signal: an outside source saying malicious is exactly the input that withdraws a reassuring verdict.
Reputation informs the report, it does not become the verdict. A VirusTotal score can withdraw a clean conclusion; it cannot, on its own, assert a malicious one. The analysis still has to earn that claim itself.
Managed .NET is a known blind spot, stated here, not buried.
The false negative above was not bad luck. It was a managed .NET/CIL assembly, and that file class is genuinely harder for this pipeline. Ghidra decompiles native code, so on a .NET sample it reads the loader stub rather than the managed logic where the behavior lives. On that sample only 2 of 20 decompiled functions came back for interpretation, FLOSS emitted its own disclaimer that it does not deobfuscate .NET strings, and bad-instruction-data warnings were present.
None of that is hidden. The coverage-warnings gate turns each of those signals into a plain-English line in the report, precisely so that "we found nothing" on a managed binary is never mistaken for "there is nothing there." A page that tells a security buyer where the tool is weaker, and how a report would show it, is worth more than one that claims uniform capability.
On your own machine, free.
The desktop app and the CLI, with a local model by default so sample-derived text stays on the box. You supply the hypervisor the sandbox needs, nothing leaves your network except what you choose to send, and the sample itself never leaves your hardware at all. No credits, no plan, no account required for the CLI. For a security team that isn't permitted to upload a sample to a third party, that isn't a workaround, it's the point.
Get Ryvx →The parts a skeptical analyst would poke at, stated first.
Triage, not a replacement
This is automated triage, not a substitute for a reverse engineer. It identifies, extracts and generates leads at speed. It does not replace the analyst who reads the decompilation on the samples that matter.
Managed .NET/CIL is weaker
The page above says so with the specifics. The report surfaces the coverage gap rather than papering over it, but the gap is real.
The YARA rule is a starting point
It needs an analyst's eye before it goes into production. A generated rule shipped unread is the failure the generator's own docstring warns about.
No quoted detection rate
We have not independently measured one. Quoting a number we cannot show the working for would be exactly the behavior the contradiction gate exists to prevent.
Samples are analyzed in a disposable, network-isolated VM and never persisted to the host disk. How the isolation and reporting work in detail is on the reverse-engineering docs, and the wider security stance is on the security policy page.