Skip to main content
REVERSE ENGINEERING

Triage a suspicious binary without it ever touching your disk

Hand Ryvx a sample and it runs the static battery inside a hardware-backed VM the file cannot escape, then writes a report an analyst can hand a client: identification, a capability read, IOCs, a starter YARA rule, and hash reputation.

The strongest thing on this page is not a capability. It is the machinery that made the tool withdraw its own conclusion the day it was wrong.

WHEN THE TOOL WAS WRONG

A false alarm costs an analyst an hour. False reassurance costs them the incident.

A malware tool has two ways to fail, and they are not equally expensive. Crying wolf wastes an hour. Telling an analyst a live sample is clean when it is not loses the incident. So this report is gated in both directions, and the gate that matters most exists because Ryvx made exactly that mistake and now catches it.

Sample 699ec052... was analysed at Tier 3, full mode, twice. Both reports concluded it appeared to be a legitimate .NET application with no suspicious indicators. VirusTotal flags that same hash with 59 of 74 engines. The reputation lookup had already run, and its answer sat in report.json, unread by anything that formed a conclusion.

59 / 74
engines flagged the hash
0
suspicious indicators the reports found
sample 699ec052..., Tier 3 full mode, analysed twice; VirusTotal 59 of 74; the reputation answer was already in report.json
  • 01
    Reputation contradiction
    gate_reputation_contradiction
    When an independent source says malicious and this analysis said nothing of the kind, it says so loudly, at the top, before anyone reads the reassuring part. The verdict is withdrawn to insufficient_evidence, never flipped to malicious on someone else's say-so.
  • 02
    Unsupported verdict
    _gate_unsupported_verdict
    The other direction, and older: a suspicious or malicious label the analysis cannot support is withdrawn. Both directions are gated, because a tool that only guards against crying wolf is guarding the cheaper failure.
  • 03
    Coverage warnings
    tool_coverage_warnings
    Surfaces tools that said, in their own output, that they could not really look at this file. An absence-of-findings conclusion drawn on top of a tool that opted out is not evidence of absence, and the report says so where a reader will see it.
  • 04
    Withdrawal threshold
    _REPUTATION_WITHDRAW_MIN_ENGINES = 5
    A threshold, so no single engine flips a verdict: five independent engines before a reputation source is treated as contradicting. A MalwareBazaar hit withdraws regardless of count, because getting onto a malware repository is a categorically different signal from one engine's heuristic firing.
What the gate does

When an independent source contradicts a reassuring verdict, the gate withdraws that verdict to insufficient_evidence and states the contradiction at the top of the report, above the part a hurried reader trusts. It is the same safe direction the unsupported-verdict gate already chose: withdraw the claim without asserting its opposite.

What it deliberately does not do

It never asserts malicious on reputation alone. A VirusTotal score is other vendors' judgement, not this tool's analysis, and this product's whole stance is that a claim needs its own evidence. So a clean verdict is withdrawn, never flipped to malicious, and a verdict that already agrees with reputation is left untouched.

HOW DEEP TO GO

Three tiers, from static triage to decompilation.

Most samples are identified at Tier 1 without ever being run. The deeper tiers exist for the ones that hide behind a check or need their code read directly. Each tier is a real job type, and the id below is the one ryvx re-analyze runs internally.

Tier 1Static triage
re_triage

The static tool battery, no execution. Raw strings, FLOSS-decoded strings, radare2 auto-analysis and capa capability detection, run against the sample inside the VM. Enough to fingerprint a sample and file its IOCs without ever executing it. No detection rate is claimed: the sample this page opens with was not caught by this battery, which is why the gates above exist.

strings, FLOSS, radare2, capa
Tier 2Solve
re_solve

A symbolic solve pass with angr, and a Wine path for Windows binaries, for samples where the answer is behind a check the static battery cannot see through on its own.

angr, Wine
Tier 3Decompile
re_decompile

Ghidra decompilation for the samples that need the code read, not just its surface described. This is also where the managed-.NET limitation below bites, and the report is built to notice.

Ghidra

The Tier 1 battery

A fixed static battery, run identically every time so the output is comparable across samples. Four tools, none of which executes the file:

  • stringsraw extracted strings, and FLOSS-decoded strings for the obfuscated ones
  • radare2auto-analysis: architecture, sections, imports and the function graph
  • capacapability detection: what the binary can do, mapped to named behaviours

The captured tool output, not the sample, is what an LLM ever reads. And because that text can carry C2 domains, embedded credentials and victim data, RE jobs default to a local model: sample-derived text does not leave the machine unless you pass a flag that says it may.

A VM boundary, not a container

The sample runs inside a real QEMU virtual machine, one fresh VM per job, always force-destroyed after. The rootfs is mounted readonly=on and the guest has -nic none: no writable disk to persist to, no network to reach. Hardware acceleration is KVM on Linux or WHPX on Windows.

The sample bytes never touch the host disk. This is a hardware-backed boundary, and it is not optional: the sandbox refuses to run without a hypervisor and has no container-only fallback to weaker isolation. If the box cannot give it a real VM, the job does not silently downgrade, it stops.

A container shares the host kernel. For a file whose entire purpose may be to break out of wherever you put it, that is the wrong boundary, so it is not the one used.

Artifacts you can act on

From the Tier 1 output, Ryvx extracts IOCs and generates a starter YARA rule keyed to the sample's hash and strings. Deterministic, so the same sample yields the same rule, and something you can drop into a detection pipeline rather than a paragraph you have to translate.

A starter rule, not a finished one, and the generator is honest about why that distinction matters: a rule that never matches is merely useless, a rule that matches everything is worse. It is a lead an analyst tightens, not a signature to ship unread.

Hash reputation, kept in its lane

The sample's hash is checked against MalwareBazaar and VirusTotal. This is where the contradiction gate above draws its independent signal: an outside source saying malicious is exactly the input that withdraws a reassuring verdict.

Reputation informs the report, it does not become the verdict. A VirusTotal score can withdraw a clean conclusion; it cannot, on its own, assert a malicious one. The analysis still has to earn that claim itself.

WHERE IT IS WEAKER, AND HOW YOU WOULD NOTICE

Managed .NET is a known blind spot, stated here, not buried.

The false negative above was not bad luck. It was a managed .NET/CIL assembly, and that file class is genuinely harder for this pipeline. Ghidra decompiles native code, so on a .NET sample it reads the loader stub rather than the managed logic where the behaviour lives. On that sample only 2 of 20 decompiled functions came back for interpretation, FLOSS emitted its own disclaimer that it does not deobfuscate .NET strings, and bad-instruction-data warnings were present.

None of that is hidden. The coverage-warnings gate turns each of those signals into a plain-English line in the report, precisely so that "we found nothing" on a managed binary is never mistaken for "there is nothing there." A page that tells a security buyer where the tool is weaker, and how a report would show it, is worth more than one that claims uniform capability.

WHERE IT RUNS

On your own machine, free.

The desktop app and the CLI, with a local model by default so sample-derived text stays on the box. You supply the hypervisor the sandbox needs, nothing leaves your network except what you choose to send, and the sample itself never leaves your hardware at all. No credits, no plan, no account required for the CLI. For a security team that isn't permitted to upload a sample to a third party, that isn't a workaround, it's the point.

Get Ryvx →
WHAT WE ARE NOT GOING TO CLAIM

The parts a sceptical analyst would poke at, stated first.

Triage, not a replacement

This is automated triage, not a substitute for a reverse engineer. It identifies, extracts and generates leads at speed. It does not replace the analyst who reads the decompilation on the samples that matter.

Managed .NET/CIL is weaker

The page above says so with the specifics. The report surfaces the coverage gap rather than papering over it, but the gap is real.

The YARA rule is a starting point

It needs an analyst's eye before it goes into production. A generated rule shipped unread is the failure the generator's own docstring warns about.

No quoted detection rate

We have not independently measured one. Quoting a number we cannot show the working for would be exactly the behaviour the contradiction gate exists to prevent.

Samples are analysed in a disposable, network-isolated VM and never persisted to the host disk. How the isolation and reporting work in detail is on the reverse-engineering docs, and the wider security stance is on the security policy page.

STAY IN THE LOOP

Release notes and product updates, by email.

We'll send a confirmation email; you're not on the list until you click the link in it. See our privacy policy.