Skip to main content

Docs / Reverse Engineering

Reverse Engineering

Hand Ryvx a binary sample and get a triaged, PoC-gated analysis back, without the sample ever touching your network.

What it does, and what it runs on

This capability analyses a single potentially hostile binary: a suspected malware sample, an unlabelled executable from an incident, a crackme. It is deliberately separate infrastructure from the rest of Ryvx, which sandboxes run_shell against a cooperative pentest target inside a Docker container. Here the payload itself is the threat, so a container isn't the right boundary: the isolation is a hardware-backed microVM (Firecracker/QEMU on Linux, QEMU (WHPX) on Windows, QEMU on macOS), one fresh VM per job, always force-destroyed after. The guest has no virtual network device attached at all, not a firewalled one: on Linux the QEMU backend passes an explicit -nic none, and the Firecracker backend simply never issues the API call that would attach a NIC. A second, independent layer sits inside the guest itself: the tool battery runs in a rootful Podman container started with --network none --cap-drop ALL. The sample has no network path to reach at either layer, regardless of what it does once executed.

The sample's own bytes never leave your machine, with or without any flag. What can leave, and only with your say-so, is sample-derived text: decoded strings, and capa/FLOSS tool output, sent to an LLM for interpretation. That's refused by default unless the model you've configured is local (Ollama, or an API base on 127.0.0.1/localhost) or you explicitly pass --allow-egress. A second, optional channel is a reputation lookup (MalwareBazaar/VirusTotal/YARAify): only the sample's SHA-256 is ever sent, never the sample.

The tiers

Three tiers, all free. Which one you pick is the core decision each time you run one: how much of the sample do you need dissected.

  • Tier 1: static triage (re_triage). Runs the fixed tool battery inside the guest: file, strings, radare2 auto-analysis, capa's MITRE-mapped capability detection, and floss's obfuscated-string recovery, and interprets the output: an executive summary, behaviour assessment, recommended actions, IOCs, and a MITRE ATT&CK mapping. Fastest tier, and the one every other tier runs on top of. Pick this when you just need a fast read on an unknown sample.
  • Tier 2: solve (re_solve). The same battery, then an attempt to find the sample's valid input, built for crackme/license-check-style binaries. A free automated pass first (angr, symbolic execution), escalating to an LLM-driven interactive pass only if that fails. Neither pass trusts its own claim of success: a "solved" verdict requires the process to exit 0 and a specific success string, named before the run, to appear in real stdout; a bare exit code is not evidence, since most crackmes print "Wrong" and still return 0. Pick this when you have a binary that's gating on some input and you want to know what input passes.
  • Tier 3: decompile (re_decompile). The same battery plus a full Ghidra headless decompilation of every function, in its own sandbox session. Ghidra's raw output is dominated by CRT/runtime startup boilerplate the compiler pulled in, not the payload, so it isn't pasted whole into the model's context: a separate, deterministic, offline ranking pass scores every function with fixed rules a human can read and audit, and only the functions that look interesting get interpreted. Slowest tier: a large binary can take up to ~40 minutes. Pick this when you need to read decompiled source, not just a behavioural summary.

Running it

Everything here runs on your own machine, in the isolated microVM described above: free, no account or credits required.

python -m ryvx re-analyze /path/to/sample.bin --tier 1
python -m ryvx re-analyze /path/to/crackme.bin --tier 2 --mode solve
python -m ryvx re-analyze /path/to/sample.bin --tier 3 --check-reputation

--tier selects depth (1/2/3, matching the tiers above); --mode (quick/full/solve) selects how the tool output is interpreted: --tier 2 alone already implies --mode solve, but --tier 3 does not imply solve mode; pass both if you want a decompile and a solve attempt together. A sample already on disk is the positional argument; a remote one can be fetched straight into memory instead with --sample-url <https URL> (optionally pinned with --expect-sha256); it's never written to this machine's disk either way. --zip-password infected extracts a ZIP-wrapped sample the way MalwareBazaar and vx-underground both serve them; without it a ZIP-wrapped sample is analysed as a ZIP, not unwrapped. --allow-egress permits sample-derived tool output to reach a non-local model, off by default, and the command prints exactly what would leave the machine before anything runs, no flag required to find out.

What you get back

Every tier produces a report with the same shell: an executive summary/behaviour assessment (full mode) or a fast verdict and one-line reasoning (quick mode); extracted IOCs, split into two buckets: deterministic ones regex-extracted from the sample's own captured strings, and separate model-suggested ones marked explicitly "unverified"; a MITRE ATT&CK tag list; a reputation section per configured provider (a miss there is never treated as proof of benignity, only "not in this database"); and a raw, per-tool output panel you can expand. A generated YARA rule is both copyable and downloadable as rule.yar directly from the job page. Tier 2 adds a Solve section (solved/not solved, the candidate input if one was found, transcript turn count). Tier 3 adds a Decompile section: function count, an in-page excerpt of the decompiled C (capped at 4,000 characters for display; the full text, up to 100,000 characters, is in the downloadable report.json), and a flag if a sample turned out to be .NET, since native decompilation only ever covers a managed assembly's native bootstrap stub, not the real IL payload, so that flag matters more than the decompile output itself when it's set. Both report.json and rule.yar are downloadable from the finished job page.

The guest image

ryvx re-analyze needs a Debian bookworm root filesystem carrying the static-analysis battery, Ghidra, and angr, plus a kernel and initrd to boot it under QEMU. No prebuilt image is published right now: the earlier release was withdrawn. ryvx setup says so and prints the steps to build the image locally. If you point RE_VM_IMAGE_MANIFEST_URL at your own build, every part is SHA-256-checked against that manifest before anything is installed, and a mismatch fails closed. The manifest itself is not signature-checked, so this proves the files match the manifest, not who published it.

What the report tells you when it could not see

Static analysis can miss things, and the dangerous failure is not a false alarm, it is false reassurance. Four checks run over every report and print above the summary, so a conclusion you should not lean on says so before you read it.

  • No verdict was returned. If the model wrote an analysis but never answered whether the sample is benign, suspicious or malicious, the report says so and none is inferred on its behalf. Not judging something and judging it clean are different claims.
  • Part of the analysis did not run. When a tool reports that it does not support this file type, an absence of findings from it is not evidence of absence. FLOSS does not deobfuscate .NET strings and says so; Ghidra decompiles native code, so on a managed assembly it reads the loader stub rather than the logic. Both are named on the report, along with how much of the decompilation was returned for interpretation.
  • An independent source disagrees. With the reputation check enabled, a reassuring verdict that a third-party source contradicts is withdrawn to insufficient evidence. It is never restated as malicious: that would be asserting another vendor's judgement as our own finding.
  • The verdict cites nothing. A suspicious or malicious label with no IOC, MITRE technique, family indicator or anti-analysis observation behind it is withdrawn the same way, in the other direction. The gate runs both ways.

The reputation check is off by default, because a hash lookup tells a third party that someone holds that exact file, which matters for an internal or proprietary binary. Worth enabling for anything you suspect is real malware. When it is off, the report says it was not checked rather than staying silent about it.

Limits

  • Tier 2 (re_solve) has run on the production box and solved: a MinGW-w64 crackme, recovered password at tier 2a with no model escalation, 2026-09-05. That is one sample of one shape. It is no longer untested, and it is not yet a rate: treat a solve on a binary unlike that one as unproven until you have verified it yourself.
  • A .NET sample's decompile result is close to useless without the is_dotnet flag: native decompilation of a managed assembly covers only its native bootstrap stub, not the actual IL payload it loads.
  • Decompiled C is truncated at 100,000 characters in the report; a very large or heavily-inlined binary's full output can exceed that.

Next

  • Verification: the same "PoC or it didn't happen" discipline Tier 2's solve verdict is built on.
  • Download: the desktop app and CLI. The guest image is not published there at the moment.

← Back to Docs
STAY IN THE LOOP

Release notes and product updates, by email.

We'll send a confirmation email; you're not on the list until you click the link in it. See our privacy policy.