by 2akouwu
Stop your AI from making things up — it proposes, deterministic tools decide, every claim checked against ground truth with evidence. Grounded facts and context survive resets. Reverse engineering is the proving ground. MCP server + CLI.
# Add to your Claude Code skills
git clone https://github.com/2akouwu/reverifyLast scanned: 9/3/2026
{
"issues": [],
"status": "PASSED",
"scannedAt": "2026-09-03T08:34:09.196Z",
"npmAuditRan": true,
"pipAuditRan": true,
"promptInjectionRan": true
}reverify is an open-source ai agents skill for AI coding assistants such as Claude Code, Codex CLI, and ChatGPT, built by 2akouwu. Stop your AI from making things up — it proposes, deterministic tools decide, every claim checked against ground truth with evidence. Grounded facts and context survive resets. Reverse engineering is the proving ground. MCP server + CLI. It has 1,026 GitHub stars.
Yes. reverify passed SkillsLLM's automated security scan — a dependency vulnerability audit plus prompt-injection heuristics — with no high-severity issues. You can read the full report in the Security Report section on this page.
Clone the repository with "git clone https://github.com/2akouwu/reverify" and add it to your Claude Code skills directory (see the Installation section above).
reverify is primarily written in Python. It is open-source under 2akouwu on GitHub, so you can review or fork the full source.
Yes. SkillsLLM lists many other AI Agents skills you can browse and compare side by side. Open the AI Agents category from the badge at the top of this page, or use the Related Skills and comparison links further down to weigh reverify against similar tools.
No comments yet. Be the first to share your thoughts!
⚠️ Third-Party Software Notice
This skill is third-party open-source software developed and hosted independently on GitHub. SkillsLLM is an informational directory and does not control or maintain the underlying repository.
Any security checks, ratings, or warnings displayed by SkillsLLM are automated and limited in scope. They do not constitute a security certification or guarantee that the software is safe, error-free, or free from malicious code, vulnerabilities, compromised dependencies, or prompt-injection risks.
Review the source code, permissions, dependencies, and configuration before installing or running any third-party skill. Use is at your own risk. To the maximum extent permitted by applicable law, SkillsLLM is not liable for losses arising from third-party software.
AI is confident and often wrong: it invents an API, a struct field, an offset, or what a function does, and says it like fact. Reverify makes a deterministic tool the judge — the model proposes a claim, the tool checks it against the actual artifact, and it comes back VERIFIED / REFUTED with evidence. The model never gets to assert a fact on its own.
Two things it does today:
reverify verify, or the MCP
server your agent already talks to).reverify rollover
hands the session off to a file and starts a fresh one, so long tasks don't drift or need
/clear. Works in Claude Code, Codex, Gemini CLI and OpenCode.The hardest place to prove the first point is binary reverse engineering, where hallucination is
worst — so that's where the numbers come from. On 71 real Windows system files the AI's textbook
answer was wrong 97% of the time; reverify caught every one and never accepted a wrong
claim (0 of 71; the same gate runs in CI on Linux and macOS every push, and an independent
aarch64 run found the same) (EXAMPLE.md, BENCHMARK.md;
python benchmarks/prologue_prior.py).
Language models are great at reading code and unreliable at reverse engineering. Ask a model to reconstruct a struct or an algorithm from a binary and it will confidently invent offsets, sizes, and behavior. In binary analysis this hallucination problem is far worse than in source code, and "did the model just make that up?" is the single biggest blocker to using AI for real RE.
Reverify pairs a language model with a deterministic, pure-Python RE toolkit and makes the toolkit the judge. The model proposes; the tools verify. A hypothesis about a structure or an algorithm is only reported once it has been checked against the actual bytes — disassembled, pattern-matched, or executed in the emulator — so the output is grounded in the binary instead of the model's imagination.
pip install "reverify[full]" the toolkit upgrades
itself in place to capstone (disassembly), unicorn (real CPU emulation), lief
(PE/ELF/Mach-O) and Z3 (proofs); pip install "reverify[angr]" adds angr for
function boundaries, the call graph and cross-references. Not installed? It falls back to
the pure-Python core. reverify backends shows what's active.reverify equiv <reference> <candidate> --lang python (or C) runs a
candidate implementation and a reference over shared inputs and checks they agree, so an AI's
rewrite or refactor is tested, not trusted — a refutation comes back with the input and both
outputs. The same rigour, aimed at ordinary source code.Reverify is for authorized reverse engineering — malware analysis, CTF, interoperability research, and software you own or are permitted to analyze. See SECURITY.md.
# Install the CLI + MCP server from PyPI:
pip install reverify # pure-Python core; or "reverify[full]" for capstone+unicorn+lief
reverify auto sample.bin --json
# Or run straight from a checkout — pure standard library, nothing to install:
python reverify/cli.py auto sample.bin --json
python reverify/cli.py parse-pe sample.exe --json
python reverify/cli.py disasm 90505831C0C3 --arch x86_64
This is what the name is about. A claim is any hypothesis about the binary; the
deterministic tools are the judge and hand back VERIFIED, REFUTED, or
INCONCLUSIVE together with the bytes they actually observed:
reverify verify sample.bin --claim '{
"kind": "instructions", "offset": 4096,
"mnemonics": ["push", "mov", "sub"], "note": "function prologue"
}'
# Check a reconstructed routine actually computes what the model claimed:
reverify verify - --claim '{
"kind": "emulate_result", "code": "b805000000b90300000001c8c3",
"arch": "x86", "expect_registers": {"eax": 8}
}'
Claims can be batched from a JSON file (--claims-file claims.json); the CLI exits
non-zero if anything is refuted, so an agent or CI job can gate on a grounded
reconstruction. Claim kinds: bytes_at, u16_at / u32_at / u64_at (typed reads, no
endianness math), pattern_present, string_present, instructions (mnemonics and
optionally operands), emulate_result, behavior_equiv, prove_equiv, protobuf_field,
import_present, export_present, section_present, and the semantic kinds
function_at, calls, references, reachable_from_entry (see
The semantic layer). Offsets are file offsets unless a claim says
"space": "rva" or "va"; the verifier translates through the section table and echoes
all three addresses in the evidence, and a refuted bytes_at reports where the expected
bytes actually are. Set "observe": true (or omit expected) to have the tools read a
value instead of asserting one, and "depends_on": [...] so a refuted root invalidates
the claims built on it.
"Every claim verified" is trivially reachable: assert that the file starts with MZ and
that .text exists. So Reverify also weighs how much a verified set actually says. Each
result carries a weight — zero for claims that merely restate the fact sheet the model was
shown, for duplicates, for inline code/data that does not occur in the binary
(self-referential), and for echoes of the tools' own previous output; otherwise it is
measured from the binary itself — how often the expected content occurs in this file and
how much entropy it has — so zero padding, a ubiquitous prologue, or a pattern that matches
everywhere weigh almost nothing even though they verify, and emulation must actually execute
non-degenerate code. A reconstruction is grounded only when nothing
is refuted and the verified weight reaches --min-information (default 1.0). This follows
the CORE refinement of FActScore: credit only claims that are factual, informative and
non-repetitive. reverify reconstruct --samples N draws several proposals per round and
lets the verifier — not the model's confidence — select among them.
EXAMPLE.md walks through one run on kernel32.dll — the model
proposes the textbook prologue from prior, the verifier refutes it with the real
bytes, and the model corrects to grounded, with no API key and no specific model.
BENCHMARK.md is the reproducible measurement behind the numbers above.
Everything above is checkable without trusting the author:
objdump and hand-verified Intel vectors, the emulator against Unicorn, the semantic
engine against the export table — plus fuzzing that a malformed file never crashes the
reader and that a wrong claim is never VERIFIED. All of it runs in CI on Linux, Windows
and macOS, with and without the engines; a nightly job fuzzes 20k inputs.benchmarks/results/, and a
third-party aarch64 replication is in BENCHMARK.md.