From recognition to reconstruction
Ten Windows programs, paired with and without symbols. Agents recovered substantial behavior after symbol removal, while exact execution chains and configuration remained a boundary.
V3 · Windows PEOpen benchmark Version 3 · August 2026
Can AI agents reason about software without its source code?
AgentRE evaluates models on unseen binaries, identifies where they fail, and turns those failures into training signals for better reverse engineering agents.
Ten Windows programs, each evaluated with and without symbols. Six models, 20 binaries, 120 episode slots.
| # | Model | Main / 1 | Bonus / 1 | Total / 2 | Valid answers |
|---|---|---|---|---|---|
| 01 | Kimi K3 | 0.6757 | 0.6861 | 1.3618 | 20/20 |
| 02 | Claude Opus 5 | 0.6147 | 0.6638 | 1.2785 | 18/20 |
| 03 | GPT-5.6 Sol | 0.5859 | 0.6298 | 1.2157 | 18/20 |
| 04 | DeepSeek V4 Flash | 0.5125 | 0.6782 | 1.1906 | 19/20 |
| 05 | Gemini 3.6 Flash | 0.5400 | 0.5653 | 1.1052 | 20/20 |
| 06 | Claude Fable 5 | 0.0000 | 0.0000 | 0.0000 | 0/20 |
Main measures average performance on 18 core artifacts. Bonus covers the two synthetic-worm artifacts. Total adds the two scores; it is not an average across all 20 artifacts. Refusals and budget exits remain part of the result.
Read the complete V3 reportMost AI coding benchmarks begin with readable code and a clear task. Security teams often begin with neither. They get a compiled program, limited context, and one question: what does this actually do?
AgentRE measures whether an agent can turn that uncertainty into a defensible explanation. Agents recover behaviors, configurations, infrastructure, API chains, and execution details through tool use. Correct findings earn credit; unsupported claims are penalized against fixed expert ground truth.
Why AgentRE mattersEvery agent gets the same constrained environment. Every claim meets the same fixed ground truth.
A source-free, anonymized binary in an isolated workspace, with ordinary static-analysis tools and no internet access.
Tool calls, evidence, time, coverage, failures, and final structured claims are captured for review.
Fixed expert ground truth rewards correct detail and penalizes unsupported claims. The same answer always receives the same score.
Ten Windows programs, paired with and without symbols. Agents recovered substantial behavior after symbol removal, while exact execution chains and configuration remained a boundary.
V3 · Windows PEA broader model field ran the same Linux ladder. Gemini 3.1 Flash Lite led with a 0.5618 Main score and full 13/13 coverage. Fewer unsupported claims mattered more than deeper reasoning on this suite.
V2 · Linux ELFThirteen synthetic Linux challenges established a source-free baseline with constrained tools and deterministic scoring. Completion was not the same as trust; every model scored zero on the bonus challenge.
V1 · Linux ELFV1 and V2 share the Linux suite. V3 uses a new platform, paired artifacts, tasks, and rubric. Scores are not directly comparable across the two suites.
The public benchmark establishes a baseline. Private evaluations and client-specific training turn measured failures into improvements.
Held-out and client-specific binaries expose capability gaps across models, checkpoints, and agents. Get capability maps and failure analysis for real cybersecurity workflows.
Explore private evalsVerified outcomes, expert demonstrations, and complete tool trajectories become deterministic reward signals and expert-curated datasets for client-specific training and re-evaluation.
Explore RL DataInterested in evaluating or training your agent?
Get in touch