Open benchmark Version 3 · August 2026

AgentRE-Bench

Can AI agents reason about software without its source code?

AgentRE evaluates models on unseen binaries, identifies where they fail, and turns those failures into training signals for better reverse engineering agents.

33
distinct binary challenges
Linux + Windows
two binary ecosystems
Deterministic
replayable scoring
Open
methodology & transcripts

Latest results

Ten Windows programs, each evaluated with and without symbols. Six models, 20 binaries, 120 episode slots.

Version 3 leaderboard · ranked by Total
#ModelMain / 1Bonus / 1Total / 2Valid answers
01Kimi K30.67570.68611.361820/20
02Claude Opus 50.61470.66381.278518/20
03GPT-5.6 Sol0.58590.62981.215718/20
04DeepSeek V4 Flash0.51250.67821.190619/20
05Gemini 3.6 Flash0.54000.56531.105220/20
06Claude Fable 50.00000.00000.00000/20

Main measures average performance on 18 core artifacts. Bonus covers the two synthetic-worm artifacts. Total adds the two scores; it is not an average across all 20 artifacts. Refusals and budget exits remain part of the result.

Read the complete V3 report

Key findings

Explore the findings and limitations

Why reverse engineering?

Most AI coding benchmarks begin with readable code and a clear task. Security teams often begin with neither. They get a compiled program, limited context, and one question: what does this actually do?

AgentRE measures whether an agent can turn that uncertainty into a defensible explanation. Agents recover behaviors, configurations, infrastructure, API chains, and execution details through tool use. Correct findings earn credit; unsupported claims are penalized against fixed expert ground truth.

Why AgentRE matters

How it works

Every agent gets the same constrained environment. Every claim meets the same fixed ground truth.

  1. Present unfamiliar software

    A source-free, anonymized binary in an isolated workspace, with ordinary static-analysis tools and no internet access.

  2. Record the investigation

    Tool calls, evidence, time, coverage, failures, and final structured claims are captured for review.

  3. Score what can be proved

    Fixed expert ground truth rewards correct detail and penalizes unsupported claims. The same answer always receives the same score.

Explore the methodology and transcripts

Benchmark history

V3Current

From recognition to reconstruction

Ten Windows programs, paired with and without symbols. Agents recovered substantial behavior after symbol removal, while exact execution chains and configuration remained a boundary.

V3 · Windows PE
V2

Calibration becomes the advantage

A broader model field ran the same Linux ladder. Gemini 3.1 Flash Lite led with a 0.5618 Main score and full 13/13 coverage. Fewer unsupported claims mattered more than deeper reasoning on this suite.

V2 · Linux ELF
V1

Establishing the baseline

Thirteen synthetic Linux challenges established a source-free baseline with constrained tools and deterministic scoring. Completion was not the same as trust; every model scored zero on the bonus challenge.

V1 · Linux ELF

V1 and V2 share the Linux suite. V3 uses a new platform, paired artifacts, tasks, and rubric. Scores are not directly comparable across the two suites.

From evaluation to training

The public benchmark establishes a baseline. Private evaluations and client-specific training turn measured failures into improvements.

Private evaluations

Held-out and client-specific binaries expose capability gaps across models, checkpoints, and agents. Get capability maps and failure analysis for real cybersecurity workflows.

Explore private evals

RL + SFT training

Verified outcomes, expert demonstrations, and complete tool trajectories become deterministic reward signals and expert-curated datasets for client-specific training and re-evaluation.

Explore RL Data

Interested in evaluating or training your agent?

Get in touch

Referenced in research