Version 1 · Linux ELF · February 13, 2026

Version 1 · Linux ELF

The baseline: binary reasoning becomes measurable.

The first AgentRE-Bench run asked three frontier models to investigate 13 synthetic Linux binaries with no source code, a 25-call tool budget, and deterministic scoring.

13Linux ELF tasks
3models reported
25tool calls per task
0bonus scores above zero

01 Published results

A wide gap appeared immediately.

Claude Opus 4.6 led the first suite. DeepSeek V3.2 was close on score while making fewer unsupported technique claims. GPT-5.2 completed every task but paid the largest hallucination penalty.

Archive noteThese values are preserved from the final V1 homepage snapshot. The current repository does not retain V1 raw reports or transcripts, so this page does not describe the result as independently replay-verified.
Main score leaderboardScore · submissions · unsupported claims/task
Claude Opus 4.611/13 · 3.31 claims/task
0.487
DeepSeek V3.28/13 · 1.77 claims/task
0.476
GPT-5.213/13 · 6.23 claims/task
0.190

02 What V1 taught us

Answering every task did not mean understanding every task.

The first benchmark established the core trust problem AgentRE continues to measure: coverage, correctness, and calibration are different dimensions of capability.

Coverage

Completion can hide weak evidence

GPT submitted 13 of 13 answers, yet its final score was lowest and it averaged the most unsupported technique claims.

Calibration

Restraint became measurable

The scorer's false-claim penalty separated agents that recovered defensible facts from those that produced broader, plausible reports.

Frontier

The hardest reconstruction failed universally

All three models scored zero on Level 13, the encrypted metamorphic bonus task. Exact recovery—not recognition—was already the wall.

03 Evaluation design

A controlled test of long-horizon tool use.

Purpose-built binaries kept ground truth known and scoring repeatable while the task sequence moved from plaintext behavior toward encryption, obfuscation, and anti-analysis.

InputAn anonymized compiled ELF binary
Tool surfaceStatic Linux analysis tools only
Budget25 calls before final submission
ScoringWeighted facts with false-claim penalties

Next Version 2 · April 2026

Calibration becomes the advantage.

The same ELF ladder met a broader model field—and the original story changed.

Continue to Version 2