Completion can hide weak evidence
GPT submitted 13 of 13 answers, yet its final score was lowest and it averaged the most unsupported technique claims.
Version 1 · Linux ELF · February 13, 2026
The baseline: binary reasoning becomes measurable.
The first AgentRE-Bench run asked three frontier models to investigate 13 synthetic Linux binaries with no source code, a 25-call tool budget, and deterministic scoring.
01 Published results
Claude Opus 4.6 led the first suite. DeepSeek V3.2 was close on score while making fewer unsupported technique claims. GPT-5.2 completed every task but paid the largest hallucination penalty.
02 What V1 taught us
The first benchmark established the core trust problem AgentRE continues to measure: coverage, correctness, and calibration are different dimensions of capability.
GPT submitted 13 of 13 answers, yet its final score was lowest and it averaged the most unsupported technique claims.
The scorer's false-claim penalty separated agents that recovered defensible facts from those that produced broader, plausible reports.
All three models scored zero on Level 13, the encrypted metamorphic bonus task. Exact recovery—not recognition—was already the wall.
03 Evaluation design
Purpose-built binaries kept ground truth known and scoring repeatable while the task sequence moved from plaintext behavior toward encryption, obfuscation, and anti-analysis.