Lower overclaiming tracked stronger scores
The ranking by hallucination rate closely resembled the final ranking. Plausible but unsupported reports were expensive.
Version 2 · Linux ELF · April 28–29, 2026
Calibration becomes the advantage.
A broader frontier-model field ran the same 13-task Linux suite. A small, non-thinking model led the final result by completing every task and making fewer unsupported claims.
01 Main score leaderboard
Gemini 3.1 Flash Lite led Main and Total with full coverage. The result is specific to this suite and run; it does not imply that small models are universally better at reverse engineering.
Main averages Levels 1–12. Total adds the Level 13 bonus, so DeepSeek V4 Pro ranks second by Total despite placing fourth on Main. GLM 5.1 and Gemini 3.1 Pro are omitted because they did not complete enough of the suite for a ranked comparison.
02 What V2 taught us
V2 separated the metadata and recognition layer from the deeper execution layer, shaping the Windows tasks that followed.
The ranking by hallucination rate closely resembled the final ranking. Plausible but unsupported reports were expensive.
On the same ELF suite, the leading Main result moved from 0.487 in V1 to 0.5618 in V2. This is descriptive across two benchmark generations.
Level 13's best bonus result reached 0.2709, but no run recovered its decoded command-and-control value or decoded strings.
03 Evaluation design
Keeping the core V1 suite stable made the change interpretable. Expanded instrumentation captured reasoning time, failures, coverage, and tool-use patterns alongside correctness.