Version 2 · Linux ELF · April 28–29, 2026

Version 2 · Linux ELF

Calibration becomes the advantage.

A broader frontier-model field ran the same 13-task Linux suite. A small, non-thinking model led the final result by completing every task and making fewer unsupported claims.

13Linux ELF tasks
6ranked model runs
13/13leader submissions
1.92leader false claims/task

01 Main score leaderboard

Better judgment beat more deliberation on this suite.

Gemini 3.1 Flash Lite led Main and Total with full coverage. The result is specific to this suite and run; it does not imply that small models are universally better at reverse engineering.

Main score leaderboardMain · Total · unsupported claims/task
Gemini 3.1 Flash LiteTotal .6668 · 1.92 claims/task
0.5618
Kimi K2.6Total .4966 · 2.08 claims/task
0.4966
DeepSeek V4 FlashTotal .4486 · 3.38 claims/task
0.4486
DeepSeek V4 ProTotal .6476 · 2.69 claims/task
0.3767
Claude Opus 4.7Total .5119 · 5.62 claims/task
0.3673
GPT-5.5Total .2551 · 6.31 claims/task
0.2551

Main averages Levels 1–12. Total adds the Level 13 bonus, so DeepSeek V4 Pro ranks second by Total despite placing fourth on Main. GLM 5.1 and Gemini 3.1 Pro are omitted because they did not complete enough of the suite for a ranked comparison.

02 What V2 taught us

The frontier moved—but exact reconstruction still resisted.

V2 separated the metadata and recognition layer from the deeper execution layer, shaping the Windows tasks that followed.

Calibration

Lower overclaiming tracked stronger scores

The ranking by hallucination rate closely resembled the final ranking. Plausible but unsupported reports were expensive.

Progress

The best Main score rose 15.4%

On the same ELF suite, the leading Main result moved from 0.487 in V1 to 0.5618 in V2. This is descriptive across two benchmark generations.

Boundary

Recognition improved before decryption

Level 13's best bonus result reached 0.2709, but no run recovered its decoded command-and-control value or decoded strings.

03 Evaluation design

The same ladder, a more revealing model field.

Keeping the core V1 suite stable made the change interpretable. Expanded instrumentation captured reasoning time, failures, coverage, and tool-use patterns alongside correctness.

Input13 source-free Linux ELF binaries
Model fieldSix ranked runs; two additional DNFs
Constraint25 static-analysis tool calls per task
New signalReasoning time and operational failures

Next Version 3 · August 2026

From Linux recognition to Windows reconstruction.

A new platform, paired stripped artifacts, six models, and 120 independently replayed score rows.

Continue to Version 3