Kimi K3
Led the core benchmark and bonus category with valid answers for all 20 artifacts.
Version 3 evaluates six frontier models on 20 paired Windows PE artifacts using static-analysis tools only.
Kimi K3 achieved the strongest overall performance. Every answer, tool transcript, score, binary, ground truth, and checksum is published for independent review.
Leading agents can orient themselves inside unfamiliar binaries and identify broad Windows behavior. Useful performance still depends on exact recovery, evidence discipline, tool planning, refusals, and completing the task within a fixed budget.
Led the core benchmark and bonus category with valid answers for all 20 artifacts.
Removing symbols caused a moderate aggregate reduction across complete stripped and unstripped pairs.
Broad techniques were easier to identify than keys, configuration, infrastructure, and complete API chains.
Refusals and tool-budget exhaustion reduced useful coverage independently of analytical knowledge.
The solid bar shows average performance on the 18 core artifacts. The shaded extension shows results from the two synthetic-worm artifacts, scored separately as a bonus. Total adds the two figures together; it is not an average across all 20 artifacts.
Terminal refusals and tool-budget exits remain part of the observed result.
Only Kimi and Gemini returned valid answers for all 20 artifacts.
No model recovered the actual Level 23 command-and-control endpoint.
| Rank | Model | Main /1 | Bonus /1 | Valid | Terminal outcome | Total /2 |
|---|---|---|---|---|---|---|
| 01 | Kimi K3 | 0.6757 | 0.6861 | 20/20 | None | 1.3618 |
| 02 | Claude Opus 5 | 0.6147 | 0.6638 | 18/20 | 2 refusals | 1.2785 |
| 03 | GPT-5.6 Sol | 0.5859 | 0.6298 | 18/20 | 2 budget exits | 1.2157 |
| 04 | DeepSeek V4 Flash | 0.5125 | 0.6782 | 19/20 | 1 budget exit | 1.1906 |
| 05 | Gemini 3.6 Flash | 0.5400 | 0.5653 | 20/20 | None | 1.1052 |
| 06 | Claude Fable 5 | 0.0000 | 0.0000 | 0/20 | 20 refusals | 0.0000 |
Across 46 model/level pairs with valid answers on both variants, mean score fell from 0.6428 unstripped to 0.6012 stripped. Static PE evidence still preserved much of the behavior models needed.
Every stripped artifact removed the symbol names available in its paired original.
Aggregate binary size fell sharply, along with the number of sections.
The average relative decline was 6.47%, so stripping made analysis harder without making the binaries opaque.
| Model | Pairs | Unstripped | Stripped | Change |
|---|---|---|---|---|
| Kimi K3 | 10 | 0.7122 | 0.6413 | −0.0709 |
| Claude Opus 5 | 8 | 0.7313 | 0.6720 | −0.0593 |
| GPT-5.6 Sol | 9 | 0.6731 | 0.6387 | −0.0344 |
| DeepSeek V4 Flash | 9 | 0.5369 | 0.5525 | +0.0156 |
| Gemini 3.6 Flash | 10 | 0.5705 | 0.5145 | −0.0560 |
| Level | Pairs | Unstripped | Stripped | Change |
|---|---|---|---|---|
| 14 · DLL Injection | 5 | 0.8617 | 0.7327 | −0.1290 |
| 15 · APC Injection | 4 | 0.6064 | 0.5757 | −0.0307 |
| 16 · Code Cave | 5 | 0.7672 | 0.7219 | −0.0453 |
| 17 · Process Hollowing | 5 | 0.4623 | 0.5096 | +0.0473 |
| 18 · Hell’s Gate | 5 | 0.5119 | 0.4649 | −0.0469 |
| 19 · Reflective DLL | 4 | 0.6119 | 0.4114 | −0.2004 |
| 20 · Remote PE | 5 | 0.6678 | 0.7054 | +0.0376 |
| 21 · Ghost Hollowing | 4 | 0.5484 | 0.6502 | +0.1018 |
| 22 · Advanced Evasion | 4 | 0.7082 | 0.5750 | −0.1332 |
| 23 · Synthetic Worm | 5 | 0.6628 | 0.6264 | −0.0364 |
These are descriptive differences from one run per model and artifact, not confidence-bounded causal estimates.
Coverage, precision, speed, refusal behavior, and budget discipline changed the practical value of each model’s run.
Strongest overall and best on decoded strings and anti-analysis. Complete, but the slowest successful run at 9.74 summed episode-hours.
Best injection-detail mean and strong technique recall. Two mid-analysis refusals reduced useful coverage.
Best single score with no unsupported technique claims. Both Level 22 variants ended at the tool-call limit.
Strong bonus and encryption results, with greater variance, one budget exit, and the most unsupported claims.
Fastest complete run and strong metadata recovery. Exact decoded-string and technique detail were weaker.
The technique family was recognized, but required implementation details were missing.
APIs, techniques, configuration values, or execution steps were reported without recovered evidence.
Individual components were found but not connected into the correct execution sequence.
Low-value exploration consumed the fixed call budget before a complete answer was produced.
An early hypothesis was reported before critical details were verified.
A valid defensive analysis task ended without a gradable answer.
Models were better at identifying the shape of a program than recovering the details required for a final incident conclusion.
Agents worked from isolated binaries with fixed tools and budgets. The benchmark does not measure dynamic analysis, debugging, or interactive decompiler workflows.
Each model-artifact pair was evaluated once, so results characterize the recorded runs rather than a stable population mean.
Deterministic grading is possible because behavior is known, but the suite does not reproduce the full diversity of real-world malware.
Each result can be followed from input artifact to agent transcript, tool output, final answer, deterministic grader decision, and aggregate score.
One anonymized PE per isolated workspace. No source, Internet, debugger, decompiler, compilation, or dynamic execution.
No LLM judge. Terminal refusals and budget exits remain in the aggregates. Main is the broad score; Level 23 supplies the separate Bonus.
All 120 answers were rescored, all six aggregates reconciled, and the 366-file checksum ledger and artifact hashes passed verification.
Current agents can accelerate orientation and hypothesis generation. Exact reconstruction still requires stronger tool planning, longer-horizon evidence tracking, better calibration, efficient budget use, and expert verification.
Apply them to triage, prioritization, and evidence discovery while retaining analyst review for configuration, attribution, and consequential conclusions.
Optimize for evidence-grounded technical facts, efficient tool paths, uncertainty reporting, and completion reliability—not broader descriptions alone.