Awards
First place receives $1,000 plus one month of the recipient's choice of ChatGPT Pro or Claude Max 20x. Second place receives one month of the recipient's choice of ChatGPT Pro or Claude Max 5x.
A simple 33 days competition for original Linux x86-64 ELF binaries that meaningfully challenge frontier AI agents using static analysis.
Entrants keep binaries, checksums, metadata, and ground truth in private GitHub repositories. Any source code and build instructions requested from finalist or prize-eligible entries are provided privately and remain confidential. The public site shows only safe registration status and sanitized results.
First place receives $1,000 plus one month of the recipient's choice of ChatGPT Pro or Claude Max 20x. Second place receives one month of the recipient's choice of ChatGPT Pro or Claude Max 5x.
Submissions must target Linux x86-64 ELF, remain under 10 MB, and be suitable for static reverse engineering.
Public source release is not required. Private source, binaries, repository URLs, ground truth, evaluation notes, and full model transcripts remain confidential unless the entrant chooses to publish them.
Calendar dates are set in the central challenge configuration. The Season 1 operating model is fixed around submission, validation, evaluation, and announcement windows.
agentrebench as a collaborator.agentre-season-1-final and register the resolved 40-character commit SHA.Season 1 focuses on meaningful static analysis difficulty rather than aggressive anti-analysis. A strong entry makes the model recover intent through evidence, not through guessing or infrastructure failure.
Indirect data flow, cross-function reasoning, state-machine reconstruction, and meaningful control-flow analysis.
Sparse but recoverable clues, custom protocol recovery, algorithm identification, and compiler optimization effects.
Private ground truth should explain important functions, key behaviors, traps, partially correct findings, and scoring expectations.
AgentRE scores how much required behavior each model correctly recovers. The primary ranking is lowest average official model correctness, reported as entry difficulty score: 100 - average official model correctness. The first-place entry must also be lowest-scoring on at least 4 of 6 official models.
Season 1 uses the same six frontier-model runs represented on the public AgentRE leaderboard at competition opening. Only AgentRE-verified runs count toward rankings and awards.
Tool crashes, API errors, malformed inputs, missing dependencies, challenge-unrelated infrastructure abuse, and unrelated safety refusals are invalid runs. A model's failure to solve a valid difficult binary within the standard budget is a valid outcome.
Official scoring uses a fixed AgentRE toolset from agentre-bench-tools:latest. Models may not generate and execute arbitrary Python or code in other programming languages. Allowing custom parsers, symbolic executors, deobfuscators, or other new tools during evaluation would make model comparisons unfair and difficult to reproduce.
AgentRE may reject unsafe, copied, trivial, broken, unverifiable, or low-quality submissions. Security-relevant behavior must be inert, isolated, safe, clearly disclosed in private ground truth, and useful for reverse-engineering evaluation.
Entries must not perform credential theft, credential collection, persistence intended for deployment, propagation, destructive payloads, ransomware behavior, external downloads, real command-and-control, exploitation of third-party systems, personal-data collection, or build scripts/workflows intended to compromise infrastructure.
This summary is loaded from public-safe registration JSON files merged through pull requests. No private repository URLs, binaries, source code, ground truth, private evaluation notes, or full model transcripts are stored in this public website repository.
Private repositories prevent competitors and future models from seeing binaries, source code, challenge techniques, and ground truth during judging.
No. AgentRE reviews submissions privately. Public releases require separate participant approval.
No automatic publication occurs. Public release of source or binaries requires separate participant permission.
Yes. One entrant or team may submit up to 3 entries and may win only one prize.
Yes before the final deadline. AgentRE evaluates only the commit resolved from the final tag before the deadline.
The commit resolved from agentre-season-1-final, recorded as a full 40-character SHA in the public-safe registration PR.
No public source release is required. AgentRE may request source code and build instructions privately from finalist or prize-eligible entries to verify ownership, safety, reproducibility, intended behavior, and ground-truth accuracy. Those materials remain confidential unless you choose to publish them.
Yes, when it supports meaningful static reverse-engineering difficulty and is disclosed in the private ground truth.
No real malware. Controlled simulation may be accepted only when inert, isolated, safe, clearly disclosed, and useful for evaluation.
First place receives $1,000 plus one month of the recipient's choice of ChatGPT Pro or Claude Max 20x. Second place receives one month of the recipient's choice of ChatGPT Pro or Claude Max 5x.
If no entry passes validation, evaluation cannot produce a reliable result, or no entry satisfies the 4-of-6 first-place requirement, AgentRE may decline to issue either award.
Submission grants only the limited permission needed to inspect, clone, rebuild, evaluate, store an internal evaluation copy, score, and announce results. Commercial use, training use, benchmark inclusion, dataset resale, or redistribution requires a separate explicit agreement.
Yes. The registration JSON includes display flags for the public leaderboard. The pull request itself is public.
Scores are published after judging and review, with award recipients announced on September 3, 2026.
Season 1 can be archived. Sanitized summaries and optional public repository links may be added only after participant approval.