What did the model recover?
Behaviors, API chains, configuration, exact values, infrastructure, and execution details.
Private evals Held out by design
Know what your model can do before the data is familiar.
AgentRE designs confidential reverse-engineering evaluations around your model, tools, and target workflows. Fresh binaries and deterministic ground truth help distinguish generalization from benchmark recall.
Designed for model developers, security teams, and research organizations evaluating consequential cyber workflows.
Public benchmarks create a shared baseline, but their tasks can eventually appear in training data. A model can improve on the leaderboard without proving it will generalize to new software.
AgentRE builds a fresh evaluation surface around the capability your organization needs: unfamiliar artifacts, realistic constraints, and facts that can be verified independently.
Every engagement decomposes performance into the technical and operational dimensions that determine whether an agent can increase cyber capability in practice.
Behaviors, API chains, configuration, exact values, infrastructure, and execution details.
Completion, refusals, unsupported claims, calibration, and consistency across repeated runs.
Tool selection, evidence quality, investigation strategy, time, and budget efficiency.
New binaries, stripped variants, altered workflows, and artifacts outside familiar distributions.
The evaluation begins with the client question—not a generic task list. We translate that question into a controlled benchmark and an evidence package the technical team and leadership can use.
Specify the target capability, model or agent stack, operating constraints, and decision criteria.
Create fresh artifacts, evaluation tasks, expert ground truth, and a scoring rubric aligned to the scope.
Evaluate models, checkpoints, and tool configurations under identical, controlled conditions.
Translate the results into capability maps, episode evidence, failure priorities, and next actions.
The same underlying evidence is organized for different audiences: a clear capability view for decision-makers and episode-level detail for the teams building the system.
Evaluation artifacts remain separate from development and training data. When training follows, AgentRE uses a distinct task pool for learning and fresh private binaries for re-evaluation. That separation shows whether the capability generalized instead of merely memorizing the task.
AgentRE Private evaluations
Tell us the model, workflow, and capability question. We will scope a fresh private benchmark and evidence package around it.