Private evals Held out by design

Private evaluations

Know what your model can do before the data is familiar.

AgentRE designs confidential reverse-engineering evaluations around your model, tools, and target workflows. Fresh binaries and deterministic ground truth help distinguish generalization from benchmark recall.

Designed for model developers, security teams, and research organizations evaluating consequential cyber workflows.

Evaluation briefClient-specific
01
Target capabilityThe behaviors, decisions, and workflows your team needs to trust.
02
Fresh corpusHeld-out artifacts developed outside the public benchmark.
03
Controlled runYour models, checkpoints, agents, and tools under repeatable constraints.
04
Decision-ready evidenceCapability maps, failure modes, and complete episode traces.
Fresh artifactsSeparate from the public suite
Tailored scopeBuilt around the client decision
Fixed ground truthDeterministic, replayable scoring
Complete tracesEvidence behind every score

Private evaluation of model generalization

Public benchmarks create a shared baseline, but their tasks can eventually appear in training data. A model can improve on the leaderboard without proving it will generalize to new software.

AgentRE builds a fresh evaluation surface around the capability your organization needs: unfamiliar artifacts, realistic constraints, and facts that can be verified independently.

Capabilities and behaviors measured

Every engagement decomposes performance into the technical and operational dimensions that determine whether an agent can increase cyber capability in practice.

01 / Recovery

What did the model recover?

Behaviors, API chains, configuration, exact values, infrastructure, and execution details.

02 / Reliability

Can the team depend on it?

Completion, refusals, unsupported claims, calibration, and consistency across repeated runs.

03 / Tool use

How did it reach the answer?

Tool selection, evidence quality, investigation strategy, time, and budget efficiency.

04 / Generalization

Does it survive novelty?

New binaries, stripped variants, altered workflows, and artifacts outside familiar distributions.

Private evaluation workflow

The evaluation begins with the client question—not a generic task list. We translate that question into a controlled benchmark and an evidence package the technical team and leadership can use.

  1. Define

    Specify the target capability, model or agent stack, operating constraints, and decision criteria.

  2. Build

    Create fresh artifacts, evaluation tasks, expert ground truth, and a scoring rubric aligned to the scope.

  3. Run

    Evaluate models, checkpoints, and tool configurations under identical, controlled conditions.

  4. Deliver

    Translate the results into capability maps, episode evidence, failure priorities, and next actions.

Evaluation reports and technical artifacts

The same underlying evidence is organized for different audiences: a clear capability view for decision-makers and episode-level detail for the teams building the system.

  1. 01
    Executive capability mapWhat works, what fails, and where risk remains.
  2. 02
    Model and checkpoint comparisonControlled comparisons across the systems and configurations in scope.
  3. 03
    Failure-mode taxonomyExact recovery gaps, hallucinations, refusals, inefficiencies, and incomplete investigations.
  4. 04
    Episode-level evidenceStructured answers, tool calls, transcripts, scores, and artifacts for audit and debugging.
  5. 05
    Regression planA repeatable way to measure whether changes improve capability without hiding new failure modes.

Separation between evaluation and development

Evaluation artifacts remain separate from development and training data. When training follows, AgentRE uses a distinct task pool for learning and fresh private binaries for re-evaluation. That separation shows whether the capability generalized instead of merely memorizing the task.

AgentRE Private evaluations

Measure cyber capability on fresh held-out data.

Tell us the model, workflow, and capability question. We will scope a fresh private benchmark and evidence package around it.