RL Data Verified training signals

RL + SFT training

Turn cyber failure modes into training signal.

AgentRE transforms deterministic outcomes, complete tool-use trajectories, and expert demonstrations into high-value data for RL and SFT—then helps clients train the reverse-engineering capabilities their models need.

Client-specific data and training programs for models that must investigate, reason, use tools, and stay grounded in technical evidence.

Signal stackVerifiable
01
TaskFresh artifact, target capability, controlled tools, and explicit constraints.
02
TrajectoryComplete investigation path: tool calls, evidence, decisions, and final answer.
03
RewardGranular credit for correct facts, efficient process, and calibrated conclusions.
04
VerificationFresh held-out re-evaluation shows whether the capability generalized.
Outcome rewardsFacts checked against ground truth
Process tracesEvery observable tool call and response retained
Expert demonstrationsCurated examples for SFT
Held-out re-evaluationImprovement tested on fresh tasks

Training signal from the full investigation

A model can identify the right behavior but miss an exact key, choose the right tool but exhaust its budget, or recover most of a binary and then add unsupported claims.

AgentRE separates those outcomes so training can reward the capability that matters: correct recovery, disciplined investigation, efficient tool use, and restraint under uncertainty.

The data captured in each episode

The evaluator captures more than whether an answer looks plausible. Each episode becomes a structured training asset with precise positive and negative feedback.

01 / Outcomes

Verified technical facts

Techniques, exact values, execution chains, configuration, endpoints, and behavioral details.

02 / Process

Investigation quality

Tool choice, evidence use, sequencing, coverage, efficiency, and budget discipline.

03 / Negative signal

What should not be reinforced

Incorrect detail, unsupported claims, unnecessary refusal, wasted actions, and ungrounded certainty.

04 / Demonstrations

What good analysis looks like

Expert-curated solutions, successful trajectories, and structured outputs for supervised learning.

Using the data for RL and SFT

Different capability gaps need different forms of supervision. AgentRE connects both approaches to the same evidence standard and held-out evaluation loop.

Reinforcement learning

Optimize against granular, deterministic rewards.

Reward exact behavior recovery, efficient use of analysis tools, complete investigation, and calibrated uncertainty. Penalize wrong details and unsupported claims without relying on an LLM judge.

Supervised fine-tuning

Teach the investigation patterns the task requires.

Use expert demonstrations and curated successful trajectories to teach tool strategy, evidence standards, reasoning structure, and output formats before or alongside RL.

Client training workflow

Training starts with a measured failure mode and ends with a fresh test. This closes the loop between data generation and real capability improvement.

  1. Baseline

    Run a private evaluation to locate the model’s specific capability and reliability gaps.

  2. Design

    Define client-specific curricula, reward dimensions, demonstrations, and development tasks.

  3. Train

    Use RL, SFT, or a combined program on data kept separate from the final evaluation set.

  4. Re-evaluate

    Measure the resulting model on fresh held-out binaries and compare the full evidence trail.

Training data and reporting deliverables

The task distribution, evidence schema, and learning signal are designed around the client capability—not repackaged from a generic corpus.

  1. 01
    Verified episode datasetsArtifacts, prompts, environments, outcomes, and complete tool-use trajectories.
  2. 02
    Reward specificationGranular, replayable dimensions for facts, process, reliability, and efficiency.
  3. 03
    Expert SFT examplesCurated investigations and structured outputs that demonstrate the target behavior.
  4. 04
    Task curriculumProgressive development tasks aligned to the model’s current capability boundary.
  5. 05
    Pre/post capability reportA fresh, held-out comparison that shows what changed and what remains unresolved.

Verification and held-out evaluation

Without a trustworthy verifier, post-training can reinforce answers that only sound plausible. AgentRE anchors every signal to fixed technical ground truth, retains the complete episode trace, and tests the resulting model on fresh data.

AgentRE RL Data

Train against evidence. Verify the improvement.

Start with a private baseline. We will define the signal, build the curriculum, support RL and SFT, and measure the result on held-out tasks.