Model Evaluation & Red Teaming

Evaluation must cover average quality and severe failures, then connect findings to deployment decisions.

Manage capability, safety, security, bias, and tool-use failures through threat models and evidence-based release gates.

01

Problem definition

Average-case testing misses rare but consequential failures and adversarial paths.

Red-team findings without owners and closure criteria become reports while risk reaches production.

  • A model touches sensitive data, external tools, or customer decisions
  • Risk regression must be tested for every model or policy update
  • No threat model or risk-acceptance owner exists
  • There is no authority to remediate or block findings
02

Model Evaluation & Red Teaming: system plate

  1. 01

    State 1

    Begin with a threat model connecting assets, actors, permissions, and harm paths.

  2. 02

    State 2

    Combine automated attack generation with domain-expert manual exploration.

  3. 03

    State 3

    Maintain a finding ledger with reproducible inputs, environment, impact, and mitigation.

  4. 04

    State 4

    Do not close a release gate without retest evidence and residual-risk acceptance.

FEEDBACKFailed acceptance returns evidence to the first controlled stage: Threat coverage and attack surface.

A decision and validation view for Model Evaluation & Red Teaming; labels describe architecture, not a measured deployment result.
  1. The workflow begins with Begin with a threat model connecting assets, actors, permissions, and harm paths..
  2. It reaches an acceptance decision through Threat coverage and attack surface.

03

Design method

Evaluation must cover average quality and severe failures, then connect findings to deployment decisions.

  1. 01

    Stage 1

    Begin with a threat model connecting assets, actors, permissions, and harm paths.

  2. 02

    Stage 2

    Combine automated attack generation with domain-expert manual exploration.

  3. 03

    Stage 3

    Maintain a finding ledger with reproducible inputs, environment, impact, and mitigation.

  4. 04

    Stage 4

    Do not close a release gate without retest evidence and residual-risk acceptance.

04

Application scenarios

Hypothetical application scenario

A model touches sensitive data, external tools, or customer decisions

Average-case testing misses rare but consequential failures and adversarial paths.

APPROACH
Begin with a threat model connecting assets, actors, permissions, and harm paths.
BOUNDARY
Red teaming does not prove safety; it expands the known failure surface.
Hypothetical application scenario

Risk regression must be tested for every model or policy update

Red-team findings without owners and closure criteria become reports while risk reaches production.

APPROACH
Combine automated attack generation with domain-expert manual exploration.
BOUNDARY
Attack success in a test environment is not generalized to all production conditions.

05

Design choices

DecisionGainCostWatch
A model touches sensitive data, external tools, or customer decisionsBegin with a threat model connecting assets, actors, permissions, and harm paths.Red teaming does not prove safety; it expands the known failure surface.Threat coverage and attack surface
Risk regression must be tested for every model or policy updateCombine automated attack generation with domain-expert manual exploration.Attack success in a test environment is not generalized to all production conditions.Finding reproduction and remediation rate
06

Model Evaluation & Red Teaming: system plate

  1. No threat model or risk-acceptance owner exists

    Red teaming does not prove safety; it expands the known failure surface.

    SIGNAL
    Threat coverage and attack surface
    MITIGATION
    Combine automated attack generation with domain-expert manual exploration.
  2. There is no authority to remediate or block findings

    Attack success in a test environment is not generalized to all production conditions.

    SIGNAL
    Finding reproduction and remediation rate
    MITIGATION
    Maintain a finding ledger with reproducible inputs, environment, impact, and mitigation.
A decision and validation view for Model Evaluation & Red Teaming; labels describe architecture, not a measured deployment result.
  1. The workflow begins with Begin with a threat model connecting assets, actors, permissions, and harm paths..
  2. It reaches an acceptance decision through Threat coverage and attack surface.

07

Validation plan

MeasureMethodPass conditionCaveat
Threat coverage and attack surfaceBegin with a threat model connecting assets, actors, permissions, and harm paths.Repeated runs satisfy the acceptance threshold agreed during discoveryRed teaming does not prove safety; it expands the known failure surface.
Finding reproduction and remediation rateCombine automated attack generation with domain-expert manual exploration.Repeated runs satisfy the acceptance threshold agreed during discovery—
Residual-risk approval traceabilityMaintain a finding ledger with reproducible inputs, environment, impact, and mitigation.Repeated runs satisfy the acceptance threshold agreed during discovery—

08

Constraints and failure conditions

No threat model or risk-acceptance owner exists

Red teaming does not prove safety; it expands the known failure surface.

There is no authority to remediate or block findings

Attack success in a test environment is not generalized to all production conditions.

10

Durable deliverables

Model Evaluation & Red Teaming decision record
Evaluation must cover average quality and severe failures, then connect findings to deployment decisions.Client-owned · Patty-reviewed
Validation harness and acceptance criteria
Threat coverage and attack surface · Finding reproduction and remediation rate · Residual-risk approval traceabilityJointly maintained
Operations and recovery runbook
Red teaming does not prove safety; it expands the known failure surface. · Attack success in a test environment is not generalized to all production conditions.Operating-team owned

11

Terminology

Model Evaluation & Red Teaming
Manage capability, safety, security, bias, and tool-use failures through threat models and evidence-based release gates.
Acceptance criterion
Threat coverage and attack surface
Operating boundary
Red teaming does not prove safety; it expands the known failure surface.

REFERENCES

References and primary material

  1. NIST AI 600-1

    Primary material for the method and terminology.

  2. MITRE ATLAS

    Primary material for the method and terminology.