Model Evaluation & Red Teaming

Evaluation must cover average quality and severe failures, then connect findings to deployment decisions.

Manage capability, safety, security, bias, and tool-use failures through threat models and evidence-based release gates.

§ 01

Problem definition

The operating conditions that justify Model Evaluation & Red Teaming

Average-case testing misses rare but consequential failures and adversarial paths.

Red-team findings without owners and closure criteria become reports while risk reaches production.

  • A model touches sensitive data, external tools, or customer decisions
  • Risk regression must be tested for every model or policy update
  • No threat model or risk-acceptance owner exists
  • There is no authority to remediate or block findings
PLATE 01

Model Evaluation & Red Teaming: system plate

  1. 01

    State 1

    Begin with a threat model connecting assets, actors, permissions, and harm paths.

  2. 02

    State 2

    Combine automated attack generation with domain-expert manual exploration.

  3. 03

    State 3

    Maintain a finding ledger with reproducible inputs, environment, impact, and mitigation.

  4. 04

    State 4

    Do not close a release gate without retest evidence and residual-risk acceptance.

FEEDBACKFailed acceptance returns evidence to the first controlled stage: Threat coverage and attack surface.

A decision and validation view for Model Evaluation & Red Teaming; labels describe architecture, not a measured deployment result.
  1. The workflow begins with Begin with a threat model connecting assets, actors, permissions, and harm paths..
  2. It reaches an acceptance decision through Threat coverage and attack surface.

§ 03

Design method

Fix the boundary and acceptance criteria before implementation.

Evaluation must cover average quality and severe failures, then connect findings to deployment decisions.

  1. 01

    Stage 1

    Begin with a threat model connecting assets, actors, permissions, and harm paths.

    Review artifact 1
  2. 02

    Stage 2

    Combine automated attack generation with domain-expert manual exploration.

    Review artifact 2
  3. 03

    Stage 3

    Maintain a finding ledger with reproducible inputs, environment, impact, and mitigation.

    Review artifact 3
  4. 04

    Stage 4

    Do not close a release gate without retest evidence and residual-risk acceptance.

    Review artifact 4

§ 04

Application scenarios

Hypothetical workloads make the applicability boundary concrete.

Hypothetical application scenario

A model touches sensitive data, external tools, or customer decisions

Average-case testing misses rare but consequential failures and adversarial paths.

APPROACH
Begin with a threat model connecting assets, actors, permissions, and harm paths.
BOUNDARY
Red teaming does not prove safety; it expands the known failure surface.
Hypothetical application scenario

Risk regression must be tested for every model or policy update

Red-team findings without owners and closure criteria become reports while risk reaches production.

APPROACH
Combine automated attack generation with domain-expert manual exploration.
BOUNDARY
Attack success in a test environment is not generalized to all production conditions.

§ 05

Design choices

Review gains and costs in the same table.

DecisionGainCostWatch
A model touches sensitive data, external tools, or customer decisionsBegin with a threat model connecting assets, actors, permissions, and harm paths.Red teaming does not prove safety; it expands the known failure surface.Threat coverage and attack surface
Risk regression must be tested for every model or policy updateCombine automated attack generation with domain-expert manual exploration.Attack success in a test environment is not generalized to all production conditions.Finding reproduction and remediation rate
PLATE 02

Model Evaluation & Red Teaming: system plate

  1. No threat model or risk-acceptance owner exists

    Red teaming does not prove safety; it expands the known failure surface.

    SIGNAL
    Threat coverage and attack surface
    MITIGATION
    Combine automated attack generation with domain-expert manual exploration.
  2. There is no authority to remediate or block findings

    Attack success in a test environment is not generalized to all production conditions.

    SIGNAL
    Finding reproduction and remediation rate
    MITIGATION
    Maintain a finding ledger with reproducible inputs, environment, impact, and mitigation.
A decision and validation view for Model Evaluation & Red Teaming; labels describe architecture, not a measured deployment result.
  1. The workflow begins with Begin with a threat model connecting assets, actors, permissions, and harm paths..
  2. It reaches an acceptance decision through Threat coverage and attack surface.

§ 07

Validation plan

Agree on measurement conditions before publishing a result.

MeasureMethodPass conditionCaveat
Threat coverage and attack surfaceBegin with a threat model connecting assets, actors, permissions, and harm paths.Repeated runs satisfy the acceptance threshold agreed during discoveryRed teaming does not prove safety; it expands the known failure surface.
Finding reproduction and remediation rateCombine automated attack generation with domain-expert manual exploration.Repeated runs satisfy the acceptance threshold agreed during discovery
Residual-risk approval traceabilityMaintain a finding ledger with reproducible inputs, environment, impact, and mitigation.Repeated runs satisfy the acceptance threshold agreed during discovery

§ 08

Constraints and failure conditions

Conditions for not applying the capability are part of the design.

No threat model or risk-acceptance owner exists

Red teaming does not prove safety; it expands the known failure surface.

There is no authority to remediate or block findings

Attack success in a test environment is not generalized to all production conditions.

§ 10

Durable deliverables

Artifacts remain with the operating organization after the engagement.

Model Evaluation & Red Teaming decision record
Evaluation must cover average quality and severe failures, then connect findings to deployment decisions.Client-owned · Patty-reviewed
Validation harness and acceptance criteria
Threat coverage and attack surface · Finding reproduction and remediation rate · Residual-risk approval traceabilityJointly maintained
Operations and recovery runbook
Red teaming does not prove safety; it expands the known failure surface. · Attack success in a test environment is not generalized to all production conditions.Operating-team owned

§ 11

Terminology

Use shared terms with explicit operating meaning.

Model Evaluation & Red Teaming
Manage capability, safety, security, bias, and tool-use failures through threat models and evidence-based release gates.
Acceptance criterion
Threat coverage and attack surface
Operating boundary
Red teaming does not prove safety; it expands the known failure surface.

REFERENCES

References and primary material

  1. NIST AI 600-1

    Primary material for the method and terminology.

  2. MITRE ATLAS

    Primary material for the method and terminology.

Begin by determining whether Model Evaluation & Red Teaming is the justified next step.

We define scope and validation against representative work, data and infrastructure boundaries, and explicit failure conditions.

Request a technical review