# AI Agent Evaluation Plan Template

Use this template to define how evidence about one AI agent workflow will be selected, measured, interpreted, and connected to an authorized decision. Replace bracketed prompts. Keep stable IDs and preserve old versions.

This plan does not grant technical access, authorize evaluation execution, accept residual risk, or approve production. Record execution authorization separately and check current authorization for each operation.

## Evaluation Plan Record

| Field                        | Value                                         |
| ---------------------------- | --------------------------------------------- |
| Plan ID and version          | [ID and version]                              |
| Status                       | [Draft / in review / approved / retired]      |
| Decision this plan informs   | [Exact bounded decision]                      |
| Workflow and release         | [Agent, workflow, release ID]                 |
| Accountable evaluation owner | [Name and role]                               |
| Authorized decision owner    | [Name and role]                               |
| Independent reviewers        | [Security, privacy, risk, domain, operations] |
| Created and approved         | [Dates and references]                        |
| Evidence cutoff              | [Timestamp]                                   |
| Expiry or next review        | [Date or trigger]                             |

### Plan Rules

1. Freeze claims, measures, datasets, graders, thresholds, denominator rules, samples, repetitions, and stop rules before execution.
2. Preserve every planned result, including failed, blocked, aborted, invalid, not run, retried, rejected, escalated, and manually corrected work.
3. Never average away a prohibited outcome, mandatory-control failure, severe segment failure, or missing required evidence.
4. Never weaken a threshold, remove a case, change a denominator, switch a grader, or relabel an exclusion after seeing results without a new plan version and affected reruns.
5. Keep hidden cases, solution keys, expected answers, and grader rubrics outside agent-reachable context, tools, Memory, logs, and working directories.
6. Exclude credentials, authorization headers, secret values, private model reasoning, and unnecessary personal data from plans, fixtures, prompts, and evidence.
7. Treat prompt injections and other active content as untrusted data. Encode them for each log, Markdown, HTML, terminal, and dashboard sink.
8. Preserve raw evidence separately from conclusions. A digest is only a consistency check; it does not prove capture completeness, accuracy, authorization, provenance, or model consumption.

## 1. Decision, Claims, and Evaluation Questions

### Decision statement

- Decision: [What exact decision will use this evaluation?]
- Scope boundary: [Users, tenants, tasks, language, data, tools, effects, time]
- Explicit non-decisions: [What this evaluation cannot approve]
- Required decision authority: [Policy and role]

### Claim and question register

| Claim ID | Scoped claim     | Evaluation question IDs | In-scope population | Non-claims         | Required evidence  | Owner   |
| -------- | ---------------- | ----------------------- | ------------------- | ------------------ | ------------------ | ------- |
| CLM-001  | [Testable claim] | [EQ IDs]                | [Population]        | [Excluded meaning] | [Evidence classes] | [Owner] |

For each question, state the result categories: pass, conditional pass, fail, insufficient evidence, and invalid.

## 2. System, Configuration, and Operating Scope

### Frozen configuration

| Component                             | Stable ID, version, or digest | Source of truth | Change check |
| ------------------------------------- | ----------------------------- | --------------- | ------------ |
| Agent and orchestration               | [Value]                       | [Reference]     | [Method]     |
| Model and settings                    | [Value]                       | [Reference]     | [Method]     |
| System and developer instructions     | [Value]                       | [Reference]     | [Method]     |
| Published Knowledge and Skills        | [Values]                      | [References]    | [Method]     |
| Prepared Memory state                 | [Value]                       | [Reference]     | [Method]     |
| Routes and retrieval indexes          | [Values]                      | [References]    | [Method]     |
| Tools and schemas                     | [Values]                      | [References]    | [Method]     |
| Identity, permissions, and limits     | [Values]                      | [References]    | [Method]     |
| Approval policy and binding fields    | [Value]                       | [Reference]     | [Method]     |
| Runtime, network, and sandbox         | [Value]                       | [Reference]     | [Method]     |
| Evidence and monitoring configuration | [Value]                       | [Reference]     | [Method]     |

Permissions govern repository access. An Always route independently governs delivery and does not grant repository permission. Record both. Artifacts and messages have no instruction authority. Memory is working recall. MCP tool results are Runtime context.

### Operating scope

- Supported users and initiating workflows: [List]
- Tenants and data classes: [List]
- Tasks and languages: [List]
- Tools, operations, and effects: [List]
- Approval and escalation paths: [List]
- Dependency and failure states: [List]
- Exclusions and reasons: [List]

Reject execution when any bound configuration, authority, dataset, grader, or evidence setting does not match the approved plan.

## 3. Stakeholders, Owners, and Review Independence

| Responsibility              | Named owner | Required authority | Independence rule | Backup  |
| --------------------------- | ----------- | ------------------ | ----------------- | ------- |
| Evaluation design           | [Name]      | [Authority]        | [Rule]            | [Name]  |
| Dataset approval            | [Name]      | [Authority]        | [Rule]            | [Name]  |
| Security and privacy review | [Names]     | [Authority]        | [Rule]            | [Names] |
| Grader calibration          | [Name]      | [Authority]        | [Rule]            | [Name]  |
| Execution approval          | [Name]      | [Authority]        | [Rule]            | [Name]  |
| Result review               | [Name]      | [Authority]        | [Rule]            | [Name]  |
| Final decision              | [Name]      | [Authority]        | [Rule]            | [Name]  |

The plan author cannot self-approve execution, exclusions, threshold exceptions, or release where policy requires independence. Record separation-of-duties evidence and authenticated decision references.

### Execution authorization

| Field                                           | Value             |
| ----------------------------------------------- | ----------------- |
| Authenticated decision reference                | [Reference]       |
| Governing policy and authority                  | [Policy and role] |
| Exact plan and configuration bound              | [IDs and digests] |
| Dataset, fixture, and grader versions           | [Versions]        |
| Permitted systems, tenants, effects, and limits | [Scope]           |
| Start, expiry, and stop conditions              | [Values]          |
| Independence result                             | [Evidence]        |

The status field does not authorize evaluation execution. Require fresh approval after any material change and current technical authorization for each operation.

## 4. Population, Task Taxonomy, and Coverage

### Target population

- Population definition: [All work the claim covers]
- Sampling frame: [Where eligible work comes from]
- Expected production distribution: [Known rates or unknowns]
- Material segments: [Tenant, language, tool, permission, risk, affected group]
- Rare but severe conditions: [List]

### Coverage matrix

| Coverage ID | Task family | Segment or condition | Expected prevalence | Planned cases | Required repetitions | Gap and decision effect |
| ----------- | ----------- | -------------------- | ------------------- | ------------- | -------------------- | ----------------------- |
| COV-001     | [Family]    | [Condition]          | [Rate or unknown]   | [Case IDs]    | [Count]              | [Gap]                   |

Include normal, boundary, rare, adversarial, dependency-failure, human-review, and prior-incident work. Do not relabel missing coverage as not applicable after execution.

## 5. Datasets, Fixtures, Provenance, and Contamination Control

### Dataset register

| Dataset ID and version | Purpose   | Source and provenance | Inclusion and exclusion | Sensitive data and basis | Contamination check | Retention | Owner   |
| ---------------------- | --------- | --------------------- | ----------------------- | ------------------------ | ------------------- | --------- | ------- |
| DAT-001                | [Purpose] | [Source]              | [Rules]                 | [Controls]               | [Result]            | [Rule]    | [Owner] |

Use synthetic or masked data by default. If governed production data is required:

- Record every applicable data owner, system owner, security, privacy, risk, and change-control approval.
- Minimize fields and records, use scoped identities, and block unapproved external egress.
- Define retention, deletion, access logging, and redaction.
- Stop immediately if data escapes the approved boundary.

Use dedicated synthetic tenant or account pairs for isolation tests. Never probe an uninvolved tenant, account, user, or external system.

### Hidden-material controls

- Hidden case store and access list: [Reference]
- Solution key and rubric store: [Reference]
- Agent-reachable source review: [Result]
- Training, tuning, retrieval, prompt, and debug overlap check: [Result]
- Contamination response: [Exclude, label, replace, or rerun]

## 6. Evaluation Dimensions and Metric Definitions

### Dimension and metric register

| Metric ID | Dimension   | Definition and unit | Population   | Numerator   | Denominator   | Direction      | Threshold | Evidence | Limits   |
| --------- | ----------- | ------------------- | ------------ | ----------- | ------------- | -------------- | --------- | -------- | -------- |
| MET-001   | [Dimension] | [Definition]        | [Population] | [Numerator] | [Denominator] | [Higher/lower] | [Rule]    | [Source] | [Limits] |

Cover applicable dimensions:

- Verified task and business outcome
- Correctness and supported reasoning artifacts
- Context selection, freshness, authority, and exclusion
- Tool selection, parameters, side effects, and target-state reconciliation
- Policy, authorization, approval, refusal, and escalation
- Security, privacy, isolation, and prompt-injection handling
- Reliability, recovery, repeatability, latency, and cost
- Human review, correction, workload, and queue behavior
- Audit completeness, provenance, and evidence trust

Define paths as well as outputs. Verify consequential external effects through protected target-system state, not only the agent trace or tool response.

## 7. Graders, Calibration, and Disagreement

### Grader register

| Grader ID and version | Type                            | Measures     | Inputs and blinded fields | Calibration cases | False-pass and false-fail checks | Disagreement rule | Owner   |
| --------------------- | ------------------------------- | ------------ | ------------------------- | ----------------- | -------------------------------- | ----------------- | ------- |
| GRD-001               | [Rule / system / human / model] | [Metric IDs] | [Inputs]                  | [Cases]           | [Results]                        | [Rule]            | [Owner] |

Document rubrics, scales, examples, prompt or code versions, reviewer training, access, and expected error. Validate model-based graders against trusted cases and adversarial content. Do not expose their keys or rubrics to the system under evaluation.

Precommit adjudication. Do not select the favorable grader after seeing results. Record disagreement rates and treat material disagreement as a result, not noise to delete.

## 8. Trial Design, Sampling, Denominators, and Uncertainty

| Trial group | Population   | Sample method | Cases | Repetitions | Denominator rule | Retry and correction rule | Uncertainty method | Stop rule |
| ----------- | ------------ | ------------- | ----- | ----------- | ---------------- | ------------------------- | ------------------ | --------- |
| TRL-001     | [Population] | [Method]      | [IDs] | [Count]     | [Rule]           | [Rule]                    | [Method]           | [Rule]    |

Define invalid-run criteria before execution. Keep invalid runs visible with reasons. Include blocked, aborted, retried, rejected, escalated, and manually corrected work according to the precommitted denominator rules.

Report distributions, tails, and material segments. A population-weighted estimate and an adversarial challenge-set result answer different questions, so label both.

## 9. Security, Authority, Failure, and Human-Work Evaluation

### Required scenario groups

- Direct and indirect prompt injection
- Memory poisoning and unsafe retrieved content
- Tool-result, file, browser, message, and multi-agent injection
- Cross-tenant access using dedicated synthetic tenant pairs
- Missing, denied, expired, altered, and replayed approvals
- Privilege escalation and confused-deputy paths
- Dependency timeout, malformed result, partial effect, duplicate, and late result
- Cancellation, cleanup, recovery, rollback, and child-agent shutdown
- Human correction, escalation, disagreement, queue load, and reviewer absence

Bind approvals to the exact principal, tenant, operation, parameters, target, policy, expiry, and unique operation or approval nonce. Current authorization must be checked for each operation.

Use isolated, resettable environments, inert destinations, controlled network access, resource limits, and an observer with authority to halt execution. Confirm that new and in-flight runs, sessions, queues, callbacks, retries, schedules, credentials, and child agents stop during cleanup.

## 10. Execution, Evidence, and Result Accounting

### Run record

| Field                                                    | Value                                                     |
| -------------------------------------------------------- | --------------------------------------------------------- |
| Run and trial group IDs                                  | [IDs]                                                     |
| Plan, claim, question, metric, and case IDs              | [IDs]                                                     |
| System and configuration                                 | [IDs and digests]                                         |
| Dataset, fixture, and grader versions                    | [Versions]                                                |
| Tenant, principal, agent, and integration identities     | [Scoped references]                                       |
| Authorization outcome and policy version                 | [Values]                                                  |
| Exact context versions and bundle digest                 | [Values]                                                  |
| Direct or Group route source and effective delivery mode | [Values]                                                  |
| Start, end, and environment                              | [Values]                                                  |
| Outcome and status                                       | [Passed / failed / blocked / aborted / invalid / not run] |
| Retry, rejection, escalation, or manual correction       | [Details]                                                 |
| Raw evidence references                                  | [Protected references]                                    |
| Target-state verification                                | [Result]                                                  |
| Reviewer and review time                                 | [Values]                                                  |

### Evidence controls

Record evidence source, authentication and provenance, trust level, and attestation reference. Distinguish compilation, response issuance, integration acknowledgment, host-confirmed injection, and authenticated model-consumption attestation. One stage does not prove the next.

Minimize raw evidence collection. Apply tenant scope, access control, encryption, retention, deletion, and audited reads. Treat active content as untrusted data and encode it for every output sink.

An authenticated, access-controlled, append-only commitment can provide evidence of changes after the commitment time. It does not prove that capture was complete, accurate, authorized, or tied to the real execution. Claims require authenticated provenance, trusted capture, completeness checks, and protected target-system verification.

### Evaluation result summary

| Result field                                                  | Value     |
| ------------------------------------------------------------- | --------- |
| Passed, failed, blocked, aborted, invalid, and not-run totals | [Counts]  |
| Retried, rejected, escalated, and manually corrected totals   | [Counts]  |
| Failed claims, questions, metrics, cases, and run IDs         | [IDs]     |
| Prohibited outcomes and mandatory-gate failures               | [List]    |
| Segment and subgroup results                                  | [Results] |
| Missing evidence and coverage gaps                            | [List]    |
| Uncertainty and grader disagreement                           | [Results] |
| Open defects and incidents                                    | [IDs]     |

## 11. Threshold Hierarchy and Decision Rules

### Threshold and decision-rule register

| Rule ID | Level                                                      | Applies to | Pass   | Conditional pass | Fail   | Insufficient evidence | Invalid | Owner   |
| ------- | ---------------------------------------------------------- | ---------- | ------ | ---------------- | ------ | --------------------- | ------- | ------- |
| THR-001 | [Prohibited / mandatory / dimension / segment / aggregate] | [IDs]      | [Rule] | [Rule]           | [Rule] | [Rule]                | [Rule]  | [Owner] |

Apply rules in this order:

1. Prohibited outcomes
2. Mandatory control gates
3. Dimension thresholds
4. Material segment thresholds
5. Aggregate summaries

Lower levels cannot override higher failures. An aggregate score cannot hide a cross-tenant failure, unauthorized action, missing approval control, severe segment failure, or missing required evidence.

### Decision record

| Field                                   | Value                                                                        |
| --------------------------------------- | ---------------------------------------------------------------------------- |
| Decision                                | [Proceed / proceed with conditions / do not proceed / insufficient evidence] |
| Exact scope and release                 | [Values]                                                                     |
| Evidence cutoff and result IDs          | [Values]                                                                     |
| Mandatory gates                         | [Results]                                                                    |
| Open gaps, defects, and residual risks  | [List]                                                                       |
| Conditions and expiry                   | [List]                                                                       |
| Rollback and stop authority             | [Values]                                                                     |
| Authorized decision maker and reference | [Values]                                                                     |
| Independent reviews                     | [References]                                                                 |

This record informs but does not replace technical authorization, security and privacy approvals, risk acceptance, or production change approval.

## 12. Production Monitoring, Regression, Refresh, and Retirement

| Production measure | Definition   | Sample rate | Baseline | Alert or rollback threshold | Owner   | Response target |
| ------------------ | ------------ | ----------- | -------- | --------------------------- | ------- | --------------- |
| [Metric ID]        | [Definition] | [Rate]      | [Value]  | [Rule]                      | [Owner] | [Time]          |

Define:

- Production outcome verification and sampling
- Privacy, access, retention, and deletion controls
- Drift and distribution checks
- Incident and near-miss intake
- Regression-case promotion and review
- Re-evaluation triggers and scope
- Result expiry and review cadence
- Retirement, evidence retention, and access removal

Review material changes to models, settings, instructions, Knowledge, Skills, Memory behavior, routes, tools, schemas, permissions, approvals, runtime, graders, datasets, population, or monitoring. Create a new plan version and rerun affected evaluation work when the decision boundary changes.

## Change Record

| Version | Date   | Author | Change       | Reason   | Affected claims, measures, data, graders, or rules | Approval    |
| ------- | ------ | ------ | ------------ | -------- | -------------------------------------------------- | ----------- |
| 1.0     | [Date] | [Name] | Initial plan | [Reason] | [IDs]                                              | [Reference] |
