# AI Agent Test Plan Template

Use this template to plan and record testing for one AI agent workflow and one bounded release. Replace every bracketed field. Delete examples that do not apply, but record why a material test category is excluded.

This plan does not grant technical access, runtime authority, pilot approval, or production authorization. Every operation still requires current authorization at execution time. Release and risk owners must make their own decisions from the test evidence and other required reviews.

Do not put production secrets, credentials, private model reasoning, or unnecessary personal data in this document or its fixtures. Use references to protected evidence where the test plan does not need the underlying content.

## Test Plan Record

| Field                           | Value                                                                                                                                                                                                                                                                                             |
| ------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Test plan ID                    | [Stable ID]                                                                                                                                                                                                                                                                                       |
| Workflow and agent              | [Name and stable IDs]                                                                                                                                                                                                                                                                             |
| Release under test              | [Immutable release or configuration ID]                                                                                                                                                                                                                                                           |
| Requirements baseline           | [Version and approval reference]                                                                                                                                                                                                                                                                  |
| Risk assessment                 | [Version and reference]                                                                                                                                                                                                                                                                           |
| Test owner                      | [Named person]                                                                                                                                                                                                                                                                                    |
| Business outcome owner          | [Named person]                                                                                                                                                                                                                                                                                    |
| Security reviewer               | [Named person / Not required under recorded exemption]                                                                                                                                                                                                                                            |
| Independent reviewer            | [Named person / Not required under recorded exemption]                                                                                                                                                                                                                                            |
| Review requirement or exemption | [Governing policy and risk tier; authorized approver; approval record; scope; expiry; separation-of-duties check]                                                                                                                                                                                 |
| Plan version                    | [Version]                                                                                                                                                                                                                                                                                         |
| Status                          | [Draft / Approved for execution / Executing / Complete / Superseded]                                                                                                                                                                                                                              |
| Approved scope                  | [Workflow, users, tenant, environment, and authority]                                                                                                                                                                                                                                             |
| Execution authorization         | [Authenticated approver identity and authority; governing policy; decision reference; exact plan version, release/configuration digest, approved case and fixture versions, and evidence configuration; scope and environment; permitted effects and limits; expiry; separation-of-duties result] |
| Evidence location               | [Protected repository reference]                                                                                                                                                                                                                                                                  |
| Planned execution window        | [Start and end]                                                                                                                                                                                                                                                                                   |
| Decision deadline               | [Date and time]                                                                                                                                                                                                                                                                                   |
| Next review or expiry           | [Date or trigger]                                                                                                                                                                                                                                                                                 |

The status field does not authorize test execution. Execution requires the authenticated decision record above and current technical authorization for every operation. Bind that decision to the exact plan version, release or configuration digest, approved cases, fixture versions, and evidence configuration. Require fresh approval after any material change. The plan author or test owner cannot self-approve execution where policy requires independence. The plan author cannot waive a required review. Any exemption must come from an authorized approver under the governing policy, state its scope and expiry, and preserve required separation of duties.

## Test Rules

- Freeze expected results, measures, thresholds, denominator rules, and severity rules before execution.
- Keep expected results separate from observed results.
- Use one stable test case ID for one independently passable behavior or control.
- Link every case to a requirement, risk, control, incident, or operating objective.
- Record every planned run as passed, failed, blocked, aborted, invalid, or not run. Do not silently remove unfavorable runs.
- Count retries, manual corrections, refusals, escalations, and partial completions under precommitted rules.
- Treat an average pass rate as insufficient when a prohibited outcome or required control fails.
- Verify consequential external effects through protected target-system state or audit records, not only through the agent trace. Never repeat an effect unless the environment is isolated and inert and the repetition has explicit approval.
- Keep test credentials, data, networks, and destinations bounded to the test environment.
- Preserve raw evidence separately from reviewer conclusions and derived scores.

## 1. Test Decision, Claims, and Objectives

### Decision this plan informs

[State the exact decision, such as whether release R-17 may enter a bounded internal pilot. Name the decision owner and evidence cutoff.]

### Claims under test

| Claim ID  | Scoped claim | Population and conditions    | Supporting test IDs | Required threshold | Decision owner |
| --------- | ------------ | ---------------------------- | ------------------- | ------------------ | -------------- |
| CLM-[###] | [Claim]      | [Who, what, where, and when] | [IDs]               | [Threshold]        | [Owner]        |

### Test objectives

1. [Objective tied to a requirement, risk, or decision]
2. [Objective]
3. [Objective]

### Non-claims

- [What this plan cannot establish]
- [Conditions, users, languages, tools, or authority outside the evidence]
- [Properties covered by another assessment]

## 2. System Under Test and Frozen Configuration

| Component                         | Stable ID or version               | Expected state                                      | Confirmation source | Material to decision? |
| --------------------------------- | ---------------------------------- | --------------------------------------------------- | ------------------- | --------------------- |
| Agent definition                  | [ID and version]                   | [State]                                             | [Reference]         | [Yes / No]            |
| Model and provider                | [Deployment and version]           | [Settings]                                          | [Reference]         | [Yes / No]            |
| System and developer instructions | [Version or digest]                | [State]                                             | [Reference]         | [Yes / No]            |
| Published Knowledge               | [IDs and versions]                 | [Authority and freshness]                           | [Reference]         | [Yes / No]            |
| Published Skills                  | [IDs, versions, package digests]   | [State]                                             | [Reference]         | [Yes / No]            |
| Working Memory                    | [ID and version or prepared state] | [State]                                             | [Reference]         | [Yes / No]            |
| Context routes                    | [Route IDs and modes]              | [Expected delivery]                                 | [Reference]         | [Yes / No]            |
| Retrieval sources                 | [Indexes and snapshots]            | [State]                                             | [Reference]         | [Yes / No]            |
| Tools and integrations            | [IDs and versions]                 | [Allowed operations]                                | [Reference]         | [Yes / No]            |
| Identity and permissions          | [Principal and grants]             | [Scope and limits]                                  | [Reference]         | [Yes / No]            |
| Approval controls                 | [Policy version]                   | [Checkpoints and binding]                           | [Reference]         | [Yes / No]            |
| Runtime and sandbox               | [Build and policy]                 | [Filesystem, network, code, and persistence limits] | [Reference]         | [Yes / No]            |
| Monitoring and evidence           | [Configuration version]            | [Signals and retention]                             | [Reference]         | [Yes / No]            |

Recording a context version does not pin an Always route or prove delivery. Permissions govern repository access, while routing independently governs automatic delivery. The harness must arrange the intended test state and confirm the compiled bundle. Response issuance or acknowledgment does not prove host injection or model consumption.

### Configuration verification before each run batch

- [ ] The release identifier matches the approved test subject.
- [ ] The environment reset completed.
- [ ] Expected identities, grants, routes, context, tools, policies, and limits are present.
- [ ] Unexpected access, context, external destinations, schedules, queues, and callbacks are absent.
- [ ] Clocks, random seeds where supported, model settings, and dependency versions are recorded.
- [ ] The evidence pipeline and target-system confirmation checks are working.

## 3. Scope and Operating Envelope

### Included

- Users and initiating principals: [Scope]
- Workflow paths: [Scope]
- Data classes and jurisdictions: [Scope]
- Languages and channels: [Scope]
- Tools, actions, and transaction limits: [Scope]
- Human checkpoints: [Scope]
- Dependency states: [Normal, degraded, and unavailable states]
- Load, latency, cost, token, and concurrency bounds: [Scope]

### Excluded

| Exclusion | Reason and risk   | Governing policy and authority | Authorized approver and decision reference | Scope and expiry   | Separation-of-duties evidence | Owner and required future test or control |
| --------- | ----------------- | ------------------------------ | ------------------------------------------ | ------------------ | ----------------------------- | ----------------------------------------- |
| [Item]    | [Reason and risk] | [Policy and authority]         | [Approver and reference]                   | [Scope and expiry] | [Reference]                   | [Owner and action]                        |

The plan author cannot self-approve an exclusion. Excluding a mandatory requirement, risk, or control must follow the governing policy's authorized risk-acceptance or exemption path.

### Unacceptable outcomes

- [Unauthorized, cross-tenant, irreversible, unsafe, deceptive, discriminatory, or uncontained outcome]
- [Secret or unnecessary personal-data exposure]
- [Action after denied, expired, missing, or mismatched approval]
- [False claim of completion when the target system did not confirm the effect]

## 4. Traceability and Coverage Model

Every planned test case must have exactly one row in this master register. Each requirement, material risk, required control, and relevant prior incident must map to at least one case or have an approved exclusion with a reason.

| Test ID   | Requirement, risk, control, or incident IDs | Suite   | Priority | Method   | Expected result     | Threshold   | Evidence    | Owner   | Status   |
| --------- | ------------------------------------------- | ------- | -------- | -------- | ------------------- | ----------- | ----------- | ------- | -------- |
| TST-[###] | [IDs]                                       | [Suite] | [P0-P3]  | [Method] | [Observable result] | [Pass rule] | [Reference] | [Owner] | [Status] |

### Coverage summary

| Coverage dimension          | Population          | Included | Excluded   | Coverage method | Gap owner | Decision impact |
| --------------------------- | ------------------- | -------- | ---------- | --------------- | --------- | --------------- |
| Requirements                | [Count and version] | [Count]  | [Count]    | [Mapping]       | [Owner]   | [Impact]        |
| Risks and controls          | [Count and version] | [Count]  | [Count]    | [Mapping]       | [Owner]   | [Impact]        |
| Workflow paths              | [Population]        | [Count]  | [Count]    | [Method]        | [Owner]   | [Impact]        |
| Users and affected groups   | [Population]        | [Sample] | [Excluded] | [Method]        | [Owner]   | [Impact]        |
| Data and context conditions | [Population]        | [Sample] | [Excluded] | [Method]        | [Owner]   | [Impact]        |
| Tools and dependency states | [Population]        | [Sample] | [Excluded] | [Method]        | [Owner]   | [Impact]        |

## 5. Test Environment, Data, and Containment

| Environment | Purpose   | Production differences | Data class                       | Credentials            | Network and destination limits | Reset method | Owner   |
| ----------- | --------- | ---------------------- | -------------------------------- | ---------------------- | ------------------------------ | ------------ | ------- |
| [Name]      | [Purpose] | [Differences]          | [Synthetic, masked, or approved] | [Scoped test identity] | [Allowlist]                    | [Method]     | [Owner] |

### Fixture and dataset register

| Dataset ID and version | Source and license | Population represented | Known gaps | Label or expected-result owner | Contamination control | Retention and disposal |
| ---------------------- | ------------------ | ---------------------- | ---------- | ------------------------------ | --------------------- | ---------------------- |
| DAT-[###]              | [Source]           | [Population]           | [Gaps]     | [Owner]                        | [Control]             | [Rule]                 |

Use synthetic, masked, or specifically approved data. Never put live secrets, credentials, unrestricted customer records, or private model reasoning in fixtures. Separate solution keys, grader rubrics, and hidden cases from the agent's accessible tools, retrieval sources, logs, and working directories. Cross-tenant and adversarial cases must use dedicated synthetic tenant or account pairs. Never probe an uninvolved tenant, account, user, or external system.

Dangerous tests require a resettable sandbox, scoped identities, controlled network access, inert or test-only destinations, resource limits, stop controls, and an owner watching execution. Define how to confirm that new and in-flight runs, credentials, sessions, queues, schedules, callbacks, retries, child agents, and external work have stopped.

Do not use production shadow, read-only, or canary testing until the lower test levels pass. Require every applicable system owner, data owner, security, privacy, risk, and change-control approval. Use synthetic or masked data by default. When approved production data is unavoidable, minimize the permitted fields and records, use scoped identities, block unapproved external egress, redact retained evidence, define retention and deletion rules, and stop immediately if data escapes the approved boundary. Enforce read-only behavior technically, bind the canary to approved users, data, systems, external effects, limits, and stop rules, and assign an observer with authority to halt execution.

## 6. Test Suite and Case Design

### Required suite categories

- Deterministic component and schema tests
- Context compilation, routing, authority, freshness, and conflict tests
- Functional outcome and trajectory tests
- Tool selection, parameters, results, retries, duplicates, and target-state tests
- Identity, permission, tenant, approval, and delegation tests
- Data quality, privacy, retention, and deletion tests
- Human review, handoff, override, appeal, and capacity tests
- Adversarial security and abuse-case tests
- Dependency failure, recovery, reconciliation, and safe-stop tests
- Multi-agent handoff, shared-state, loop, and failure-containment tests where applicable
- Performance, latency, cost, load, token, and resource-limit tests
- Accessibility, language, and user-experience tests where applicable
- Regression cases from prior incidents, defects, overrides, and production drift

### Test case record

| Field                                              | Value                                         |
| -------------------------------------------------- | --------------------------------------------- |
| Test ID and version                                | [ID and version]                              |
| Linked requirement, risk, control, or incident IDs | [IDs]                                         |
| Purpose                                            | [What failure this case can detect]           |
| Preconditions                                      | [Required state]                              |
| Initiating principal and tenant                    | [Identity and scope]                          |
| Input and fixture IDs                              | [References]                                  |
| Expected context and authority                     | [Versions, routes, and instruction authority] |
| Allowed tools and operations                       | [List]                                        |
| Forbidden actions and states                       | [List]                                        |
| Required approval or escalation                    | [Binding details]                             |
| Expected path constraints                          | [Required checkpoints and forbidden paths]    |
| Expected outcome and target-system state           | [Observable result]                           |
| Scoring method and grader version                  | [Method]                                      |
| Pass, fail, and invalid rules                      | [Rules]                                       |
| Required repetitions and stop rule                 | [Count and rule]                              |
| Evidence to retain                                 | [References]                                  |
| Cleanup and reconciliation                         | [Steps]                                       |
| Reviewer                                           | [Named person or role]                        |

Freeze the case before execution. A reviewer may mark a run invalid only under a prewritten invalidation rule and must preserve the run record and reason.

## 7. Functional, Context, Tool, and Authority Tests

### Functional and outcome tests

- Normal and boundary inputs
- Ambiguous, incomplete, conflicting, stale, and malformed inputs
- Unsupported requests and scope expansion
- Required fields, calculations, citations, and target-state confirmation
- Honest reporting of failed, blocked, partial, rejected, and escalated work

### Context tests

- Required published Knowledge and Skills and current Memory state arrive as intended.
- Content that is not authorized or eligible for delivery under the effective direct or Group routes does not appear. Test repository discovery and read access separately from routed delivery.
- Instruction authority and conflict resolution match the approved rules.
- Artifacts and messages have no instruction authority. Treat them as inputs unless an authorized instruction separately applies.
- MCP capability descriptors and approved configuration remain separate from MCP tool results, which are Runtime context.
- Secrets and private model reasoning never enter managed context, fixtures, or evidence.

### Tool and external-effect tests

- Tool selection and parameter validation
- Read versus write boundaries
- Idempotency, duplicate detection, retry, timeout, cancellation, and late result handling
- Partial writes and compensation
- Target-system confirmation and reconciliation
- Schema, authentication, rate-limit, and dependency changes

### Identity, authorization, and approval tests

- Current authorization is checked for every operation.
- Cross-tenant attempts inside dedicated synthetic tenant or account pairs, plus out-of-scope, excessive, expired, revoked, and delegated access, are denied. Never test an uninvolved tenant or identity.
- Approval is bound to the tenant or workspace, principal, unique operation or approval nonce, canonical operation, exact release, tool and schema versions, action, parameters, target, target resource version or state preconditions, amount or limit, expiry, and policy version.
- Denied, expired, missing, replayed, cross-tenant, or modified approval cannot authorize an action. Reject execution when any bound configuration, schema, parameter, target, or target-state precondition changed after approval.
- The agent stops or escalates when the required human decision is unavailable.

## 8. Measures, Repeated Trials, and Human Review

| Measure ID | Linked test IDs | Population and conditions | Numerator    | Denominator  | Sample and repetitions | Method or grader     | Threshold   | Stop threshold   | Reviewer |
| ---------- | --------------- | ------------------------- | ------------ | ------------ | ---------------------- | -------------------- | ----------- | ---------------- | -------- |
| MET-[###]  | [IDs]           | [Definition]              | [Definition] | [Definition] | [Rule]                 | [Method and version] | [Threshold] | [Immediate stop] | [Owner]  |

Record distributions, variability, tail failures, and uncertainty where they affect the decision. Do not report only an average when a rare prohibited outcome matters. Equivalent repeated trials must use the same frozen case and configuration unless variation is the factor under test.

Human and model-based graders need calibration cases, decision rules, disagreement handling, blind or independent review where required, and recorded versions. Validate automated graders against trusted examples and inspect false passes and false failures. A grader score is evidence from one method, not the decision itself.

## 9. Security, Fault, Recovery, and Multi-Agent Tests

### Security abuse cases

- Direct and indirect prompt injection
- Goal manipulation and policy conflict
- Unauthorized tool use and privilege escalation
- Cross-tenant access inside dedicated synthetic tenant or account pairs, plus confused-deputy requests
- Secret, personal-data, and context exfiltration
- Memory poisoning and poisoned retrieval, message, and tool-result content
- Approval bypass, replay, parameter change, and stale approval
- Unsafe code, command, filesystem, browser, or network use
- Recursive tool use, runaway loops, cost exhaustion, and denial of service
- Agent-to-agent trust abuse and identity spoofing

### Fault and recovery cases

- Tool timeout, malformed result, partial result, and unavailable dependency
- Lost state, duplicate callback, retry after partial completion, and clock skew
- Reviewer unavailable, queue overload, and expired hold
- Monitor, evidence, policy, or approval service unavailable
- Stop, rollback, compensation, target reconciliation, and known-good restore

### Multi-agent cases

- Handoff schema, provenance, identity, and minimum necessary context
- Delegated authority boundaries. Current authorization must be checked for each operation at each receiving agent.
- Conflicting goals, shared-state races, cycles, repeated delegation, and cascading failure
- Compromised peer messages and untrusted tool results
- Final outcome ownership and attribution across the chain

## 10. Execution Protocol and Evidence

### Execution sequence

1. Verify the frozen release, environment, data, identities, context, tools, controls, and evidence capture.
2. Record the planned case set and run identifiers before execution.
3. Execute deterministic and low-risk suites before dangerous or costly suites.
4. Stop on a defined critical failure or containment breach.
5. Preserve raw evidence, observed target state, and environment state before cleanup.
6. Classify the run without changing the expected result or threshold.
7. Reconcile external effects, reset the environment, and confirm cleanup.
8. Review invalid, blocked, aborted, retried, and manually corrected runs.

Minimize raw evidence collection. Exclude credentials, authorization headers, unnecessary personal data, and live secrets. When a protected original is required, quarantine it under access controls and retention limits, then provide a redacted review copy. Treat prompt injections, malicious tool results, and other active payloads as untrusted data, and encode them for each log, Markdown, HTML, terminal, or other rendering sink. Record who can access the original, when it expires, and how disposal is confirmed.

### Run record

| Field                                                     | Value                                                                                                                                                   |
| --------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Run and trace IDs                                         | [Stable IDs]                                                                                                                                            |
| Test ID and version                                       | [ID]                                                                                                                                                    |
| Start and end time                                        | [Times]                                                                                                                                                 |
| Executor and reviewer                                     | [Authenticated identities]                                                                                                                              |
| Release and environment                                   | [IDs]                                                                                                                                                   |
| Configuration and data snapshot                           | [References]                                                                                                                                            |
| Context compilation record                                | [Reference]                                                                                                                                             |
| Response issuance, acknowledgment, and injection evidence | [Separate references or unknown]                                                                                                                        |
| Model-consumption evidence                                | [Attesting integration or provider; authentication and provenance; trust level; attestation reference; otherwise unknown]                               |
| Tool calls and approvals                                  | [References]                                                                                                                                            |
| Target-system confirmation                                | [Reference]                                                                                                                                             |
| Observed outcome                                          | [Result]                                                                                                                                                |
| Score and method                                          | [Result and grader version]                                                                                                                             |
| Status                                                    | [Pass / Fail / Blocked / Aborted / Invalid / Not run]                                                                                                   |
| Failure category and severity                             | [Category]                                                                                                                                              |
| Cleanup and reconciliation                                | [Result]                                                                                                                                                |
| Evidence integrity reference                              | [Authenticated protected record or verified signature; anchored digest labeled post-anchor tamper evidence; unanchored digest labeled consistency-only] |

Compilation, response issuance, integration acknowledgment, host-confirmed injection, and agent consumption are different evidence stages. One stage does not prove the next. Record consumption only from direct, authenticated attestation by a trusted integration or provider, including the evidence source, authentication and provenance, trust level, and attestation reference; otherwise record it as unknown.

An anchored digest can provide evidence of changes after the commitment time when the anchor is authenticated, access-controlled, append-only or independently retained, and timestamped. It does not prove that capture was complete, truthful, from the claimed source, or unchanged before anchoring. Establish those claims separately through authenticated provenance, trusted capture, completeness checks, and timestamps. A verified signature also needs protected key provenance. A digest stored beside mutable evidence is only a consistency check because an actor who can replace the evidence may replace the digest too.

## 11. Defects, Retest, and Release Decision

### Defect register

| Defect ID | Failed test and run IDs | Severity   | Observed impact | Root cause status             | Owner   | Fix version | Required retests | State   |
| --------- | ----------------------- | ---------- | --------------- | ----------------------------- | ------- | ----------- | ---------------- | ------- |
| DEF-[###] | [IDs]                   | [Severity] | [Impact]        | [Known / Suspected / Unknown] | [Owner] | [Version]   | [IDs]            | [State] |

Never weaken a case, threshold, fixture, grader, or denominator merely to make a release pass. Changes to expected behavior require an approved requirements change and an independent explanation of why the old expectation was wrong.

### Release decision record

| Field                                                         | Value                                          |
| ------------------------------------------------------------- | ---------------------------------------------- |
| Decision                                                      | [Pass / Pass with conditions / Fail / Blocked] |
| Exact release and scope                                       | [IDs and boundaries]                           |
| Evidence cutoff                                               | [Time]                                         |
| Required suites executed                                      | [Counts and references]                        |
| Passed, failed, blocked, aborted, invalid, and not-run totals | [Counts]                                       |
| Open defects by severity                                      | [Counts and IDs]                               |
| Uncovered requirements or risks                               | [IDs and reasons]                              |
| Accepted residual risk                                        | [Authenticated decision reference]             |
| Conditions and expiry                                         | [Conditions]                                   |
| Decision maker and authority                                  | [Authenticated identity and role]              |
| Protected decision evidence                                   | [Reference]                                    |

The test decision applies only to the exact release, configuration, population, environment, and authority tested. It does not grant access, authorize deployment, accept residual risk outside the decision maker's scope, or prove behavior under untested conditions.

## 12. Regression, Monitoring, Change, and Retirement

### Regression triggers

- Agent instructions, policy, model, runtime, tools, schemas, or integrations change
- Knowledge, Skill, Memory, retrieval, or route behavior changes materially
- Identity, permissions, approval logic, transaction limits, or tenant scope changes
- New user group, language, channel, data class, geography, or workflow enters scope
- A defect, override, appeal, incident, near miss, or monitoring signal reveals a new failure mode
- A dependency or external obligation changes
- A test, dataset, grader, threshold, or evidence method changes

| Trigger   | Tests to rerun or add | Owner   | Deadline | Release or operation response                |
| --------- | --------------------- | ------- | -------- | -------------------------------------------- |
| [Trigger] | [IDs or suite]        | [Owner] | [Time]   | [Block, suspend, limit, monitor, or proceed] |

Production monitoring may reveal failures that the plan missed, but monitoring does not replace pre-release testing. Add confirmed failures, meaningful overrides, and material near misses to the regression suite with protected evidence and privacy review.

At retirement, verify that test identities, credentials, grants, routes, integrations, schedules, queues, callbacks, sandboxes, fixtures, and temporary data have been removed or handled under approved retention rules. Reconcile pending external work and preserve the required decision and test evidence.

## Change Record

| Version   | Date   | Changed sections and test IDs | Reason and source | Required retest | Author   | Reviewer   | Decision reference |
| --------- | ------ | ----------------------------- | ----------------- | --------------- | -------- | ---------- | ------------------ |
| [Version] | [Date] | [Items]                       | [Reason]          | [IDs]           | [Author] | [Reviewer] | [Reference]        |
