Test the failure
before the launch.
Most automation failures are not mysterious. They happen at the boundaries: bad inputs, unclear ownership, stale context, permission mismatch, silent retries, or a decision that should have stayed with a person.
What was measured—and what was not.
This is an original evaluation framework, not a population study. It organizes failure conditions that an implementation team can test against representative examples. No frequency, benchmark, adoption, revenue, safety rate, client result, or ranking outcome is asserted.
For each workflow, record the trigger, source of truth, identity, permissions, model or rule, tool call, state change, human approval, retry behavior, notification, and recovery path. Test ordinary cases, missing fields, duplicates, stale records, provider errors, revoked access, ambiguous outputs, conflicting sources, and irreversible actions.
Reference context: NIST AI Risk Management Framework, OWASP LLM application risk guidance, OpenAI evaluation guidance, and Cloudflare Workers documentation. These are reference points, not proof that a particular system complies with a framework.
1. Input and identity failures
Test malformed payloads, missing required fields, duplicate events, ambiguous contacts, stale identifiers, conflicting records, and input that looks valid but belongs to the wrong entity. The acceptance test is not “the model responded”; it is “the system rejected or routed the wrong input without corrupting the source of truth.”
2. Context and retrieval failures
Test stale documents, incomplete retrieval, contradictory sources, permission leakage, missing citations, and answers that sound certain when evidence is weak. A useful system makes source and uncertainty visible and routes low-confidence cases to review.
3. Tool and state-change failures
Test timeouts, rate limits, partial success, non-idempotent retries, provider schema changes, unauthorized calls, and a successful API response that does not prove the business outcome occurred. Log intent, request, response, state, owner, and reconciliation result.
4. Human control and recovery failures
Test moments where judgment, consent, legal responsibility, financial commitment, or irreversible action belongs to an accountable person. The system should show a clear approval boundary, preserve evidence needed to decide, expose the next action, and make recovery possible.
Failure taxonomy
The taxonomy turns broad concerns into observable test cases. A workflow can have more than one failure class; record the primary owner and the recovery owner separately.
| Failure class | Observable signal | Consequence | Mitigation / test |
|---|---|---|---|
| Input / identity | Missing, duplicated, stale, or mis-associated record | Wrong routing or corrupted source of truth | Schema validation, deduplication, identity checks, rejected-fixture tests |
| Context / retrieval | Unsupported answer, stale source, or permission leakage | Unreliable decision or inappropriate disclosure | Source display, freshness rules, access tests, confidence threshold, human review |
| Tool / authorization | Unexpected call, expired token, or provider rejection | Partial action, data exposure, or blocked workflow | Least privilege, explicit tool allowlist, revocation test, audit event |
| State / retry | Duplicate write, timeout after commit, or unreconciled response | Double action or false completion | Idempotency key, durable state, retry budget, reconciliation fixture |
| Human / recovery | No owner, unclear approval, or no pause / undo path | Unaccountable or irreversible outcome | Approval boundary, escalation SLA, rollback path, operator drill |
Release gate
- Every failure mode has an owner and a response.
- Every consequential action has a defined human approval rule.
- Every external system has a permission and revocation path.
- Representative and adversarial examples pass before release.
- Monitoring can distinguish a completed business outcome from a successful tool call.
Limits and next use
This framework is a practical test design aid. It does not certify safety, legal compliance, model quality, or business value. Record the examples, expected behavior, actual behavior, owner, evidence, and decision for each release.
Use it with the AI automation opportunity calculator, the law-firm intake calculator, or the AI automation service. For the public evidence boundary, see the proof ledger.