AI evaluation and pilots
Design an AI Pilot Around a Testable Decision
Predeclare the comparator, measures, evaluation cases and guardrails so the pilot can support a clear next decision, including stopping.

The method, in brief
What result would justify the next scope, and what would stop it?
Write the decision protocol before running the pilot. Use a credible comparator and representative permitted cases, record every failure and abstention, and evaluate operating effort alongside quality. Keep revise, extend and stop as legitimate outcomes.
Why this framework matters
A pilot earns its value by changing a decision with inspectable evidence, including a negative result.
Write the decision protocol before running the pilot. Use a credible comparator and representative permitted cases, record every failure and abstention, and evaluate operating effort alongside quality. Keep revise, extend and stop as legitimate outcomes.
Audience and decision
For an operational sponsor, evaluator, technology lead and risk owner deciding whether a proposed intervention should stop, change, continue in a limited setting or progress toward deployment. Use this before building the pilot so that the evaluation cannot quietly become a demonstration of the team's preferred solution.
A pilot should answer a narrow decision question within a defined operating context. It cannot prove that a system is universally safe or that a benefit transfers to every plant. Allow a design workshop of two hours, preparation of representative evaluation material and enough observation time to include relevant variation. The calendar must follow the decision and data, not a marketing promise of a fixed number of days.
Research and source ledger
Accessed 16 September 2026.
Lineage: Hypothesis testing, comparison groups and statistical uncertainty are established practices. dotSuper's combined release decision, exception accounting and operating-readiness checks are original synthesis. All third-party layouts remain separate.
| Source and date | Contribution and limit |
|---|---|
| Eric Ries, Lean Startup principles, undated | Supplies experimentation lineage. Learning requires interpreting evidence, not merely shipping a small product. |
| Alexander Osterwalder, Test Card, 5 March 2015 | Predeclared hypothesis, measure and criterion inform the protocol. The original card is not reproduced. |
| Nabila Amarsy/Strategyzer, Learning Card, 9 March 2015 | Supports recording the decision following a test. This is not validation of dotSuper's pilot template. |
| NIST, AI RMF 1.0, January 2023, and Generative AI Profile, July 2024 | Inform contextual evaluation and risk treatment. Neither provides a universal acceptable error rate. |
| NIST/SEMATECH, Confidence intervals for proportions, undated handbook page | Supports appropriate interval methods when failures or samples are small. Binomial assumptions must hold before using a binomial bound. |
| Brynjolfsson, Li and Raymond, Generative AI at Work, 2023/2025 | Demonstrates the importance of a defined work setting and differences across users. It supplies no expected effect size for this pilot. |
Existing approaches and limitations
A demo establishes that selected examples can work. It does not establish typical performance or net operational benefit. Before-and-after comparisons can be confounded by workload, staffing, seasonality and learning. A single model accuracy figure can hide severe errors, uneven performance and the cost of human review. A pilot with enthusiastic volunteers may not represent the staff who will later use the system.
The proposed design separates technical quality, workflow effect, human use, cost and risk. It predeclares the main decision criterion and guardrails. It also permits an inconclusive outcome. A pilot that cannot distinguish improvement from noise should not automatically become a rollout because the project deadline arrived.
Structure and rationale
Create a decision protocol, an evaluation set register, a results ledger and a release record. The protocol specifies hypothesis, comparator, unit, population, measures, guardrails, scope and decision rules. The register records provenance, inclusion and exclusion, relevant subgroups and separation from development material.
The results ledger includes every eligible item, including abstentions, failures, timeouts and manual recovery. The release record distinguishes observed result, interpretation, unresolved risk and authorised next scope. A technically successful pilot can still stop if operating cost or ownership is unacceptable.
Inputs and preparation
Bring a baseline workflow, a permitted representative sample, clear quality definitions, a comparator and an evaluator with domain competence. Obtain the future operator's participation and the risk owner's agreement on containment. Record the model, configuration, prompt or rules version and data version.
Choose the evaluation unit carefully. One document may contain many fields, but those fields are correlated. Ten pages from one supplier are not ten independent suppliers. Do not inflate the sample by counting related observations as independent. If statistical inference is important, plan sample size and analysis with appropriate expertise using expected variation and a decision-relevant effect.
Facilitation and use
1. Write the decision question, 15 minutes. State which intervention, for whom, under what conditions and compared with what. Define what action follows each possible result. 2. Select the comparator, 15 minutes. Use current competent practice and, when relevant, a simpler process or rules-based alternative. Keep the quality standard equal across arms. If the comparison cannot be fair, document the limitation before running it. 3. Predeclare measures, 20 minutes. Choose one primary operational measure, such as total active review time per eligible item. Add critical quality, security and harm guardrails. Include setup and exception burden separately so a narrow time metric does not hide transferred work. 4. Prepare representative cases, 20 minutes. Include ordinary work and meaningful difficult cases. Tag language, document type, supplier, shift and other relevant variation. Keep tuning cases separate from the final evaluation and control access to the evaluation answers. 5. Choose a comparison design, 15 minutes. Random assignment may be feasible for low-risk work. A paired design can reduce item variation but requires counterbalancing order and attention to memory effects. If only a before-and-after study is feasible, record likely confounders and narrow causal claims. 6. Run within containment. Begin offline or in shadow mode where appropriate. Keep consequential actions under existing controls. Record all exclusions and changes. If a serious unexpected failure occurs, follow the stopping rule rather than continuing to improve the average. 7. Analyse without hiding failures, 20 minutes. Report denominators, distributions and subgroup results. Compare total review effort and recovery, not just generation time. Describe uncertainty and distinguish observed difference from causally attributable improvement. 8. Make the next decision, 15 minutes. Record stop, revise, extend evaluation or limited release, with scope and conditions. Check ownership and fallback readiness before any operational expansion.
Measures, thresholds and assumptions
Possible measures include critical-field retention, unsupported assertion rate, abstention rate, reviewer correction time and total workflow lead time. Define each numerator and denominator. “Accuracy” without a task and failure definition is insufficient. Reviewers should resolve disagreements against explicit criteria; an AI judge alone should not determine consequential domain correctness.
Set thresholds according to business consequence and baseline. A target of 20% lower review time is an illustrative decision threshold only if the team chooses it before seeing results. Do not use the same tolerated error rate for a draft summary and a safety-critical recommendation.
Zero observed failures does not prove zero underlying risk. Under independent, identically distributed Bernoulli trials, zero failures in 100 trials gives an exact one-sided 95% upper bound of approximately 2.95%, from 1 minus 0.05 raised to the power 1/100. Correlated cases or distribution change undermine that interpretation. Severe harms may require specialised assurance beyond any small pilot.
Worked example: illustrative, not a dotSuper result
A hypothetical procurement team tests assisted comparison on 100 eligible historical quotation packages. Twenty additional packages form a development set and never enter the final reported evaluation. The team predefines missing payment or delivery terms as critical errors and evaluates source traceability.
Competent manual review is the comparator. Reviewers use a counterbalanced allocation so one person does not simply remember their previous answer. Total active time includes reading, checking and correction. The primary target is a locally chosen reduction of at least 20%, with no increase in critical errors and no unauthorised data exposure. These are example design choices, not recommended universal thresholds.
Suppose measured median active review time falls from 12 to 8 minutes, but three assisted outputs omit a critical delivery qualification. The team cannot average those failures away with the time benefit. It investigates whether source extraction, presentation or reviewer behaviour caused the misses. The current version does not meet the quality guardrail.
The decision is revise, not deploy. A change that highlights uncertain clauses is tested on a fresh evaluation set, while the original failures remain in the issue history. If the team instead sees no failures, it still reports the limited sample and tests live exception handling before expanding. A successful offline evaluation does not establish adoption, long-run economics or production safety.
Outputs, failure modes and validation limits
Produce a dated protocol, evaluation-set manifest, results with denominators, incident and exclusion log, and a signed next-decision record. Preserve negative results. A useful pilot may establish that a conventional template performs as well at lower cost.
Failures include changing success criteria after seeing results, testing on training examples, omitting abstentions, excluding difficult cases without disclosure and reporting only the average. Another is claiming cash savings from faster work without a value-capture plan. A counterexample is a rare-event workflow where a short pilot cannot observe the relevant failure often enough. Longer evaluation, simulation and domain assurance may be required.
Validate this protocol through independent review of the design and reproducibility of the analysis. The dotSuper template itself has not been evaluated as an intervention. Its purpose is to make scope and evidence inspectable, not confer statistical or safety authority.
Working files and reuse
The five-page PDF includes a visual reference, two fillable worksheet pages, an illustrative worked example and a facilitator/source guide. An expandable CSV working log is also available.
All examples are illustrative, not measured client results. Original pilot and release record informed by established experimentation. Not field validated. A small pilot cannot establish rare-event safety, long-run economics or sustained adoption. Zero observed failures does not mean zero risk. No Test Card layout is reproduced.
Original dotSuper material prepared for review. No public reuse licence has been assigned. Third-party source material retains its own terms.
The PDF is not represented as a tagged PDF/UA document. The text on this page provides a readable alternative to the diagram and method.
See the method. Keep the context.
The visual companion

Reuse: Original dotSuper material. No public reuse licence has been specified. Contact dotSuper for reuse permissions. Third-party source material retains its own terms.
Read the diagram: Pilot Design and Evaluation. Original dotSuper reference diagram. Hypothesis testing, comparators and statistical uncertainty are established. Release scope, exception accounting and operating-readiness records are dotSuper synthesis.
Before the run, declare the question, population, unit, thresholds and guardrails. Compare competent current practice with a fixed version and scope of the intervention. Measure quality, effort, recovery and cost. Retain failures and abstentions and report denominators and uncertainty. Decide to stop, revise, extend evaluation or limit release. Faster performance cannot compensate for an unmet consequential guardrail.
Original pilot and release record informed by established experimentation. Not field validated. A small pilot cannot establish rare-event safety, long-run economics or sustained adoption. Zero observed failures does not mean zero risk. No Test Card layout is reproduced.
| Measure | Illustrative value | Scope |
|---|---|---|
| Manual median active review time | 12 minutes | Reading, checking and correction |
| Assisted median active review time | 8 minutes | Same operational measure |
| Assisted outputs with critical omissions | 3 of 100 | Critical delivery qualification missing |
| Predeclared time target | At least 20% reduction | Local example threshold, not a universal target |
| Next decision | Revise | Time benefit does not resolve the quality issue |
Take it into your next working session
Keep the source credits with the file. Check the reuse terms and adapt the method to your context.
Reuse: Original dotSuper material. No public reuse licence has been specified. Contact dotSuper for reuse permissions. Third-party source material retains its own terms.
Reuse: Original dotSuper material. No public reuse licence has been specified. Contact dotSuper for reuse permissions. Third-party source material retains its own terms.
Thumbnail credit and reuse
Reuse: Original dotSuper material. No public reuse licence has been specified. Contact dotSuper for reuse permissions. Third-party source material retains its own terms.
Sources, context and limits
Keep the evidence beside the method.
- Original pilot and release record informed by established experimentation. Not field validated. A small pilot cannot establish rare-event safety, long-run economics or sustained adoption. Zero observed failures does not mean zero risk. No Test Card layout is reproduced.
- This combined dotSuper method is research-informed and has not been field validated. Workshop agreement and a completed template are not proof of effectiveness.
- Worked examples are hypothetical. Country examples and intended regional audience do not establish country-wide or region-wide effectiveness.
- Source access and adaptation limits are recorded in the source ledger. Attribution does not imply endorsement or a licence to reproduce third-party artwork.
- Lean Startup principles
theleanstartup.com · accessed Sep 16, 2026
- Test Card
Strategyzer · accessed Sep 16, 2026
- Learning Card
Strategyzer · accessed Sep 16, 2026
- AI RMF 1.0
nvlpubs.nist.gov · accessed Sep 16, 2026
- Generative AI Profile
nvlpubs.nist.gov · accessed Sep 16, 2026
- Confidence intervals for proportions
itl.nist.gov · accessed Sep 16, 2026
- Generative AI at Work
nber.org · accessed Sep 16, 2026
/ CITE OR SHARE THIS GUIDE
Make the evidence easy to verify.
When you reference this guide, link to its canonical URL. That gives readers one stable place for the evidence, limitations and future updates.
dotSuper Research Desk. (September 17, 2026). Design an AI Pilot Around a Testable Decision. dotSuper. https://dotsuper.net/feeds/applied-systems/pilot-design-and-evaluation
Bring it into the work
Start with one real decision
Bring one pilot idea and define the comparison, evidence and stopping rules.
Download the fillable framework