AI evaluation and pilots

Design an AI Pilot Around a Testable Decision

Predeclare the comparator, measures, evaluation cases and guardrails so the pilot can support a clear next decision, including stopping.

8 min readUpdated Sep 17, 2026Method and working resources
Pilot Design and Evaluation reference and practical worksheet pack.
dotSuper. Hypothesis testing, comparators and statistical uncertainty are established. Release scope, exception accounting and operating-readiness records are dotSuper synthesis.. A practical working resource.

The method, in brief

What result would justify the next scope, and what would stop it?

Write the decision protocol before running the pilot. Use a credible comparator and representative permitted cases, record every failure and abstention, and evaluate operating effort alongside quality. Keep revise, extend and stop as legitimate outcomes.

Why this framework matters

A pilot earns its value by changing a decision with inspectable evidence, including a negative result.

Write the decision protocol before running the pilot. Use a credible comparator and representative permitted cases, record every failure and abstention, and evaluate operating effort alongside quality. Keep revise, extend and stop as legitimate outcomes.

Audience and decision

For an operational sponsor, evaluator, technology lead and risk owner deciding whether a proposed intervention should stop, change, continue in a limited setting or progress toward deployment. Use this before building the pilot so that the evaluation cannot quietly become a demonstration of the team's preferred solution.

A pilot should answer a narrow decision question within a defined operating context. It cannot prove that a system is universally safe or that a benefit transfers to every plant. Allow a design workshop of two hours, preparation of representative evaluation material and enough observation time to include relevant variation. The calendar must follow the decision and data, not a marketing promise of a fixed number of days.

Research and source ledger

Accessed 16 September 2026.

Lineage: Hypothesis testing, comparison groups and statistical uncertainty are established practices. dotSuper's combined release decision, exception accounting and operating-readiness checks are original synthesis. All third-party layouts remain separate.

Research and source ledger
Source and dateContribution and limit
Eric Ries, Lean Startup principles, undatedSupplies experimentation lineage. Learning requires interpreting evidence, not merely shipping a small product.
Alexander Osterwalder, Test Card, 5 March 2015Predeclared hypothesis, measure and criterion inform the protocol. The original card is not reproduced.
Nabila Amarsy/Strategyzer, Learning Card, 9 March 2015Supports recording the decision following a test. This is not validation of dotSuper's pilot template.
NIST, AI RMF 1.0, January 2023, and Generative AI Profile, July 2024Inform contextual evaluation and risk treatment. Neither provides a universal acceptable error rate.
NIST/SEMATECH, Confidence intervals for proportions, undated handbook pageSupports appropriate interval methods when failures or samples are small. Binomial assumptions must hold before using a binomial bound.
Brynjolfsson, Li and Raymond, Generative AI at Work, 2023/2025Demonstrates the importance of a defined work setting and differences across users. It supplies no expected effect size for this pilot.

Existing approaches and limitations

A demo establishes that selected examples can work. It does not establish typical performance or net operational benefit. Before-and-after comparisons can be confounded by workload, staffing, seasonality and learning. A single model accuracy figure can hide severe errors, uneven performance and the cost of human review. A pilot with enthusiastic volunteers may not represent the staff who will later use the system.

The proposed design separates technical quality, workflow effect, human use, cost and risk. It predeclares the main decision criterion and guardrails. It also permits an inconclusive outcome. A pilot that cannot distinguish improvement from noise should not automatically become a rollout because the project deadline arrived.

Structure and rationale

Create a decision protocol, an evaluation set register, a results ledger and a release record. The protocol specifies hypothesis, comparator, unit, population, measures, guardrails, scope and decision rules. The register records provenance, inclusion and exclusion, relevant subgroups and separation from development material.

The results ledger includes every eligible item, including abstentions, failures, timeouts and manual recovery. The release record distinguishes observed result, interpretation, unresolved risk and authorised next scope. A technically successful pilot can still stop if operating cost or ownership is unacceptable.

Inputs and preparation

Bring a baseline workflow, a permitted representative sample, clear quality definitions, a comparator and an evaluator with domain competence. Obtain the future operator's participation and the risk owner's agreement on containment. Record the model, configuration, prompt or rules version and data version.

Choose the evaluation unit carefully. One document may contain many fields, but those fields are correlated. Ten pages from one supplier are not ten independent suppliers. Do not inflate the sample by counting related observations as independent. If statistical inference is important, plan sample size and analysis with appropriate expertise using expected variation and a decision-relevant effect.

Facilitation and use

1. Write the decision question, 15 minutes. State which intervention, for whom, under what conditions and compared with what. Define what action follows each possible result. 2. Select the comparator, 15 minutes. Use current competent practice and, when relevant, a simpler process or rules-based alternative. Keep the quality standard equal across arms. If the comparison cannot be fair, document the limitation before running it. 3. Predeclare measures, 20 minutes. Choose one primary operational measure, such as total active review time per eligible item. Add critical quality, security and harm guardrails. Include setup and exception burden separately so a narrow time metric does not hide transferred work. 4. Prepare representative cases, 20 minutes. Include ordinary work and meaningful difficult cases. Tag language, document type, supplier, shift and other relevant variation. Keep tuning cases separate from the final evaluation and control access to the evaluation answers. 5. Choose a comparison design, 15 minutes. Random assignment may be feasible for low-risk work. A paired design can reduce item variation but requires counterbalancing order and attention to memory effects. If only a before-and-after study is feasible, record likely confounders and narrow causal claims. 6. Run within containment. Begin offline or in shadow mode where appropriate. Keep consequential actions under existing controls. Record all exclusions and changes. If a serious unexpected failure occurs, follow the stopping rule rather than continuing to improve the average. 7. Analyse without hiding failures, 20 minutes. Report denominators, distributions and subgroup results. Compare total review effort and recovery, not just generation time. Describe uncertainty and distinguish observed difference from causally attributable improvement. 8. Make the next decision, 15 minutes. Record stop, revise, extend evaluation or limited release, with scope and conditions. Check ownership and fallback readiness before any operational expansion.

Measures, thresholds and assumptions

Possible measures include critical-field retention, unsupported assertion rate, abstention rate, reviewer correction time and total workflow lead time. Define each numerator and denominator. “Accuracy” without a task and failure definition is insufficient. Reviewers should resolve disagreements against explicit criteria; an AI judge alone should not determine consequential domain correctness.

Set thresholds according to business consequence and baseline. A target of 20% lower review time is an illustrative decision threshold only if the team chooses it before seeing results. Do not use the same tolerated error rate for a draft summary and a safety-critical recommendation.

Zero observed failures does not prove zero underlying risk. Under independent, identically distributed Bernoulli trials, zero failures in 100 trials gives an exact one-sided 95% upper bound of approximately 2.95%, from 1 minus 0.05 raised to the power 1/100. Correlated cases or distribution change undermine that interpretation. Severe harms may require specialised assurance beyond any small pilot.

Worked example: illustrative, not a dotSuper result

A hypothetical procurement team tests assisted comparison on 100 eligible historical quotation packages. Twenty additional packages form a development set and never enter the final reported evaluation. The team predefines missing payment or delivery terms as critical errors and evaluates source traceability.

Competent manual review is the comparator. Reviewers use a counterbalanced allocation so one person does not simply remember their previous answer. Total active time includes reading, checking and correction. The primary target is a locally chosen reduction of at least 20%, with no increase in critical errors and no unauthorised data exposure. These are example design choices, not recommended universal thresholds.

Suppose measured median active review time falls from 12 to 8 minutes, but three assisted outputs omit a critical delivery qualification. The team cannot average those failures away with the time benefit. It investigates whether source extraction, presentation or reviewer behaviour caused the misses. The current version does not meet the quality guardrail.

The decision is revise, not deploy. A change that highlights uncertain clauses is tested on a fresh evaluation set, while the original failures remain in the issue history. If the team instead sees no failures, it still reports the limited sample and tests live exception handling before expanding. A successful offline evaluation does not establish adoption, long-run economics or production safety.

Outputs, failure modes and validation limits

Produce a dated protocol, evaluation-set manifest, results with denominators, incident and exclusion log, and a signed next-decision record. Preserve negative results. A useful pilot may establish that a conventional template performs as well at lower cost.

Failures include changing success criteria after seeing results, testing on training examples, omitting abstentions, excluding difficult cases without disclosure and reporting only the average. Another is claiming cash savings from faster work without a value-capture plan. A counterexample is a rare-event workflow where a short pilot cannot observe the relevant failure often enough. Longer evaluation, simulation and domain assurance may be required.

Validate this protocol through independent review of the design and reproducibility of the analysis. The dotSuper template itself has not been evaluated as an intervention. Its purpose is to make scope and evidence inspectable, not confer statistical or safety authority.

Working files and reuse

The five-page PDF includes a visual reference, two fillable worksheet pages, an illustrative worked example and a facilitator/source guide. An expandable CSV working log is also available.

All examples are illustrative, not measured client results. Original pilot and release record informed by established experimentation. Not field validated. A small pilot cannot establish rare-event safety, long-run economics or sustained adoption. Zero observed failures does not mean zero risk. No Test Card layout is reproduced.

Original dotSuper material prepared for review. No public reuse licence has been assigned. Third-party source material retains its own terms.

The PDF is not represented as a tagged PDF/UA document. The text on this page provides a readable alternative to the diagram and method.

See the method. Keep the context.

The visual companion

Pilot Design and Evaluation. Before the run, declare the question, population, unit, thresholds and guardrails. Compare competent current practice with a fixed version and scope of the intervention. Measure quality, effort, recovery and cost. Retain failures and abstentions and report denominators and uncertainty. Decide to stop, revise, extend evaluation or limit release. Faster performance cannot compensate for an unmet consequential guardrail.
Pilot Design and Evaluation. Original dotSuper reference diagram. Hypothesis testing, comparators and statistical uncertainty are established. Release scope, exception accounting and operating-readiness records are dotSuper synthesis. Open full size

Credit: dotSuper. Hypothesis testing, comparators and statistical uncertainty are established. Release scope, exception accounting and operating-readiness records are dotSuper synthesis.

Reuse: Original dotSuper material. No public reuse licence has been specified. Contact dotSuper for reuse permissions. Third-party source material retains its own terms.

Read the diagram: Pilot Design and Evaluation. Original dotSuper reference diagram. Hypothesis testing, comparators and statistical uncertainty are established. Release scope, exception accounting and operating-readiness records are dotSuper synthesis.

Before the run, declare the question, population, unit, thresholds and guardrails. Compare competent current practice with a fixed version and scope of the intervention. Measure quality, effort, recovery and cost. Retain failures and abstentions and report denominators and uncertainty. Decide to stop, revise, extend evaluation or limit release. Faster performance cannot compensate for an unmet consequential guardrail.

Original pilot and release record informed by established experimentation. Not field validated. A small pilot cannot establish rare-event safety, long-run economics or sustained adoption. Zero observed failures does not mean zero risk. No Test Card layout is reproduced.

Illustrative pilot. 100 eligible evaluation packages and 20 separate development packages. Values are not a dotSuper result.
MeasureIllustrative valueScope
Manual median active review time12 minutesReading, checking and correction
Assisted median active review time8 minutesSame operational measure
Assisted outputs with critical omissions3 of 100Critical delivery qualification missing
Predeclared time targetAt least 20% reductionLocal example threshold, not a universal target
Next decisionReviseTime benefit does not resolve the quality issue

Take it into your next working session

Keep the source credits with the file. Check the reuse terms and adapt the method to your context.

Download the five-page fillable frameworkPDF · 83 KB

Credit: dotSuper. Hypothesis testing, comparators and statistical uncertainty are established. Release scope, exception accounting and operating-readiness records are dotSuper synthesis.

Reuse: Original dotSuper material. No public reuse licence has been specified. Contact dotSuper for reuse permissions. Third-party source material retains its own terms.

Download the expandable working logCSV · 1 KB

Credit: dotSuper. Hypothesis testing, comparators and statistical uncertainty are established. Release scope, exception accounting and operating-readiness records are dotSuper synthesis.

Reuse: Original dotSuper material. No public reuse licence has been specified. Contact dotSuper for reuse permissions. Third-party source material retains its own terms.

Thumbnail credit and reuse

Credit: dotSuper. Hypothesis testing, comparators and statistical uncertainty are established. Release scope, exception accounting and operating-readiness records are dotSuper synthesis.

Reuse: Original dotSuper material. No public reuse licence has been specified. Contact dotSuper for reuse permissions. Third-party source material retains its own terms.

Sources, context and limits

Keep the evidence beside the method.

  • Original pilot and release record informed by established experimentation. Not field validated. A small pilot cannot establish rare-event safety, long-run economics or sustained adoption. Zero observed failures does not mean zero risk. No Test Card layout is reproduced.
  • This combined dotSuper method is research-informed and has not been field validated. Workshop agreement and a completed template are not proof of effectiveness.
  • Worked examples are hypothetical. Country examples and intended regional audience do not establish country-wide or region-wide effectiveness.
  • Source access and adaptation limits are recorded in the source ledger. Attribution does not imply endorsement or a licence to reproduce third-party artwork.
Download the printable framework

/ CITE OR SHARE THIS GUIDE

Make the evidence easy to verify.

When you reference this guide, link to its canonical URL. That gives readers one stable place for the evidence, limitations and future updates.

Suggested citation

dotSuper Research Desk. (September 17, 2026). Design an AI Pilot Around a Testable Decision. dotSuper. https://dotsuper.net/feeds/applied-systems/pilot-design-and-evaluation

Share on LinkedIn

Bring it into the work

Start with one real decision

Bring one pilot idea and define the comparison, evidence and stopping rules.

Download the fillable framework