/ THE SHORT ANSWER
Use a proof of concept to answer a narrow technical uncertainty: can the method work on representative inputs under controlled conditions? Use a production pilot to answer the operating question: can real users rely on the bounded system inside the actual workflow, with controls, monitoring, integration, and ownership? If feasibility is already established by available technology, repeating a lab demo wastes time; test adoption, reliability, economics, and exception handling instead.
- 01Name the uncertainty each stage must retire.
- 02Do not use production users to discover basic technical feasibility.
- 03Do not treat controlled accuracy as proof of operational value.
/ dotSuper point of view
A proof of concept earns the right to design a pilot. A pilot earns the right to change the operation. Neither should be called success merely because the interface produced an impressive output.
What the evidence says
NIST’s AI RMF connects measurement to the context in which an AI system is deployed and used. Performance evidence without the relevant people, process, and risk context is incomplete.
The NIST Generative AI Profile recommends evaluation across the lifecycle and attention to confabulation, information integrity, security, human-AI configuration, and other system-level risks.
A practical decision framework
The following framework is dotSuper’s operating synthesis of the cited guidance. It is designed to make the decision inspectable, not to imitate a platform ranking formula, certification checklist, or legal test.
- PoC gate: one technical hypothesis, representative sample, baseline, and stop rule.
- Pilot gate: named users, real workflow boundary, fallback path, and accountable operator.
- Scale gate: stable value signal, acceptable risk, support model, and monitored integration.
- Stop gate: evidence shows insufficient value, unreliable inputs, or disproportionate controls.
| Step | Decision to record |
|---|---|
| 01 | PoC gate: one technical hypothesis, representative sample, baseline, and stop rule. |
| 02 | Pilot gate: named users, real workflow boundary, fallback path, and accountable operator. |
| 03 | Scale gate: stable value signal, acceptable risk, support model, and monitored integration. |
| 04 | Stop gate: evidence shows insufficient value, unreliable inputs, or disproportionate controls. |
How to put it into practice
Write the decision that will be made at the end of the stage before any build begins. Define evidence strong enough to say continue, redesign, or stop.
Keep the pilot bounded: one workflow, one user group, one source-of-truth set, and one measurable outcome. Expand only after the review shows where the system is dependable and where it is not.
- Name the accountable owner and the decision this work must enable.
- Record the current evidence, assumptions, exclusions, and next review trigger.
- Measure a useful outcome rather than treating publication or deployment as success.
What this page cannot conclude
- 01There is no universal accuracy threshold; acceptable performance depends on the decision and consequence.
- 02A successful pilot in one site or team may not generalise without additional testing.
- 03Publication, technical eligibility, or good practice cannot guarantee ranking, referral traffic, citation, adoption, or a business outcome.
Sources
Test the workflow before funding the solution.
The AI Readiness Sprint turns one operational constraint into a ranked decision, an accountable owner, and an implementation-ready first move.
Explore the readiness sprint