/ THE SHORT ANSWER
Evaluate the knowledge system, not just the chat response. Require clarity on source ingestion, document versions, permissions, retrieval evaluation, answer grounding, citation behaviour, refusal, monitoring, cost, and administration. Test with real questions that include missing, conflicting, outdated, and restricted information. The vendor should show how the system fails safely and how your team updates or removes knowledge without specialist intervention.
- 01Test retrieval separately from answer fluency.
- 02Include conflicting, stale, absent, and permissioned cases.
- 03Price the ongoing content and evaluation operation, not only the software.
/ dotSuper point of view
RAG quality is an operating property of sources, retrieval, permissions, evaluation, and ownership. A fluent demo can hide weakness in every one of those layers.
What the evidence says
The NIST Generative AI Profile identifies information integrity, confabulation, privacy, security, and human-AI configuration as risks requiring lifecycle treatment.
OWASP’s LLM application guidance highlights risks including prompt injection, sensitive-information disclosure, vector and embedding weaknesses, excessive agency, and unbounded consumption.
A practical decision framework
The following framework is dotSuper’s operating synthesis of the cited guidance. It is designed to make the decision inspectable, not to imitate a platform ranking formula, certification checklist, or legal test.
- Sources: provenance, versioning, deletion, freshness, and ownership.
- Retrieval: relevance, coverage, conflict handling, and permission enforcement.
- Generation: grounded answers, citations, refusal, and uncertainty.
- Operations: evaluation set, monitoring, change control, cost, and support.
| Step | Decision to record |
|---|---|
| 01 | Sources: provenance, versioning, deletion, freshness, and ownership. |
| 02 | Retrieval: relevance, coverage, conflict handling, and permission enforcement. |
| 03 | Generation: grounded answers, citations, refusal, and uncertainty. |
| 04 | Operations: evaluation set, monitoring, change control, cost, and support. |
How to put it into practice
Create a 30–50 question acceptance set from actual users. Include straightforward, multi-source, ambiguous, unanswerable, stale, and restricted questions, then score retrieval and final answers separately.
Ask the vendor to demonstrate an update, a permission change, a source deletion, a model change, and a rollback. These lifecycle actions determine whether the system can be safely owned.
- Name the accountable owner and the decision this work must enable.
- Record the current evidence, assumptions, exclusions, and next review trigger.
- Measure a useful outcome rather than treating publication or deployment as success.
What this page cannot conclude
- 01A short acceptance set cannot predict every production interaction.
- 02Sector-specific privacy, records, security, and residency requirements need separate review.
- 03Publication, technical eligibility, or good practice cannot guarantee ranking, referral traffic, citation, adoption, or a business outcome.
Sources
Test the workflow before funding the solution.
The AI Readiness Sprint turns one operational constraint into a ranked decision, an accountable owner, and an implementation-ready first move.
Explore the readiness sprint