98% field accuracy. Is the document right?

Evaluate whole documents, critical fields and language groups separately. A synthetic test set shows how 98% field accuracy can coexist with only 90% fully correct documents.

By dotSuper Research DeskPublished Sep 17, 2026Updated Sep 17, 20266 min read
Applied systemsCited source evidence, dotSuper operating analysis and explicitly fictional worked examplesUpdated Sep 17, 2026

/ THE SHORT ANSWER

See the method. Keep the context.

The visual companion

In 200 fictional documents with five fields each, 20 field errors occur in 20 separate documents. Field accuracy is 98%, but only 180 documents, or 90%, are entirely correct. Complete-document rates are 118/120 English, 48/60 Arabic and 14/20 mixed-language documents.
Measure whole-document success and the fields that can change the decision. All evaluation figures are fictional. Each affected document has one field error. Language-group rates are an explanation of denominators, not measured model performance. Open full size

Credit: Created for dotSuper. Original layout, charts and AI-assisted editorial illustration; supplied dotSuper brand artwork.

Reuse: Original artwork created for dotSuper. No public reuse licence has been specified. Attribution to research sources does not grant rights to their artwork or datasets.

Read the diagram: Measure whole-document success and the fields that can change the decision. All evaluation figures are fictional. Each affected document has one field error. Language-group rates are an explanation of denominators, not measured model performance.

In 200 fictional documents with five fields each, 20 field errors occur in 20 separate documents. Field accuracy is 98%, but only 180 documents, or 90%, are entirely correct. Complete-document rates are 118/120 English, 48/60 Arabic and 14/20 mixed-language documents.

There are 200 documents with five evaluated fields each: 1,000 field decisions, 980 correct fields and 20 errors. Those errors affect 20 separate documents, leaving 180 entirely correct.

English: 118 of 120 complete documents, 98.3%. Arabic: 48 of 60, 80.0%. Mixed: 14 of 20, 70.0%. These constructed subgroup rates illustrate denominators, not a real language gap.

98% field accuracy. Is the document right? Fictional example: one wrong field can make the whole document unusable. Define acceptance before selecting a model.

Fields and documents measure different things. 98% / 980 correct fields / 1,000 field decisions 90% / 180 fully correct documents / 200 documents Twenty errors are spread across 20 documents, one error per affected document. All figures are fictional.

Look inside the overall score. Illustrative complete-document results English / 118 / 120 / 98.3% Arabic / 48 / 60 / 80.0% Mixed / 14 / 20 / 70.0% Report each denominator and language group. This table does not describe a real model or prove a language gap.

All 200-document results are synthetic, with one field error in each of 20 affected documents. No model ranking or production accuracy is asserted.

Research benchmarks use their own tasks and corpora; they are not Gulf invoice-compliance tests.

The binomial zero-error illustration assumes independent, comparable trials. Correlated or unrepresentative documents can make it optimistic.

Readable research papers do not grant commercial rights to their source images or datasets.

All figure values; official-source dates and fictional calculations retain their separate labels.
DatasetObservationValues and conditions
Fictional example: document language cohortsEnglishDocuments: 120; Completely correct: 118; Complete-document accuracy: 98.3%
Fictional example: document language cohortsArabicDocuments: 60; Completely correct: 48; Complete-document accuracy: 80%
Fictional example: document language cohortsMixedDocuments: 20; Completely correct: 14; Complete-document accuracy: 70%
Fictional example: accuracy denominatorsFieldsCorrect units: 980; All units: 1,000; Rate: 98%
Fictional example: accuracy denominatorsWhole documentsCorrect units: 180; All units: 200; Rate: 90%
98% field accuracy. Is the document right? Fictional example: one wrong field can make the whole document unusable. Define acceptance before selecting a model.
Panel 1 of 3. 98% field accuracy. Open full size

Credit: Created for dotSuper. Original layout, charts and AI-assisted editorial illustration; supplied dotSuper brand artwork.

Reuse: Original artwork created for dotSuper. No public reuse licence has been specified. Attribution to research sources does not grant rights to their artwork or datasets.

Read the diagram: Panel 1 of 3. 98% field accuracy.

98% field accuracy. Is the document right? Fictional example: one wrong field can make the whole document unusable. Define acceptance before selecting a model.

All 200-document results are synthetic, with one field error in each of 20 affected documents. No model ranking or production accuracy is asserted.

Fictional example: accuracy denominators.

Measure: Fields; Correct units: 980; All units: 1,000; Rate: 98%.

Measure: Whole documents; Correct units: 180; All units: 200; Rate: 90%.

Fictional example: accuracy denominators
MeasureCorrect unitsAll unitsRate
Fields9801,00098%
Whole documents18020090%
Fields and documents measure different things. 98% / 980 correct fields / 1,000 field decisions 90% / 180 fully correct documents / 200 documents Twenty errors are spread across 20 documents, one error per affected document. All figures are fictional.
Panel 2 of 3. Fields and documents measure different things. Open full size

Credit: Created for dotSuper. Original layout, charts and AI-assisted editorial illustration; supplied dotSuper brand artwork.

Reuse: Original artwork created for dotSuper. No public reuse licence has been specified. Attribution to research sources does not grant rights to their artwork or datasets.

Read the diagram: Panel 2 of 3. Fields and documents measure different things.

Fields and documents measure different things. 98% / 980 correct fields / 1,000 field decisions 90% / 180 fully correct documents / 200 documents Twenty errors are spread across 20 documents, one error per affected document. All figures are fictional.

All 200-document results are synthetic, with one field error in each of 20 affected documents. No model ranking or production accuracy is asserted.

Fictional example: accuracy denominators.

Measure: Fields; Correct units: 980; All units: 1,000; Rate: 98%.

Measure: Whole documents; Correct units: 180; All units: 200; Rate: 90%.

Fictional example: accuracy denominators
MeasureCorrect unitsAll unitsRate
Fields9801,00098%
Whole documents18020090%
Look inside the overall score. Illustrative complete-document results English / 118 / 120 / 98.3% Arabic / 48 / 60 / 80.0% Mixed / 14 / 20 / 70.0% Report each denominator and language group. This table does not describe a real model or prove a language gap.
Panel 3 of 3. Look inside the overall score. Open full size

Credit: Created for dotSuper. Original layout, charts and AI-assisted editorial illustration; supplied dotSuper brand artwork.

Reuse: Original artwork created for dotSuper. No public reuse licence has been specified. Attribution to research sources does not grant rights to their artwork or datasets.

Read the diagram: Panel 3 of 3. Look inside the overall score.

Look inside the overall score. Illustrative complete-document results English / 118 / 120 / 98.3% Arabic / 48 / 60 / 80.0% Mixed / 14 / 20 / 70.0% Report each denominator and language group. This table does not describe a real model or prove a language gap.

All 200-document results are synthetic, with one field error in each of 20 affected documents. No model ranking or production accuracy is asserted.

Fictional example: document language cohorts.

Document language: English; Documents: 120; Completely correct: 118; Complete-document accuracy: 98.3%.

Document language: Arabic; Documents: 60; Completely correct: 48; Complete-document accuracy: 80%.

Document language: Mixed; Documents: 20; Completely correct: 14; Complete-document accuracy: 70%.

Fictional example: document language cohorts
Document languageDocumentsCompletely correctComplete-document accuracy
English12011898.3%
Arabic604880%
Mixed201470%

Take it into your next working session

Keep the source credits with the file. Check the reuse terms and adapt the method to your context.

Fictional example data: accuracy denominators (CSV)CSV · 1 KB

Credit: dotSuper original fictional worked-example data.

Reuse: Original illustrative data. No public reuse licence has been specified. These are not client measurements or research results.

Fictional example data: document language cohorts (CSV)CSV · 1 KB

Credit: dotSuper original fictional worked-example data.

Reuse: Original illustrative data. No public reuse licence has been specified. These are not client measurements or research results.

Mobile panel 1 (PNG)PNG · 470 KB

Credit: Created for dotSuper. Original layout, charts and AI-assisted editorial illustration; supplied dotSuper brand artwork.

Reuse: Original artwork created for dotSuper. No public reuse licence has been specified. Attribution to research sources does not grant rights to their artwork or datasets.

Mobile panel 2 (PNG)PNG · 104 KB

Credit: Created for dotSuper. Original layout, charts and AI-assisted editorial illustration; supplied dotSuper brand artwork.

Reuse: Original artwork created for dotSuper. No public reuse licence has been specified. Attribution to research sources does not grant rights to their artwork or datasets.

Mobile panel 3 (PNG)PNG · 106 KB

Credit: Created for dotSuper. Original layout, charts and AI-assisted editorial illustration; supplied dotSuper brand artwork.

Reuse: Original artwork created for dotSuper. No public reuse licence has been specified. Attribution to research sources does not grant rights to their artwork or datasets.

Full infographic (PDF)PDF · 1.7 MB

Credit: Created for dotSuper. Original layout, charts and AI-assisted editorial illustration; supplied dotSuper brand artwork.

Reuse: Original artwork created for dotSuper. No public reuse licence has been specified. Attribution to research sources does not grant rights to their artwork or datasets.

Full infographic (PNG)PNG · 818 KB

Credit: Created for dotSuper. Original layout, charts and AI-assisted editorial illustration; supplied dotSuper brand artwork.

Reuse: Original artwork created for dotSuper. No public reuse licence has been specified. Attribution to research sources does not grant rights to their artwork or datasets.

Full infographic (WEBP)WEBP · 231 KB

Credit: Created for dotSuper. Original layout, charts and AI-assisted editorial illustration; supplied dotSuper brand artwork.

Reuse: Original artwork created for dotSuper. No public reuse licence has been specified. Attribution to research sources does not grant rights to their artwork or datasets.

Thumbnail credit and reuse

Credit: Created for dotSuper. Original layout, charts and AI-assisted editorial illustration; supplied dotSuper brand artwork.

Reuse: Original artwork created for dotSuper. No public reuse licence has been specified. Attribution to research sources does not grant rights to their artwork or datasets.

Key takeaways
  • 01In the synthetic example, 98% field accuracy coexists with 90% complete-document accuracy.
  • 02Language-group rates describe a constructed test set, not a real model or a measured Arabic-English performance gap.
  • 03Preserve document structure, row associations and legitimate missing values when scoring outputs.
  • 04Use representative held-out documents and independently checked reference answers before making a workflow decision.

/ dotSuper point of view

A high field score can conceal an unusable document. Acceptance should reflect critical fields, complete documents, subgroup denominators and downstream consequences.
01Orient

Audience, question and answer

Evaluate the business task at field, document and workflow levels, using representative documents and independently checked answers.

Separate language, layout and image quality.

A model that reads a paragraph well can still attach a quantity to the wrong line item or confuse an invoice date with a due date.

The decision is whether a defined workflow can be assisted safely at an acceptable total effort.

It is not which model leads a general leaderboard.

This explanation applies to bilingual operating documents in several markets; it does not claim one Arabic benchmark represents Saudi and UAE paperwork equally.

02Signal

Research findings

Its breadth demonstrates why “Arabic accuracy” needs task definition.

A benchmark score is conditional on its dataset and evaluated model versions.

Heakl et al., KITAB-Bench, Findings of ACL 2025, pages 22006–22024, checked 16 September 2026 (source 2).

DocVQA introduced approximately 50,000 questions over more than 12,000 document images and emphasised questions requiring document structure.

It is useful precedent for evaluating answers in document context, but it is not a country-specific invoice compliance test.

Mathew, Karatzas and Jawahar, DocVQA, July 2020 (source 3).

Arabic document work also involves cursive text and right-to-left reading within potentially mixed-direction layouts.

The KITAB paper treats these as evaluation challenges.

Do not turn its aggregate comparisons into a vendor performance promise for photographed purchase orders.

KITAB-Bench paper, 2025 (source 1).

For experimental discipline, the NIST AI RMF highlights context, measurement and risk management rather than a single universal score.

This explanation's proposed test protocol is dotSuper's operational interpretation, not NIST certification.

NIST AI RMF 1.0, January 2023, checked 16 September 2026 (source 4).

03Prove

Define the actual unit of success

For a quality report, the critical fields may be drawing revision, measured dimension, tolerance and pass/fail evidence.

The business owner defines the consequences of each error before the test starts.

Measure character error rate only for transcription tasks.

It counts character substitutions, deletions and insertions relative to reference characters.

It does not measure whether an extracted bank account belongs to the correct supplier.

A semantically plausible paraphrase may be acceptable for a summary but unacceptable for an identifier.

Use exact comparison for identifiers where appropriate.

For names and addresses, document any normalisation rules in advance.

Preserve the raw text alongside the normalised value.

Do not silently erase leading zeros, punctuation or script differences merely to raise a score.

For numeric amounts, distinguish the value from its currency and sign.

For tables, score row association as well as field content.

Ten correct quantities attached to the wrong products are ten wrong business records.

Include subtotal, discount and credit-note cases.

Require the system to return missing or uncertain when the document does not support an answer.

04Resolve

Build a test set that can fail usefully

Stratify by Arabic, English and mixed content; native PDFs and scans; clear and degraded images; known and new supplier templates; single and multiple pages.

Document whose production mix the test represents and which cases are deliberately overrepresented for safety testing.

Keep supplier templates and near-duplicate documents grouped when splitting development and test data.

A random page split can leak the same form into both sets.

Freeze the final test set and model configuration before reporting results.

Record extraction schema, prompt, model version, preprocessing and any external lookup used.

Have a second reviewer check the reference answers for critical fields.

Resolve disagreement explicitly.

An ambiguous source should not be counted as a clear model error or a clear success.

Keep an adjudication log with the document locator and reason.

Include adversarial and operational cases: a page that instructs the system to change its rules, an illegible amount, a missing page, a duplicate invoice, conflicting totals and a new stamp obscuring text.

These are test inputs.

They do not authorise changes to the system's operating instructions.

05Orient

Worked example: 98% can conceal a weak document result

Assume 200 documents with five evaluated fields each, giving 1,000 field decisions.

Twenty field errors are spread across 20 different documents.

Field accuracy is 980/1,000, or 98%.

Fully correct documents are 180/200, or 90%.

The table is constructed so that the 2, 12 and 6 incomplete documents account for the 20 total errors, one per affected document.

It does not report a real model's language gap.

The point is that a weighted overall score can hide a weak subgroup and a small denominator.

Suppose each document saves three minutes of entry but requires one minute of routine review.

Gross released effort is 200 × 2 = 400 minutes before exception correction, test maintenance and integration.

If exceptions add five minutes on each of the 20 affected documents, net released effort becomes 300 minutes, or five hours.

This remains a capacity estimate, not a realised cash saving.

Worked example: 98% can conceal a weak document result
Illustrative document groupDocumentsFully correctComplete-document rate
English12011898.3%
Arabic604880.0%
Mixed Arabic-English201470.0%
Total20018090.0%
06Signal

Quantitative interpretation and acceptance

Twenty mixed documents are too few to make a precise general claim about all mixed-language invoices.

Expand the sample where the business risk or uncertainty matters, rather than evenly adding easy documents to improve the overall average.

For an independent Bernoulli-error approximation, observing zero errors in 100 cases still allows an upper one-sided 95% error bound of about 3%.

This follows from solving (1-p)^100 = 0.05.

Real documents may be correlated, making that simple bound optimistic.

Zero observed errors is not a guarantee of zero production errors.

Acceptance should combine field correctness, exception detection, reviewer workload and downstream error severity.

Define unacceptable errors and the required escalation in advance.

A system may be acceptable for draft indexing but unsuitable for unattended accounting entries using the same underlying model.

07Prove

Evidence and boundaries

Each source supports only the scope stated beside the claim.

Evidence and boundaries
ClaimEvidenceCaveat
Arabic evaluation requires varied tasks and document typesKITAB-Bench 2025Research dataset, not Gulf production sample
Document structure matters to answeringDocVQA 2020Different task and corpus from current invoices
Measurement should reflect use contextNIST AI RMFVoluntary framework, not local legal compliance
98% fields can mean 90% complete documentsFictional 1,000-field exampleError placement is explicitly assumed
Language subgroup rates differ in the exampleConstructed 120/60/20 splitNo model or market comparison
Zero of 100 is not proof of zero errorBinomial calculation aboveIndependence assumption may fail
08Resolve

One practical next step

Record both complete-document results and reviewer effort.

09Orient

Counterevidence and limits

Conversely, unfamiliar layouts can benefit from broader document understanding.

Evaluate both against the same business task rather than assume novelty determines suitability.

Human review also makes errors and can become perfunctory.

Measure whether reviewers actually catch seeded mistakes and whether the interface shows sufficient source context.

A workflow that routes everything to a human may be safe in one sense but economically unhelpful or operationally overloaded.

Benchmark versions age.

No current model ranking or Arabic-English production accuracy is asserted here.

The dataset's licence and any underlying document restrictions must be checked separately before commercial testing.

A research paper being freely readable does not grant rights to its source images.

What this page cannot conclude

  • 01All 200-document results are synthetic, with one field error in each of 20 affected documents. No model ranking or production accuracy is asserted.
  • 02Research benchmarks use their own tasks and corpora; they are not Gulf invoice-compliance tests.
  • 03The binomial zero-error illustration assumes independent, comparable trials. Correlated or unrepresentative documents can make it optimistic.
  • 04Readable research papers do not grant commercial rights to their source images or datasets.

Sources

  1. 01KITAB-Bench paper, 2025ACL Anthology; Heakl and co-authors · accessed Sep 16, 2026
  2. 02Heakl et al., KITAB-Bench, Findings of ACL 2025, pages 22006–22024, checked 16 September 2026ACL Anthology; Heakl and co-authors · accessed Sep 16, 2026
  3. 03Mathew, Karatzas and Jawahar, DocVQA, July 2020Mathew, Karatzas and Jawahar · accessed Sep 16, 2026
  4. 04NIST AI RMF 1.0, January 2023, checked 16 September 2026US National Institute of Standards and Technology · accessed Sep 16, 2026

Our editorial standard · Found an error? Send a correction with its source.

/ CITE OR SHARE THIS GUIDE

Make the evidence easy to verify.

When you reference this guide, link to its canonical URL. That gives readers one stable place for the evidence, limitations and future updates.

Suggested citation

dotSuper Research Desk. (September 17, 2026). 98% field accuracy. Is the document right?. dotSuper. https://dotsuper.net/feeds/applied-systems/evaluate-arabic-english-document-ai

Share on LinkedIn
ONE OPERATING QUESTION98% field accuracy. Is the document right?

/ APPLY THE THINKING

Bring this decision to a research conversation.

Define critical fields and unacceptable errors, then build a held-out sample covering language, layout, scan quality and new supplier templates. Record both complete-document results and reviewer effort.

Question for the working sessionWhat does document accuracy need to mean for this workflow?

/ Topic-led working session · 98% field accuracy. Is the document right?

Turn this question\ninto a useful first move.

Bring how this question currently shows up in your business: “What does document accuracy need to mean for this workflow?” We’ll test the page’s evidence against your context and define the smallest useful next move.

Live availability from ceo@dotsuper.net Automatically converted · your local time
  1. 01Bring the contextWhere this issue shows up in the work.
  2. 02Test the relevanceUse the evidence against your reality.
  3. 03Choose the next moveOne accountable action, clearly owned.
Live availability
  1. Date
  2. Time
  3. Booked

Syncing live times