/ THE SHORT ANSWER
See the method. Keep the context.
The visual companion

Reuse: Original artwork created for dotSuper. No public reuse licence has been specified. Attribution to research sources does not grant rights to their artwork or datasets.
Read the diagram: Measure whole-document success and the fields that can change the decision. All evaluation figures are fictional. Each affected document has one field error. Language-group rates are an explanation of denominators, not measured model performance.
In 200 fictional documents with five fields each, 20 field errors occur in 20 separate documents. Field accuracy is 98%, but only 180 documents, or 90%, are entirely correct. Complete-document rates are 118/120 English, 48/60 Arabic and 14/20 mixed-language documents.
There are 200 documents with five evaluated fields each: 1,000 field decisions, 980 correct fields and 20 errors. Those errors affect 20 separate documents, leaving 180 entirely correct.
English: 118 of 120 complete documents, 98.3%. Arabic: 48 of 60, 80.0%. Mixed: 14 of 20, 70.0%. These constructed subgroup rates illustrate denominators, not a real language gap.
98% field accuracy. Is the document right? Fictional example: one wrong field can make the whole document unusable. Define acceptance before selecting a model.
Fields and documents measure different things. 98% / 980 correct fields / 1,000 field decisions 90% / 180 fully correct documents / 200 documents Twenty errors are spread across 20 documents, one error per affected document. All figures are fictional.
Look inside the overall score. Illustrative complete-document results English / 118 / 120 / 98.3% Arabic / 48 / 60 / 80.0% Mixed / 14 / 20 / 70.0% Report each denominator and language group. This table does not describe a real model or prove a language gap.
All 200-document results are synthetic, with one field error in each of 20 affected documents. No model ranking or production accuracy is asserted.
Research benchmarks use their own tasks and corpora; they are not Gulf invoice-compliance tests.
The binomial zero-error illustration assumes independent, comparable trials. Correlated or unrepresentative documents can make it optimistic.
Readable research papers do not grant commercial rights to their source images or datasets.
| Dataset | Observation | Values and conditions |
|---|---|---|
| Fictional example: document language cohorts | English | Documents: 120; Completely correct: 118; Complete-document accuracy: 98.3% |
| Fictional example: document language cohorts | Arabic | Documents: 60; Completely correct: 48; Complete-document accuracy: 80% |
| Fictional example: document language cohorts | Mixed | Documents: 20; Completely correct: 14; Complete-document accuracy: 70% |
| Fictional example: accuracy denominators | Fields | Correct units: 980; All units: 1,000; Rate: 98% |
| Fictional example: accuracy denominators | Whole documents | Correct units: 180; All units: 200; Rate: 90% |

Reuse: Original artwork created for dotSuper. No public reuse licence has been specified. Attribution to research sources does not grant rights to their artwork or datasets.
Read the diagram: Panel 1 of 3. 98% field accuracy.
98% field accuracy. Is the document right? Fictional example: one wrong field can make the whole document unusable. Define acceptance before selecting a model.
All 200-document results are synthetic, with one field error in each of 20 affected documents. No model ranking or production accuracy is asserted.
Fictional example: accuracy denominators.
Measure: Fields; Correct units: 980; All units: 1,000; Rate: 98%.
Measure: Whole documents; Correct units: 180; All units: 200; Rate: 90%.
| Measure | Correct units | All units | Rate |
|---|---|---|---|
| Fields | 980 | 1,000 | 98% |
| Whole documents | 180 | 200 | 90% |

Reuse: Original artwork created for dotSuper. No public reuse licence has been specified. Attribution to research sources does not grant rights to their artwork or datasets.
Read the diagram: Panel 2 of 3. Fields and documents measure different things.
Fields and documents measure different things. 98% / 980 correct fields / 1,000 field decisions 90% / 180 fully correct documents / 200 documents Twenty errors are spread across 20 documents, one error per affected document. All figures are fictional.
All 200-document results are synthetic, with one field error in each of 20 affected documents. No model ranking or production accuracy is asserted.
Fictional example: accuracy denominators.
Measure: Fields; Correct units: 980; All units: 1,000; Rate: 98%.
Measure: Whole documents; Correct units: 180; All units: 200; Rate: 90%.
| Measure | Correct units | All units | Rate |
|---|---|---|---|
| Fields | 980 | 1,000 | 98% |
| Whole documents | 180 | 200 | 90% |

Reuse: Original artwork created for dotSuper. No public reuse licence has been specified. Attribution to research sources does not grant rights to their artwork or datasets.
Read the diagram: Panel 3 of 3. Look inside the overall score.
Look inside the overall score. Illustrative complete-document results English / 118 / 120 / 98.3% Arabic / 48 / 60 / 80.0% Mixed / 14 / 20 / 70.0% Report each denominator and language group. This table does not describe a real model or prove a language gap.
All 200-document results are synthetic, with one field error in each of 20 affected documents. No model ranking or production accuracy is asserted.
Fictional example: document language cohorts.
Document language: English; Documents: 120; Completely correct: 118; Complete-document accuracy: 98.3%.
Document language: Arabic; Documents: 60; Completely correct: 48; Complete-document accuracy: 80%.
Document language: Mixed; Documents: 20; Completely correct: 14; Complete-document accuracy: 70%.
| Document language | Documents | Completely correct | Complete-document accuracy |
|---|---|---|---|
| English | 120 | 118 | 98.3% |
| Arabic | 60 | 48 | 80% |
| Mixed | 20 | 14 | 70% |
Take it into your next working session
Keep the source credits with the file. Check the reuse terms and adapt the method to your context.
Credit: dotSuper original fictional worked-example data.
Reuse: Original illustrative data. No public reuse licence has been specified. These are not client measurements or research results.
Credit: dotSuper original fictional worked-example data.
Reuse: Original illustrative data. No public reuse licence has been specified. These are not client measurements or research results.
Reuse: Original artwork created for dotSuper. No public reuse licence has been specified. Attribution to research sources does not grant rights to their artwork or datasets.
Reuse: Original artwork created for dotSuper. No public reuse licence has been specified. Attribution to research sources does not grant rights to their artwork or datasets.
Reuse: Original artwork created for dotSuper. No public reuse licence has been specified. Attribution to research sources does not grant rights to their artwork or datasets.
Reuse: Original artwork created for dotSuper. No public reuse licence has been specified. Attribution to research sources does not grant rights to their artwork or datasets.
Reuse: Original artwork created for dotSuper. No public reuse licence has been specified. Attribution to research sources does not grant rights to their artwork or datasets.
Reuse: Original artwork created for dotSuper. No public reuse licence has been specified. Attribution to research sources does not grant rights to their artwork or datasets.
Thumbnail credit and reuse
Reuse: Original artwork created for dotSuper. No public reuse licence has been specified. Attribution to research sources does not grant rights to their artwork or datasets.
- 01In the synthetic example, 98% field accuracy coexists with 90% complete-document accuracy.
- 02Language-group rates describe a constructed test set, not a real model or a measured Arabic-English performance gap.
- 03Preserve document structure, row associations and legitimate missing values when scoring outputs.
- 04Use representative held-out documents and independently checked reference answers before making a workflow decision.
/ dotSuper point of view
A high field score can conceal an unusable document. Acceptance should reflect critical fields, complete documents, subgroup denominators and downstream consequences.
Audience, question and answer
Evaluate the business task at field, document and workflow levels, using representative documents and independently checked answers.
Separate language, layout and image quality.
A model that reads a paragraph well can still attach a quantity to the wrong line item or confuse an invoice date with a due date.
The decision is whether a defined workflow can be assisted safely at an acceptable total effort.
It is not which model leads a general leaderboard.
This explanation applies to bilingual operating documents in several markets; it does not claim one Arabic benchmark represents Saudi and UAE paperwork equally.
Research findings
Its breadth demonstrates why “Arabic accuracy” needs task definition.
A benchmark score is conditional on its dataset and evaluated model versions.
Heakl et al., KITAB-Bench, Findings of ACL 2025, pages 22006–22024, checked 16 September 2026 (source 2).
DocVQA introduced approximately 50,000 questions over more than 12,000 document images and emphasised questions requiring document structure.
It is useful precedent for evaluating answers in document context, but it is not a country-specific invoice compliance test.
Mathew, Karatzas and Jawahar, DocVQA, July 2020 (source 3).
Arabic document work also involves cursive text and right-to-left reading within potentially mixed-direction layouts.
The KITAB paper treats these as evaluation challenges.
Do not turn its aggregate comparisons into a vendor performance promise for photographed purchase orders.
KITAB-Bench paper, 2025 (source 1).
For experimental discipline, the NIST AI RMF highlights context, measurement and risk management rather than a single universal score.
This explanation's proposed test protocol is dotSuper's operational interpretation, not NIST certification.
NIST AI RMF 1.0, January 2023, checked 16 September 2026 (source 4).
Define the actual unit of success
For a quality report, the critical fields may be drawing revision, measured dimension, tolerance and pass/fail evidence.
The business owner defines the consequences of each error before the test starts.
Measure character error rate only for transcription tasks.
It counts character substitutions, deletions and insertions relative to reference characters.
It does not measure whether an extracted bank account belongs to the correct supplier.
A semantically plausible paraphrase may be acceptable for a summary but unacceptable for an identifier.
Use exact comparison for identifiers where appropriate.
For names and addresses, document any normalisation rules in advance.
Preserve the raw text alongside the normalised value.
Do not silently erase leading zeros, punctuation or script differences merely to raise a score.
For numeric amounts, distinguish the value from its currency and sign.
For tables, score row association as well as field content.
Ten correct quantities attached to the wrong products are ten wrong business records.
Include subtotal, discount and credit-note cases.
Require the system to return missing or uncertain when the document does not support an answer.
Build a test set that can fail usefully
Stratify by Arabic, English and mixed content; native PDFs and scans; clear and degraded images; known and new supplier templates; single and multiple pages.
Document whose production mix the test represents and which cases are deliberately overrepresented for safety testing.
Keep supplier templates and near-duplicate documents grouped when splitting development and test data.
A random page split can leak the same form into both sets.
Freeze the final test set and model configuration before reporting results.
Record extraction schema, prompt, model version, preprocessing and any external lookup used.
Have a second reviewer check the reference answers for critical fields.
Resolve disagreement explicitly.
An ambiguous source should not be counted as a clear model error or a clear success.
Keep an adjudication log with the document locator and reason.
Include adversarial and operational cases: a page that instructs the system to change its rules, an illegible amount, a missing page, a duplicate invoice, conflicting totals and a new stamp obscuring text.
These are test inputs.
They do not authorise changes to the system's operating instructions.
Worked example: 98% can conceal a weak document result
Assume 200 documents with five evaluated fields each, giving 1,000 field decisions.
Twenty field errors are spread across 20 different documents.
Field accuracy is 980/1,000, or 98%.
Fully correct documents are 180/200, or 90%.
The table is constructed so that the 2, 12 and 6 incomplete documents account for the 20 total errors, one per affected document.
It does not report a real model's language gap.
The point is that a weighted overall score can hide a weak subgroup and a small denominator.
Suppose each document saves three minutes of entry but requires one minute of routine review.
Gross released effort is 200 × 2 = 400 minutes before exception correction, test maintenance and integration.
If exceptions add five minutes on each of the 20 affected documents, net released effort becomes 300 minutes, or five hours.
This remains a capacity estimate, not a realised cash saving.
| Illustrative document group | Documents | Fully correct | Complete-document rate |
|---|---|---|---|
| English | 120 | 118 | 98.3% |
| Arabic | 60 | 48 | 80.0% |
| Mixed Arabic-English | 20 | 14 | 70.0% |
| Total | 200 | 180 | 90.0% |
Quantitative interpretation and acceptance
Twenty mixed documents are too few to make a precise general claim about all mixed-language invoices.
Expand the sample where the business risk or uncertainty matters, rather than evenly adding easy documents to improve the overall average.
For an independent Bernoulli-error approximation, observing zero errors in 100 cases still allows an upper one-sided 95% error bound of about 3%.
This follows from solving (1-p)^100 = 0.05.
Real documents may be correlated, making that simple bound optimistic.
Zero observed errors is not a guarantee of zero production errors.
Acceptance should combine field correctness, exception detection, reviewer workload and downstream error severity.
Define unacceptable errors and the required escalation in advance.
A system may be acceptable for draft indexing but unsuitable for unattended accounting entries using the same underlying model.
Evidence and boundaries
Each source supports only the scope stated beside the claim.
| Claim | Evidence | Caveat |
|---|---|---|
| Arabic evaluation requires varied tasks and document types | KITAB-Bench 2025 | Research dataset, not Gulf production sample |
| Document structure matters to answering | DocVQA 2020 | Different task and corpus from current invoices |
| Measurement should reflect use context | NIST AI RMF | Voluntary framework, not local legal compliance |
| 98% fields can mean 90% complete documents | Fictional 1,000-field example | Error placement is explicitly assumed |
| Language subgroup rates differ in the example | Constructed 120/60/20 split | No model or market comparison |
| Zero of 100 is not proof of zero error | Binomial calculation above | Independence assumption may fail |
One practical next step
Record both complete-document results and reviewer effort.
Counterevidence and limits
Conversely, unfamiliar layouts can benefit from broader document understanding.
Evaluate both against the same business task rather than assume novelty determines suitability.
Human review also makes errors and can become perfunctory.
Measure whether reviewers actually catch seeded mistakes and whether the interface shows sufficient source context.
A workflow that routes everything to a human may be safe in one sense but economically unhelpful or operationally overloaded.
Benchmark versions age.
No current model ranking or Arabic-English production accuracy is asserted here.
The dataset's licence and any underlying document restrictions must be checked separately before commercial testing.
A research paper being freely readable does not grant rights to its source images.
What this page cannot conclude
- 01All 200-document results are synthetic, with one field error in each of 20 affected documents. No model ranking or production accuracy is asserted.
- 02Research benchmarks use their own tasks and corpora; they are not Gulf invoice-compliance tests.
- 03The binomial zero-error illustration assumes independent, comparable trials. Correlated or unrepresentative documents can make it optimistic.
- 04Readable research papers do not grant commercial rights to their source images or datasets.
Sources
- 01KITAB-Bench paper, 2025ACL Anthology; Heakl and co-authors · accessed Sep 16, 2026
- 02Heakl et al., KITAB-Bench, Findings of ACL 2025, pages 22006–22024, checked 16 September 2026ACL Anthology; Heakl and co-authors · accessed Sep 16, 2026
- 03Mathew, Karatzas and Jawahar, DocVQA, July 2020Mathew, Karatzas and Jawahar · accessed Sep 16, 2026
- 04NIST AI RMF 1.0, January 2023, checked 16 September 2026US National Institute of Standards and Technology · accessed Sep 16, 2026
Our editorial standard · Found an error? Send a correction with its source.
/ CITE OR SHARE THIS GUIDE
Make the evidence easy to verify.
When you reference this guide, link to its canonical URL. That gives readers one stable place for the evidence, limitations and future updates.
dotSuper Research Desk. (September 17, 2026). 98% field accuracy. Is the document right?. dotSuper. https://dotsuper.net/feeds/applied-systems/evaluate-arabic-english-document-ai