OCR vs VLM vs Parser: Which approach processes documents reliably?
Research overview, 11 October 2026. This article analyses officially published model documentation and benchmark results. It does not present new independently run model experiments. Figures remain tied to the benchmark release, evaluation protocol and metric stated by their source. Recheck the latest checkpoints before future publication.
Short answer: there is no universal OCR winner
A digital PDF with a reliable text layer often does not need a large vision model. Complex scanned tables may benefit from specialised document VLMs. General vision-language models may help interpret invoice fields, but plausible output is not the same as a correct financial record. In an ERP workflow, the critical measures are correct required fields, safe abstention, verifiable evidence and cost per accepted transaction.
We therefore separate parsers and conventional OCR, specialised document VLMs, general-purpose vision models and managed extraction services. Only an end-to-end test against identical business criteria can compare complete pipelines.
For buying decisions and operating models, see AI document processing. Our practical email-to-PDF-to-ERP workflow guide explains the controlled implementation.
Four categories with different jobs
The categories must be evaluated separately before comparing end-to-end outcomes.
| Category | Examples | Primary task | What the score does not prove |
|---|---|---|---|
| Conventional parsers and OCR | PDF text layers, Tesseract, Docling with a specified backend | Recognise text, reading order and document structure | Correct invoice fields or amounts |
| Specialised document / OCR VLMs | GLM-OCR, DeepSeek-OCR 2, Qwen-OCR, PaddleOCR-VL | Convert images to text, Markdown, tables and structured elements | Safe ERP posting |
| General vision-language models | Qwen3.8-27B, Gemma 4 31B | Interpret document imagery and produce structured invoice fields | Better OCR simply because a model has more parameters |
| Managed extraction services | Google Document AI, Azure AI Document Intelligence | Extract fields and tables through a managed API | Provider confidence equals business correctness |
Categories overlap technically. Docling is a framework, not a single OCR model. An OCR service may use multiple internal models. Cloud services and local models have different ownership, deployment and pricing. Each benchmark must identify the actual pipeline, not just a marketing category.
What official benchmarks actually measure
OmniDocBench measures document parsing, not accounting accuracy
OmniDocBench evaluates aspects of document parsing. Its published tables may include both specialised and general vision models. Results are meaningful only with the same benchmark version, split and evaluation metric.
Illustrative values documented in a public OmniDocBench evaluation:
| Model | Type | Reported overall value |
|---|---|---|
| DeepSeek-OCR 2 | Specialised document VLM | 90.25 |
| Qwen3-VL-235B | General VLM, older generation | 89.78 |
Source: OmniDocBench project leaderboard. This is not a Qwen3.8 score and not a leaderboard for Austrian or German invoices. Before reproducing these numbers, freeze the specific leaderboard revision, dataset version and evaluation implementation.
GLM-OCR: a vendor-reported figure
Z.ai reports 94.62 on OmniDocBench v1.5 for GLM-OCR (0.9B parameters). Its official project describes a document-processing system involving layout analysis and parallel recognition. This is a published provider result, not an independently reproduced test in this article.
Source: Official GLM-OCR repository.
Gemma 4: a different metric
Google's Gemma 4 model card reports an average edit distance of 0.131 on OmniDocBench 1.5 for Gemma 4 31B; lower is better. It must not be rebranded as “87% accuracy” or compared by subtracting it from GLM-OCR's 94.62. Without matched inputs, evaluator and metric, there is no valid direct ranking.
Practical inference: a compact specialist may perform very well at document recognition while a general-purpose vision model may be valuable for semantic field interpretation. Neither statement proves which one handles a particular invoice correctly.
Current candidates and what to test
| Candidate | Category | Appropriate test | Evidence limit |
|---|---|---|---|
| GLM-OCR | Specialist document OCR | Scans and complex tables | Official v1.5 figure; no own invoice field evaluation |
| DeepSeek-OCR 2 | Specialist document OCR | Reading order and document-to-Markdown | Official project/benchmark; no proven DACH invoice accuracy |
| Qwen-OCR / qwen-vl-ocr | Specialised OCR API | Document text and extraction | Record API/model revision; not interchangeable with a local Qwen VLM |
| PaddleOCR-VL | Specialist document OCR | Text, layout and tables | Pin exact checkpoint and release before testing |
| Qwen3.8-27B | General vision model | Direct invoice-to-JSON extraction | No comparable own invoice benchmark |
| Gemma 4 31B | General vision model | Visual field identification and reasoning | Vendor card figures, not an invoice test |
| PDF parser / Tesseract | Deterministic / conventional | Digital PDF and OCR baseline | Requires a separate field extractor |
| Google Document AI / Azure Document Intelligence | Managed extraction | Standard invoice fields over an API | Specify processor, API release, region and pricing |
DeepSeek-OCR 2 was announced on 27 January 2026. Exact model cards, checkpoints, availability and licences must be confirmed before a new experiment. Historical Qwen3-VL results cannot be attributed to Qwen3.8. Names alone do not specify quantisation, prompt, image resolution or runtime.
Why invoice processing requires different error analysis
Near-perfect OCR text can still produce the wrong accounting record. Consider these designed evaluation cases, not observed model error rates:
| Situation | Failure mode | Required control |
|---|---|---|
| “Invoice total” and “amount due” are both present | Wrong semantic amount | Field definition, document status and source location |
| Credit note resembles an invoice | Wrong sign or document class | Classification and business rules |
| Net, tax and gross amounts conflict | Misaligned rows or rounding | Tax and sum reconciliation |
| Invoice number resembles a customer number | Plausible incorrect identifier | Original evidence and context validation |
| Same invoice arrives by two emails | Duplicate posting despite correct extraction | Durable idempotency and duplicate detection |
| Document asks to change payment details | Business-data change or prompt injection | Supplier master-data check; no model-controlled write actions |
For these risks, the relevant outcome is not merely readable Markdown but an accurately identified, validated and appropriately approved transaction.
Compare five complete pipelines instead of mixing raw model scores
A conceptual reference workflow, not a measured result.
| Pipeline | Potential advantage | Main limitation |
|---|---|---|
| Text-layer PDF + deterministic rules | Simple and auditable | Layout diversity and field semantics |
| OCR + fixed field extractor | Works with scans; modular | Recognition errors propagate |
| Document VLM + field extractor | May preserve complex layout better | Extra components and interface complexity |
| General vision model → structured JSON | Combined visual and semantic interpretation | Confidently wrong values and compute cost |
| Hybrid routing + controlled review | Uses simple paths first and escalates exceptions | Orchestration and evaluation effort |
For a first pilot, parse structured electronic invoices directly. Inspect the text layer of digital PDFs. Escalate difficult scans to OCR/VLM and keep extracted values as proposals until schema, arithmetic, supplier and duplicate checks succeed. The workflow architecture guide covers controlled approval, retries and safe ERP writes.
Decision metrics that matter to an IT buyer
A single accuracy figure is inadequate for invoice fields. Google Document AI's evaluation documentation explains precision, recall and F1 for extracted entities.
| Metric | What it answers |
|---|---|
| OCR CER / WER | How many characters or words are transcribed incorrectly? |
| Normalised field exact match | Which required fields match ground truth exactly? |
| Precision / recall / F1 | What is the balance of correct, extra and missing entities? |
| Document pass rate | What share of documents has every required field correct? |
| Silent critical error rate | How often would materially wrong values pass without review? |
| Review rate and active review time | How much human work remains? |
| End-to-end p50 / p95 latency | What is processing time including retries and validation? |
| Cost per accepted document | What is the full cost including inference, operations and human review? |
The cheapest inference is not necessarily the cheapest document workflow. Rework and manual verification can dominate operating costs.
A reproducible German-language invoice benchmark
A practical proposed pilot design contains 100 approved or synthetic documents, with 60 development examples and 40 untouched test examples. This experiment has not yet been run. Such a small test cannot establish rare-error reliability; larger samples would be required.
Report separate results for text-layer PDFs, scanned invoices, smartphone receipts, multi-page line-item invoices and structured e-invoices. Include credit notes, multiple tax rates, mixed currencies, ambiguous totals, partially unreadable data and duplicate submissions.
Freeze the normalised target schema and adjudicated ground truth first. Pin checkpoint revisions, preprocessing, prompts, quantisation, hardware, API region, retries and the scoring implementation. Repeated model runs on the same invoice are not additional independent documents.
Publish new rankings only when raw evaluation reports, denominators, failure cases and executable configurations are available.
Conclusion: select the processing path, not the model with the nicest leaderboard
Official benchmarks are useful for shortlisting. They do not prove that a system can safely post Austrian or German invoices to an ERP.
Start with deterministic parsing for genuinely structured inputs. Evaluate specialised document VLMs for difficult scans and complex pages. Include general vision models such as Qwen3.8-27B and Gemma 4 31B where semantic field interpretation matters. Decide using the number of safely completed documents, review burden and fully loaded costs.
The valuable output is a correct, controlled business transaction – not beautiful Markdown.
Compare integration options, review boundaries and deployment choices in the AI document processing decision guide.
Primary sources and evidence status
- OmniDocBench – benchmark and published results. Public project data; not our own evaluation.
- Z.ai – GLM-OCR. Vendor-reported OmniDocBench v1.5 result.
- DeepSeek – OCR 2. Official project release.
- Qwen – Qwen3.8 model family. Official release timeline and model family documentation.
- Google DeepMind – Gemma 4 model card. Model documentation and vendor benchmarks.
- Google Cloud – evaluating document processors. Precision, recall and F1 methodology.
Transparency: This is a secondary analysis of official sources, not an independent reproduction of their scores. Differences in versions, metrics, input distributions and end-to-end design prohibit a universal ranking.