All posts
Document AIOCRVision Language ModelsBenchmark

OCR vs VLM vs Parser: Which approach processes documents reliably?

9 min readThomas Stermole
Four document-processing categories: parser, document VLM, general vision model and managed extraction.

Research overview, 11 October 2026. This article analyses officially published model documentation and benchmark results. It does not present new independently run model experiments. Figures remain tied to the benchmark release, evaluation protocol and metric stated by their source. Recheck the latest checkpoints before future publication.

Short answer: there is no universal OCR winner

A digital PDF with a reliable text layer often does not need a large vision model. Complex scanned tables may benefit from specialised document VLMs. General vision-language models may help interpret invoice fields, but plausible output is not the same as a correct financial record. In an ERP workflow, the critical measures are correct required fields, safe abstention, verifiable evidence and cost per accepted transaction.

We therefore separate parsers and conventional OCR, specialised document VLMs, general-purpose vision models and managed extraction services. Only an end-to-end test against identical business criteria can compare complete pipelines.

For buying decisions and operating models, see AI document processing. Our practical email-to-PDF-to-ERP workflow guide explains the controlled implementation.

Four categories with different jobs

Diagram showing four distinct document AI model categories and their tasks.The categories must be evaluated separately before comparing end-to-end outcomes.

CategoryExamplesPrimary taskWhat the score does not prove
Conventional parsers and OCRPDF text layers, Tesseract, Docling with a specified backendRecognise text, reading order and document structureCorrect invoice fields or amounts
Specialised document / OCR VLMsGLM-OCR, DeepSeek-OCR 2, Qwen-OCR, PaddleOCR-VLConvert images to text, Markdown, tables and structured elementsSafe ERP posting
General vision-language modelsQwen3.8-27B, Gemma 4 31BInterpret document imagery and produce structured invoice fieldsBetter OCR simply because a model has more parameters
Managed extraction servicesGoogle Document AI, Azure AI Document IntelligenceExtract fields and tables through a managed APIProvider confidence equals business correctness

Categories overlap technically. Docling is a framework, not a single OCR model. An OCR service may use multiple internal models. Cloud services and local models have different ownership, deployment and pricing. Each benchmark must identify the actual pipeline, not just a marketing category.

What official benchmarks actually measure

OmniDocBench measures document parsing, not accounting accuracy

OmniDocBench evaluates aspects of document parsing. Its published tables may include both specialised and general vision models. Results are meaningful only with the same benchmark version, split and evaluation metric.

Illustrative values documented in a public OmniDocBench evaluation:

ModelTypeReported overall value
DeepSeek-OCR 2Specialised document VLM90.25
Qwen3-VL-235BGeneral VLM, older generation89.78

Source: OmniDocBench project leaderboard. This is not a Qwen3.8 score and not a leaderboard for Austrian or German invoices. Before reproducing these numbers, freeze the specific leaderboard revision, dataset version and evaluation implementation.

GLM-OCR: a vendor-reported figure

Z.ai reports 94.62 on OmniDocBench v1.5 for GLM-OCR (0.9B parameters). Its official project describes a document-processing system involving layout analysis and parallel recognition. This is a published provider result, not an independently reproduced test in this article.

Source: Official GLM-OCR repository.

Gemma 4: a different metric

Google's Gemma 4 model card reports an average edit distance of 0.131 on OmniDocBench 1.5 for Gemma 4 31B; lower is better. It must not be rebranded as “87% accuracy” or compared by subtracting it from GLM-OCR's 94.62. Without matched inputs, evaluator and metric, there is no valid direct ranking.

Practical inference: a compact specialist may perform very well at document recognition while a general-purpose vision model may be valuable for semantic field interpretation. Neither statement proves which one handles a particular invoice correctly.

Current candidates and what to test

CandidateCategoryAppropriate testEvidence limit
GLM-OCRSpecialist document OCRScans and complex tablesOfficial v1.5 figure; no own invoice field evaluation
DeepSeek-OCR 2Specialist document OCRReading order and document-to-MarkdownOfficial project/benchmark; no proven DACH invoice accuracy
Qwen-OCR / qwen-vl-ocrSpecialised OCR APIDocument text and extractionRecord API/model revision; not interchangeable with a local Qwen VLM
PaddleOCR-VLSpecialist document OCRText, layout and tablesPin exact checkpoint and release before testing
Qwen3.8-27BGeneral vision modelDirect invoice-to-JSON extractionNo comparable own invoice benchmark
Gemma 4 31BGeneral vision modelVisual field identification and reasoningVendor card figures, not an invoice test
PDF parser / TesseractDeterministic / conventionalDigital PDF and OCR baselineRequires a separate field extractor
Google Document AI / Azure Document IntelligenceManaged extractionStandard invoice fields over an APISpecify processor, API release, region and pricing

DeepSeek-OCR 2 was announced on 27 January 2026. Exact model cards, checkpoints, availability and licences must be confirmed before a new experiment. Historical Qwen3-VL results cannot be attributed to Qwen3.8. Names alone do not specify quantisation, prompt, image resolution or runtime.

Why invoice processing requires different error analysis

Near-perfect OCR text can still produce the wrong accounting record. Consider these designed evaluation cases, not observed model error rates:

SituationFailure modeRequired control
“Invoice total” and “amount due” are both presentWrong semantic amountField definition, document status and source location
Credit note resembles an invoiceWrong sign or document classClassification and business rules
Net, tax and gross amounts conflictMisaligned rows or roundingTax and sum reconciliation
Invoice number resembles a customer numberPlausible incorrect identifierOriginal evidence and context validation
Same invoice arrives by two emailsDuplicate posting despite correct extractionDurable idempotency and duplicate detection
Document asks to change payment detailsBusiness-data change or prompt injectionSupplier master-data check; no model-controlled write actions

For these risks, the relevant outcome is not merely readable Markdown but an accurately identified, validated and appropriately approved transaction.

Compare five complete pipelines instead of mixing raw model scores

Diagram of controlled document routing from structured input to validation and approved ERP export.A conceptual reference workflow, not a measured result.

PipelinePotential advantageMain limitation
Text-layer PDF + deterministic rulesSimple and auditableLayout diversity and field semantics
OCR + fixed field extractorWorks with scans; modularRecognition errors propagate
Document VLM + field extractorMay preserve complex layout betterExtra components and interface complexity
General vision model → structured JSONCombined visual and semantic interpretationConfidently wrong values and compute cost
Hybrid routing + controlled reviewUses simple paths first and escalates exceptionsOrchestration and evaluation effort

For a first pilot, parse structured electronic invoices directly. Inspect the text layer of digital PDFs. Escalate difficult scans to OCR/VLM and keep extracted values as proposals until schema, arithmetic, supplier and duplicate checks succeed. The workflow architecture guide covers controlled approval, retries and safe ERP writes.

Decision metrics that matter to an IT buyer

A single accuracy figure is inadequate for invoice fields. Google Document AI's evaluation documentation explains precision, recall and F1 for extracted entities.

MetricWhat it answers
OCR CER / WERHow many characters or words are transcribed incorrectly?
Normalised field exact matchWhich required fields match ground truth exactly?
Precision / recall / F1What is the balance of correct, extra and missing entities?
Document pass rateWhat share of documents has every required field correct?
Silent critical error rateHow often would materially wrong values pass without review?
Review rate and active review timeHow much human work remains?
End-to-end p50 / p95 latencyWhat is processing time including retries and validation?
Cost per accepted documentWhat is the full cost including inference, operations and human review?

The cheapest inference is not necessarily the cheapest document workflow. Rework and manual verification can dominate operating costs.

A reproducible German-language invoice benchmark

A practical proposed pilot design contains 100 approved or synthetic documents, with 60 development examples and 40 untouched test examples. This experiment has not yet been run. Such a small test cannot establish rare-error reliability; larger samples would be required.

Report separate results for text-layer PDFs, scanned invoices, smartphone receipts, multi-page line-item invoices and structured e-invoices. Include credit notes, multiple tax rates, mixed currencies, ambiguous totals, partially unreadable data and duplicate submissions.

Freeze the normalised target schema and adjudicated ground truth first. Pin checkpoint revisions, preprocessing, prompts, quantisation, hardware, API region, retries and the scoring implementation. Repeated model runs on the same invoice are not additional independent documents.

Publish new rankings only when raw evaluation reports, denominators, failure cases and executable configurations are available.

Conclusion: select the processing path, not the model with the nicest leaderboard

Official benchmarks are useful for shortlisting. They do not prove that a system can safely post Austrian or German invoices to an ERP.

Start with deterministic parsing for genuinely structured inputs. Evaluate specialised document VLMs for difficult scans and complex pages. Include general vision models such as Qwen3.8-27B and Gemma 4 31B where semantic field interpretation matters. Decide using the number of safely completed documents, review burden and fully loaded costs.

The valuable output is a correct, controlled business transaction – not beautiful Markdown.

Compare integration options, review boundaries and deployment choices in the AI document processing decision guide.

Primary sources and evidence status

Transparency: This is a secondary analysis of official sources, not an independent reproduction of their scores. Differences in versions, metrics, input distributions and end-to-end design prohibit a universal ranking.

Next step

Sounds relevant for your company?

In a no-obligation initial call, we clarify within 30 minutes whether and where getting started is worthwhile for you — honestly and without sales pressure.

Request an initial call