All posts
Document AIAPI integrationAutomation

Email → PDF → ERP: automate document workflows reliably

11 min readThomas Stermole
Document workflow with separate stages for intake, approval and controlled ERP transfer.

A supplier sends an invoice as a PDF. Someone opens the attachment, finds the purchase order, checks the amounts and enters the record into the ERP. Automation needs to cover that entire handoff. A plausible JSON response does not complete the transaction.

Short answer: Keep intake, validation, approval and system access in a controlled workflow. Use parsers, OCR and, where appropriate, language or vision models for extraction. Write only approved, business-validated data to the ERP and handle repeated requests with durable operation identifiers.

This guide describes a reference architecture, not an already delivered ERP integration or a BelegHelfer feature. The AI document processing decision guide covers suitable documents, deployment choices and buying versus building.

A workflow with controlled transitions

From the inbox to a confirmed ERP transaction

  1. 01

    Intake

    Isolate attachments, verify file type, preserve the original and assign an operation ID.

  2. 02

    Parsing / OCR

    Use structured data and the text layer first; process page images when needed.

  3. 03

    Extraction

    Produce a versioned proposal with document class, fields and source locations.

  4. 04

    Validation

    Check schema, amounts, required fields, master data and duplicates.

  5. 05

    Approval

    Show discrepancies alongside the original; record the decision and revision.

  6. 06

    System integration

    Transfer the approved revision idempotently and confirm the ERP record ID.

  7. 07

    Logging

    Record status, errors, versions, processing time and consumption throughout the workflow.

AI produces proposals. Rules and approvals control transitions. Each stage has its own possible failure state.

Logging accompanies every stage; it does not start after the ERP write. The original file, extracted proposal, corrected revision and ERP response remain separate artifacts connected by one operation ID.

1. Intake: emails are untrusted inputs

Limit intake to a defined mailbox or upload channel. Store attachments in isolation first. Check the actual file type, file-size and page limits, and permitted formats. File extensions alone are insufficient. Encrypted, damaged or unsupported files need a visible failure state. Malware scanning and a resource-limited parser process belong in the protection design.

A message ID helps recognise the same email. A hash of the original file recognises identical attachments. Neither replaces business-level duplicate checks: an invoice can arrive again with a different filename, while forwarding changes the email identifier.

Document text remains input data. An instruction inside a PDF to replace a bank account or send data to another URL must not change workflow rules or tool permissions. The model receives no direct ERP write access. These controls follow the principles in the OWASP prompt injection prevention guide.

2. Parsing and OCR: use the simplest workable method first

Check for machine-readable structured data first. For an XML invoice, a suitable format parser and validation are more useful than rendering it as an image and guessing the same values again. A digital PDF may contain a text layer, but its presence guarantees neither correct reading order nor complete tables.

Scans need text recognition or image processing. OCR recognises text from images. A vision-language model (VLM) processes images together with language and can identify fields in their visual context. The categories overlap: specialised OCR systems can themselves use VLMs. A parser processes the structure of a file format; a PDF pipeline may combine parsing, layout models and OCR.

Choose the extraction path from the input

  1. 01

    Is reliable structured input available?

    • Yes: parse it directly

      Validate the format and business fields; do not re-read existing structured values with vision.

    • No: inspect the PDF text layer

      Check reading order and table completeness before relying on its output.

  2. 02

    Does the text layer preserve the required information?

    • Yes: parser plus field rules

      Use the simplest workable extraction path, with evidence and validation.

    • No: test OCR or a document VLM

      Evaluate scans and difficult layout separately; add a field extractor where needed.

  3. 03

    Is the critical value genuinely unreadable?

    • Yes: recapture or human review

      Leave the value unknown; a model must not invent a required field.

    • No: extract a verifiable proposal

      Semantic interpretation can be tested with a general vision model under the same controls.

A routing recommendation, not a measured model ranking. Validate every path against the same required fields.

Docling documents different OCR engines and VLM pipelines. That does not establish a universal quality ranking. Test your own document classes, including cropped pages, rotated scans and repeated table headers. The OCR, VLM and parser research comparison explains model categories and the limits of official benchmark scores.

3. Extraction: a narrow task with traceable output

Define the output contract before writing a prompt. For invoices, it may include document type, issuer, invoice number, date, currency, gross amount and tax lines. Orders need different fields and matching rules. Unknown values remain explicitly empty; a model must not satisfy required fields with plausible additions.

Each critical field needs both a normalised value and a source location, such as a page and text span or image region. The Google Document AI reference describes text and page anchors as possible provenance. Verify separately whether your chosen processor provides suitable evidence for every required field.

Multi-stage extraction is an architecture variant to evaluate, not an automatic improvement. If a later assembler overwrites an already correct currency with EUR, the failure lies in the contract. A larger model does not correct that design error.

4. Validation: valid syntax does not prove correct business data

CheckDeterministic implementationOn failure
Output contractTypes, permitted classes, required fields and no extra action fieldsReject the proposal
AmountsDecimal arithmetic, currency and documented rounding toleranceShow the original and discrepancy
Tax linesCheck totals and associations; separate exceptional casesRequire business review
Business partnerMatch approved master data; no automatic partner creationAsk for selection or clarification
Purchase orderCompare reference, line items and approved tolerancesRoute to a discrepancy workflow
DuplicateTenant + issuer ID + invoice number; date and amount as supporting signalsShow the existing transaction
Bank accountFlag changes against the approved master recordSeparate approval; no automatic payment

Consistent arithmetic does not prove correct content: incorrectly recognised numbers can still add up. New suppliers, changed bank details and unknown document classes remain independent review reasons. The workflow does not replace tax or legal assessment.

5. Approval: correction time is part of product quality

A useful review interface shows the original alongside the proposal, highlights missing or conflicting fields and explains why review is required. It records who corrected and approved which revision. A later content change invalidates that approval and requires a new decision.

A value such as “confidence: 0.99” generated by an LLM is not a measured error probability. Microsoft explains how confidence signals should be interpreted and tested on your own documents. Combine them with business rules and measured errors. In the first pilot, review every proposal; automatic approval needs separate evidence for a narrowly defined scope.

6. ERP transfer: handle duplicates and retries together

A timeout does not mean the ERP stored nothing. Blindly repeating the request may create a second invoice. A robust integration therefore stores the transfer job durably and associates it with a stable idempotency key. One possible key combines tenant, operation and approved revision.

If the ERP supports idempotency keys, follow its documented contract. Otherwise, use a unique external reference and a reconciliation query. If neither success nor failure can be established after a timeout, mark the operation “transfer outcome unknown” and reconcile it before writing again. A local lock alone does not guarantee one execution in a remote system.

After an ERP timeout: reconcile before retrying

  1. 01

    Can the existing operation be confirmed by a stable reference?

    • Record confirmed

      Store the ERP record ID and completed state; do not create it again.

    • No reliable confirmation

      Separate a proven failure from an unknown outcome; an empty immediate search may be insufficient.

  2. 02

    Is a safe retry supported by the ERP contract?

    • Yes: retry within a fixed budget

      Use the same idempotency key and approved revision; reconcile the resulting state.

    • No or uncertain: hold the write

      Keep the outcome unknown and request reconciliation. Do not invent a new key or write blindly.

The remote system may already have saved the record. A retry is permitted only under the target system's verified contract.
FailureHandlingRetry?
Provider temporarily unavailableBounded retries with backoff; then an error queueYes, within the retry budget
Extraction violates the contractRecord the failure class and review manuallyNo endless prompt loop
Business validation failsReview task with reason and originalAfter correction
ERP rejects business dataFix mapping or master data; approve the new revisionAfter approval
ERP timeout after a possible writeReconcile using the external referenceOnly after clarification
Approval revokedStop transfer; handle any already completed write separatelyNo

A transactional outbox can connect local approval state with the transfer job. The target adapter still needs to handle uncertain remote outcomes. An LLM must not select arbitrary destination URLs, database commands or ERP actions.

7. Logging: preserve diagnostic value without excessive data

Useful operation records include state transitions, technical error class, pipeline/model/schema version, approved revision, ERP reference, duration and consumption. Do not place sensitive invoice text, API keys or complete provider responses into standard logs indiscriminately. Originals and protected diagnostic artifacts need their own access and deletion rules.

Measure three distinct times: machine processing, active human review and total elapsed time including queues. Faster extraction can produce a worse overall workflow when it requires more corrections.

Economics: cost per completed transaction

Worked assumption, not a measurement or price quotation: At 1,000 transactions per month, four minutes of manual work each and internal labour cost of €40/hour, baseline work costs about €2,667. After automation, an assumed 1.5 minutes of review costs €1,000. Adding €200 for technology and €400 for operations leaves about €1,067 in theoretical monthly benefit. With assumed setup cost of €6,000, simple payback would be about 5.6 months.

At three minutes of remaining review, the same assumptions leave only about €67 per month and roughly 90 months to pay back setup. This sensitivity shows why measured rework can change the decision. Released staff time is not automatically a cash saving. A complete calculation must also include transition, training, outages and opportunity costs.

The official Google pricing page distinguishes processor types and billing units. A low OCR page price is not the cost of a completed ERP workflow. Compare the same tasks including classification, extraction, repeated requests, review and operations.

A bounded pilot with defined stop criteria

Choose one document class, one intake channel and one target system. Measure existing manual processing time and common exceptions first. A proposed initial sample of 50–100 representative cases may reveal frequent problems; it does not establish a rare-error rate. Separate development examples from an untouched test set and, where possible, separate supplier layouts between them.

Compare a simple parser/OCR path with an additional AI variant. Use identical inputs, target fields and scoring rules. Evaluate individual fields and then the complete transaction. For variable line-item tables, precision, recall and F1 are more informative than a single “accuracy”; the Google evaluation documentation explains the distinction.

Illustrative pilot gates, to adapt before starting: no unapproved ERP writes; no second record when repeating the same approved operation; complete failure states; at least 30% less active handling time and positive monthly net benefit under conservative assumptions. These are proposed targets, not achieved results.

Stop or narrow the scope if correction work removes the benefit, the ERP cannot support safe transfer or critical errors do not reliably reach review. More models or connectors are not automatically the answer.

The next useful decision

Use the document process assessment to record document class, volume, rework and integration limits. Check existing accounting, ERP and document-management functions before developing a new service.

For the broader architecture choice, see AI agent versus workflow and planning an AI pilot. If the documents primarily support questions and knowledge retrieval, connecting a DMS to RAG is the relevant guide. Evaluate an existing prototype against the production-readiness criteria.

For Austrian self-employed businesses preparing receipts, BelegHelfer is a product concept to explore. Public status checked on 11 October 2026: pre-launch with pre-registration, no immediate app access, and no confirmed price or launch date. Its site shows a design preview with sample data; direct DATEV, BMD or ERP integrations are not promised. The ERP workflow described here is not an available BelegHelfer feature. Product information is in German.

Next step

Sounds relevant for your company?

In a no-obligation initial call, we clarify within 30 minutes whether and where getting started is worthwhile for you — honestly and without sales pressure.

Request an initial call