All articles

Custom Software11 min read

Intelligent Document Processing for Invoices, Forms and Contracts: How It Works

Intelligent document processing reads invoices, forms and contracts, extracts the fields you need, checks them against rules and your own records, and sends doubtful cases to a person. This guide explains each stage, where it fails and how to measure it on your own documents.

Written byUsama AsifPublished

Intelligent document processing (IDP) is software that reads business documents such as invoices, forms and contracts, works out what each one is, extracts the fields you need, checks them against rules and your own records, sends doubtful cases to a person and posts approved data into your ERP or workflow system. The software extracts the data; people stay accountable for decisions.

The extraction model gets most of the attention, but it is only one of six stages. Most of the value, and most of the risk, sits in validation, exception handling and the connection to the system the data ends up in. This guide covers each stage, the trade-offs between extraction methods, and how to test a system on your own documents before you rely on it.

How does intelligent document processing work?

An IDP system runs every document through the same pipeline: capture, classify, extract, validate, human review and post. Each stage has its own failure modes, so design and test them separately.

StageWhat happensOutputWhat typically goes wrong
1. CaptureDocuments arrive from a shared inbox, scanner, upload portal or supplier portalA file with source, sender and received timeDuplicates from forwarded emails, multiple documents in one PDF, password-protected files
2. ClassifyThe system decides what each document is: invoice, credit note, statement, remittance, purchase order, contract, application formA document type and, where useful, a layout or form versionStatements treated as invoices; one PDF holding three invoices not split
3. ExtractFields and tables are read: header fields, line items, parties, dates, clausesStructured data with a confidence score and page location per fieldWrong column read, totals taken from the wrong line, values that are not on the page
4. ValidateRules check arithmetic, master data, duplicates, purchase orders and datesPass, or a list of named exceptionsRules too loose to catch errors, or so strict everything becomes an exception
5. Human reviewPeople resolve exceptions with the document image beside the extracted dataCorrected, approved data and a reason codeReviewers rubber-stamp; corrections are not recorded
6. PostApproved data is created in the ERP, accounting or case system with the image attachedA draft bill, record or case, linked to the source filePosting failures that nobody sees; data posted twice

Two design points run through every stage. First, keep a link from every extracted value back to the page and position it came from, so reviewers and auditors can check it. Second, the pipeline should post drafts or hold records for approval, not release payments. Approval belongs to your normal workflow, which our approval workflow design guide covers in detail.

What is the difference between OCR and intelligent extraction?

OCR (optical character recognition) turns an image of a page into text. It does not know which of the numbers on an invoice is the total. Intelligent extraction adds that understanding, using templates, layout-aware models, large language models (LLMs) or a mix.

In plain words:

  • Template or zonal extraction reads a field from a fixed position on a known layout. It is precise on forms you control and breaks when the layout moves.
  • Layout-aware models learn from labelled examples where fields tend to sit and what labels sit near them, such as the number next to "Invoice No." or the columns of a line-item table. They cope with many supplier layouts of a known document type.
  • LLM-based extraction reads the text, and sometimes the page image, and fills in a schema you define: vendor name, invoice number, lines, totals. It handles unfamiliar layouts and long documents such as contracts well. The catch is that it can return a plausible-looking value that is not on the page, so its output needs the same validation as everything else.
ApproachFinds a field byStrong onWeak onYou need
Templates or zonal OCRFixed position per layoutYour own forms, stable government or bank formsAny layout change; many suppliersA template per layout and someone to maintain them
Layout-aware modelsLearned positions and nearby labelsInvoices and receipts from many suppliersNew document types, unusual tablesLabelled samples of your documents
LLM extractionReading the content against a schemaVaried layouts, contracts, free-text clausesUnsupported values, cost and speed per page, hidden instructions in documentsA strict output schema, source grounding and downstream rules

Most production systems combine methods: OCR or the PDF's own text layer for reading, a model for extraction, and deterministic rules for validation.

One LLM-specific risk deserves a line in your design. The OWASP Top 10 for LLM Applications (2025) lists prompt injection as its first risk (LLM01) and notes that indirect injection can arrive through external content such as files. An invoice could carry hidden text telling the model to mark it approved. The defence is structural: the model only extracts, it never approves, and every value still passes rules it cannot influence.

Which validation rules catch extraction errors?

Validation turns a clever extractor into a dependable system. It checks the extracted data against arithmetic, your master data and your business rules, and anything that fails goes to a person whatever the model's confidence.

Validation checklist for supplier invoices:

  • Arithmetic. Quantity times unit price equals each line amount; lines plus tax plus freight minus discount equals the total; tax equals the taxable base times the rate for that tax code, within a rounding tolerance finance sets.
  • Vendor match. Identify the vendor by tax registration number or bank account against the vendor master, not by name alone.
  • Bank details. If the bank account differs from the vendor master, hold the invoice and route it to a verification step. Never update bank details from an invoice.
  • Duplicates. Same vendor and invoice number; or same vendor, amount and date under a different number.
  • Purchase order. The PO exists, is open, belongs to this vendor and uses the same currency.
  • Dates. Invoice date not in the future and inside an open accounting period; due date consistent with the vendor's payment terms.
  • Tax. Tax registration number in a valid format for the country; tax code permitted for the vendor and item.
  • Contracts and forms. Mandatory fields present, dates in a logical order (start before end), parties matching your customer or counterparty records.

How do confidence thresholds and exception queues work?

Extraction models return a confidence score for each field. A threshold decides which values pass without review, and an exception queue holds everything else with a reason attached.

Three rules keep this honest:

  1. Set thresholds per field, not per document. A wrong description costs little; a wrong amount or bank account costs a lot.
  2. Calibrate on your own documents. A score of 0.9 is not automatically a 90 percent chance of being right. Check what each score band actually means on a labelled sample.
  3. Rules override confidence. A high-confidence total that fails the arithmetic check is still an exception.
FieldCost if wrongStarting policyAlways review when
Total and tax amountsWrong payment or tax returnHigh threshold; must pass arithmeticAny rule fails
Vendor identityPaying the wrong partyMatched to vendor master onlyNo unique match
Bank detailsFraud or misdirected paymentNever taken from the documentAlways differs from master: hold
Invoice number and dateDuplicates, wrong periodMedium thresholdDuplicate or closed-period check fails
Line descriptionsPoor reportingLower thresholdLine cannot be matched to the PO

A good exception queue shows the page image beside the extracted fields with the source highlighted, gives every exception a named reason, records what the reviewer changed, and has an owner and a target turnaround. Recorded corrections become test data for the next round of tuning.

Start strict. In the first weeks, route every document to review in shadow mode, compare the system's output with what reviewers enter, and loosen thresholds field by field only when the evidence supports it.

How does three-way match work in accounts payable?

Three-way match compares the supplier invoice with the purchase order (what was ordered, at what price) and the goods receipt (what actually arrived). The invoice is cleared for approval only when all three agree within tolerances finance has set. Services without a receipt usually use a two-way match of invoice against PO.

IDP's job is to extract the invoice lines cleanly enough for matching. The matching logic itself usually lives in the ERP or a dedicated matching service, and you should decide which early, because it shapes the integration. Common rules:

  • Price and quantity tolerances set as company policy, per category if needed
  • An invoice that arrives before the goods receipt is held, not rejected
  • Partial deliveries matched line by line, leaving the rest of the PO open
  • Variances outside tolerance routed to the buyer or receiver, with the reason shown

Connecting extraction to the ERP's PO and receipt data is integration work; our ERP integration services page explains how we approach it, and the system integration approaches guide compares the options.

What are the limits: handwriting, scans and poor-quality documents?

Extraction quality falls with input quality, and some inputs stay unreliable whatever the model.

  • Handwriting. Block capitals in boxes and tick boxes are far more reliable than joined-up writing. A system can detect that a signature is present; it cannot verify it is genuine.
  • Scan and photo quality. Skewed pages, low resolution, phone photos with shadows, stamps over totals and faded thermal receipts all reduce accuracy.
  • Tables across pages. Line items that continue onto the next page, with repeated headers and subtotals, are a common source of errors.
  • Languages and scripts. Bilingual documents, right-to-left scripts and different numeral systems need their own test samples.

The cheapest fix is often upstream: ask suppliers for PDFs generated by their software rather than scans (these carry a text layer), or for structured electronic invoices, which remove the need for extraction altogether.

How should document data be handled?

Invoices carry bank details, expense claims carry personal data, and contracts are confidential, so data handling is part of the design, not an afterthought.

  • Know where documents are processed and stored, including which region
  • If a third-party AI service is used, check its terms on data retention and whether your data may be used for training
  • Encrypt documents in transit and at rest; restrict access by role
  • Keep images and extracted data only as long as your record-retention policy requires
  • Keep sensitive values out of application logs
  • Record who viewed or changed which field, and when

Our security and data protection page explains our approach. The design is meant to support your legal obligations; confirming them remains with your own advisers.

How do you measure IDP accuracy on your own documents?

Accuracy figures quoted by vendors were measured on someone else's documents. The only number that matters is performance on yours, measured against a labelled test set.

Build the test set from real documents across your suppliers and channels, include the ugly ones, and have two people label the correct values independently. Then measure:

MetricWhat it tells you
Field accuracy for critical fieldsHow often amounts, dates, vendor and PO number are exactly right
Straight-through rateShare of documents posted with no human edit
Exception rate by reasonWhere review time goes and which rules need tuning
False acceptanceWrong values that passed without review: the metric that matters most
Review time per exceptionWhether the queue is workable for your team

A system with a lower straight-through rate and near-zero false acceptance is usually better than the reverse. Our guide on how to evaluate an AI pilot sets out how to run this kind of test before committing to a full rollout.

Worked example: an illustrative accounts payable flow

The following scenario is illustrative; the company, documents and figures are invented.

A distributor receives supplier invoices by email. One morning, a PDF arrives from an existing supplier.

  1. Capture. The file is saved with sender and time. Its hash does not match any earlier file, so it is not a duplicate upload.
  2. Classify. The system labels it an invoice (not a statement) and finds one invoice in the file.
  3. Extract. Vendor tax number, invoice number INV-4471, date, PO 10234, three lines, subtotal 1,200.00, tax 120.00 at the 10 percent rate configured for that tax code, total 1,320.00.
  4. Validate. Arithmetic passes. The vendor is matched by tax number, and the bank account agrees with the vendor master. PO 10234 is open. The goods receipt shows lines 1 and 2 received in full, but only 40 of 50 units on line 3, while the invoice bills 50. That is outside the quantity tolerance, so the invoice goes to the exception queue with the reason "quantity exceeds receipt".
  5. Human review. The buyer checks with the warehouse: the remaining 10 units arrived yesterday but were not booked. The warehouse posts the receipt, the match is re-run and passes.
  6. Post. A draft supplier bill is created in the ERP with the PDF attached and every field linked to its source, then enters the normal approval route.

A second invoice in the same batch shows a bank account that differs from the vendor master. It is held and sent to the vendor-master team, who verify the change by calling the supplier on the number already on file, not the one printed on the invoice.

When is IDP the wrong tool?

IDP solves the problem of unstructured documents arriving in volume. If that is not your problem, something simpler will usually do better.

  • Low volume. If a clerk handles the documents in a few hours a week, careful manual entry with good validation may cost less than building and maintaining a pipeline.
  • The data already exists in structured form. If suppliers or partners can send electronic invoices, EDI, CSV files or API calls, integrate those directly.
  • You control the form. Replace the paper or PDF form with a web form. Capturing data at source beats extracting it later.
  • The decision needs judgement. IDP can find renewal dates and liability clauses in contracts; a lawyer still decides what they mean.
  • The upstream process is broken. If purchase orders are not raised before buying, three-way match cannot work. Fix the process first, which is often a workflow automation project in its own right.

How Timeline Digital approaches document processing

We design IDP around the full pipeline: extraction, the validation rules your finance or operations team already applies, a usable exception queue and a clean connection to your ERP or case system. Our document processing service page describes the work. Timeline Digital starts with a free pilot of 2 to 3 key modules before the full project, so you can test the pipeline on your own documents and measure the results before committing further.

Frequently asked questions

What is the difference between OCR and intelligent document processing?

OCR converts an image of a page into text. Intelligent document processing goes further: it classifies the document, extracts specific fields such as invoice number, totals and line items, validates them against rules and your master data, routes doubtful cases to a person and posts approved data into your ERP or workflow system. OCR is usually one component inside an IDP pipeline.

Can intelligent document processing run without human review?

For a share of documents, yes, once evidence supports it. Straight-through processing should apply only when every critical field is above its confidence threshold and every validation rule passes. Anything else goes to an exception queue. Start with all documents reviewed in shadow mode, measure accuracy on your own documents, then relax thresholds field by field. Bank details should never be taken from a document without verification.

Can IDP read handwritten forms?

Partly. Block capitals written in boxes and tick boxes are read far more reliably than joined-up handwriting, and accuracy varies with writing and scan quality. Signatures can be detected as present but not verified as genuine. If you control the form, replacing it with a web or mobile form usually beats extracting handwriting later. Test handwritten samples separately before relying on them.

How do you measure the accuracy of document extraction?

Build a labelled test set from your own documents, including poor scans and unusual suppliers, with correct values agreed by two people. Measure field accuracy for critical fields, the straight-through rate, the exception rate by reason, review time, and above all false acceptance: wrong values that passed without review. Vendor accuracy figures were measured on other documents and do not transfer.

Is it safe to use large language models to extract invoice data?

It can be, with the right design. LLMs handle varied layouts well but can return values that are not on the page, and OWASP lists prompt injection, including instructions hidden in files, as the top LLM application risk. Limit the model to extraction, require every value to point to its source on the page, validate with deterministic rules, and check the provider terms on data retention.

Where does three-way matching happen in an IDP setup?

Usually in the ERP or a dedicated matching service rather than in the extraction step. IDP extracts invoice header and line data accurately enough to match; the matching logic then compares the invoice with the purchase order and goods receipt using tolerances finance has set. Deciding early where matching runs shapes the integration design and avoids building the same rules twice.

Topics in this article

  • Intelligent Document Processing
  • Invoice Automation
  • Accounts Payable
  • Three-Way Match
  • AI Development
  • Workflow Automation

Start a conversation

Tell us how your business works.

Describe what is slowing your team down. We will help you work out what to build, and how a free pilot lets you judge our work before the full project.

Prefer WhatsApp? Start a chat

What happens next

  1. You send a short brief

    The problem, the people involved and any target date. A senior engineer replies within 4 business hours.

  2. We understand your workflow

    A first call about how your business works today. An NDA can be signed before you share details.

  3. You test a free pilot

    You choose 2 to 3 key modules and we build them first, so you judge real software before the full project.