Back to Blog
AI Implementation

AI document processing: invoices, forms and contracts explained

How AI document processing works for invoices, forms and contracts: the pipeline, what cloud services offer, real accuracy evidence and human review steps.

K

Klevere AI Team

AI Implementation

5 October 202612 min read

Most finance and operations teams still have a person retyping figures from PDFs into an accounting system. The government has now said the UK will introduce mandatory e-invoicing for all VAT invoices from 2029, which makes structured, machine-readable documents the direction of travel. Until every supplier catches up, though, businesses will keep receiving scans, photographs, emailed PDFs and awkward spreadsheets, and something has to read them.

That something is AI document processing: software that reads invoices, forms and contracts, pulls out the fields you care about, checks them against your rules and passes clean data to the system that needs it. This guide explains how it works, what the major cloud services actually do, where accuracy breaks down and how to build a process you can trust with real money.

Quick answer

AI document processing combines optical character recognition with language and layout models to turn unstructured documents into structured data. Good systems return a confidence score per field, validate results against business rules, and send doubtful items to a human review queue. The goal is not zero humans. It is humans only where they add value.

What AI document processing actually does

Older automation relied on templates. You told the software that on supplier A's invoice the total sits at a fixed position, and it broke when supplier A changed their layout. Modern document AI reads the page more like a person does. It identifies text, works out the layout, and uses context to decide that a number near the words amount due is the total payable.

The cloud vendors describe this in similar terms. Amazon says Textract extracts relevant data from almost any invoice or receipt without the need for any templates or configuration. Google says its Invoice Parser can extract up to 46 generic entities from invoices, including invoice number, supplier name, invoice amount, tax amount, invoice date and due date. Microsoft states that its invoice model currently supports invoices in 27 languages and accepts JPEG, PNG, PDF and TIFF files.

You will see several overlapping terms for this work: intelligent document processing, AI invoice processing, AI data extraction and document automation. They describe the same pipeline seen from different angles. Whatever the label, the stages are consistent, and understanding them is the fastest way to judge a product or a proposal.

The pipeline, stage by stage

1. Ingestion and classification

Documents arrive by email, shared folder, supplier portal or scanner. The first job is to work out what each file is. An invoice, a credit note, a delivery note and a terms-and-conditions attachment may all arrive in one PDF. Classification and splitting stop the wrong model reading the wrong page, and they matter more than most buyers expect.

2. Reading the page

OCR converts pixels to characters. Layout analysis then recovers structure: tables, columns, headers, stamps and handwriting. This stage is where poor scans hurt. Blurred photographs, skewed pages and shadows reduce the quality of everything downstream, so capture quality is a legitimate part of your project, not an afterthought.

3. Field and line-item extraction

Extraction identifies the header fields (supplier, number, dates, totals, tax) and the line items (description, quantity, unit price, product code). Amazon documents that AnalyzeExpense standardises varied labels into a consistent taxonomy, so that bill number, invoice number and receipt number all map to one field. That normalisation is what lets one downstream process handle thousands of different layouts.

4. Validation against rules

Extraction on its own is a guess. Validation turns a guess into something you can post. Typical checks include whether line items sum to the net total, whether tax matches the stated rate, whether the supplier exists in your master data, whether the invoice number has been seen before, and whether the amount sits within a purchase order tolerance. A figure that reads cleanly but fails arithmetic is exactly the case humans should see.

5. Exception handling and human review

Documents that fail validation or carry low confidence go to a queue. A reviewer sees the original document beside the extracted fields, corrects what is wrong and approves. Good designs record each correction, because corrections are the best training and testing data you will ever have.

6. Posting and audit trail

Approved data is posted to the accounting system, ERP or case management tool, with the source document attached and a record of who or what approved it. If you use Xero, our guide to Xero AI integration covers how extracted data typically reaches bills and contacts.

What the major services offer, and what they do not

The three big cloud providers all sell pre-trained invoice extraction, and they are a sensible starting point for many projects. Each has trade-offs worth understanding before you commit.

Amazon Textract returns a confidence score for every piece of data it detects, which is the foundation of any routing logic. Its documentation shows examples such as 97.1 for a label and 99.9 for a value. Because it needs no templates, onboarding a new supplier does not require configuration.

Microsoft Azure Document Intelligence extracts customer and vendor details, due dates, amounts and line items, and its page also lists related document types such as utility bills, sales orders and purchase orders. It publishes file limits, with larger files allowed on the paid tier than on the free tier, and a monthly free allowance of 500 pages on its pricing page. Check that page for current rates in your region and currency, because they vary.

Google Document AI offers a pre-trained Invoice Parser and, separately, a custom extractor with generative AI for documents that fall outside the standard shapes. One detail worth knowing if you read older tutorials: Google's own deprecations table lists its Human in the Loop feature as deprecated from 16 January 2024. If you choose Google, plan to build your own review queue rather than relying on a built-in one.

None of these services is a finished business process. They read documents. They do not know your chart of accounts, your approval limits, your supplier master data or what to do when an invoice references a purchase order that does not exist. That surrounding logic is where most of the work, and most of the value, sits.

How accurate is it really?

Vendor marketing tends to quote a single accuracy number. In practice accuracy depends on the document, the field and the capture quality, and it is worth testing on your own files before you believe anyone.

Independent evidence shows the spread. A recent academic evaluation of Amazon Textract on 118 receipts from Australian vendors found vendor name accuracy of 68.7 per cent for text-based receipts, with problems including multiple string variants for the same vendor, partial parsing of non-English product names and missing fields on blurred or angled photographs. Totals were detected consistently. Receipts are harder than clean digital invoices, so treat that figure as a floor for messy inputs rather than a prediction for yours, but it illustrates why field-level measurement matters.

Confidence scores deserve similar scepticism. Researchers at Amazon Web Services built ConfBench, a benchmark of 1,346 document variants made by degrading 75 verified invoices, to test how far model-reported confidence can be trusted. Their paper, Can You Trust the Confidence?, found that calibration varies dramatically across models, from near-perfect to severely overconfident, and that combining OCR text with the page image gave more accurate confidence estimates. The practical lesson is simple. Never assume a confidence of 95 means 95 per cent of such values are right. Measure it against your own corrected data and set thresholds from evidence.

A sound way to measure is to take a sample of a few hundred real documents, have a person key the correct values once, and score the system field by field. Report three numbers: the share of fields extracted correctly, the share of documents that pass validation with no human touch, and the number of errors that reach the ledger undetected. The last figure is the one that costs money.

Where it goes wrong

Most failures are predictable, which means most can be designed around.

  • Supplier name variants. The same supplier appears as a trading name, a legal name and a logo, which breaks matching against master data. Fuzzy matching plus a confirmed alias list fixes most of it.
  • Date ambiguity. A date written 07/04/25 means different things in the UK and the US. Set the expected locale per supplier rather than per system.
  • Multi-page and multi-document files. Remittance advice stapled to an invoice, or three invoices in one PDF, will confuse a model that assumes one document per file.
  • Line items on invoice-style documents. The same Textract evaluation noted that item arrays stayed empty on some invoice-style receipts even though summary fields were extracted. Always check line-item coverage separately from header fields.
  • Poor scans. Photographs taken at an angle in poor light remain the biggest single driver of errors, so improve capture at source where you can.
  • Confident mistakes. A model can be wrong and sure about it. That is why arithmetic and master-data checks sit alongside confidence thresholds, not behind them.
  • For a wider view of how production agents fail and how to contain the damage, see our piece on what can go wrong with AI agents in production.

    Human review that actually works

    A review queue only protects you if reviewers genuinely review. The UK Information Commissioner's Office warns that non-meaningful human review is caused by automation bias or a lack of interpretability, and says reviewers need appropriate knowledge, experience, authority and independence to challenge outputs. Its audit guidance also recommends documenting target accuracy rates and acceptable tolerances, and keeping a log of when a human overrides the AI together with the reasons.

    Those points translate neatly into design choices for a finance workflow. Show the reviewer the source image next to the extracted value, rather than only a list of fields. Highlight the specific fields that triggered the review so attention goes where the risk is. Seed the queue occasionally with known-wrong items to check people are paying attention. Log every override. And review the log monthly, because patterns in corrections tell you which supplier, field or rule needs fixing.

    Invoices usually contain personal data of some kind, from sole-trader supplier details to employee expense claims. Retention, access control and processor agreements therefore matter. Our guide to GDPR and custom AI agents in the UK and EU sets out the questions to settle before documents touch a model.

    Contracts and forms are a different problem

    Invoices are the friendliest case because they share a fairly standard set of fields. Contracts and free-form documents are harder, and it helps to be honest about the difference.

    Forms such as onboarding packs, claims or applications have fixed questions but varying layouts, handwriting and checkboxes. Extraction works well when fields are consistent and degrades with handwriting and poor scans. The review queue does more of the work here.

    Contracts are about clauses, not fields. The useful tasks are finding specific terms, comparing them against a playbook and flagging deviations, such as a notice period, a liability cap or an auto-renewal date. Language models are strong at this kind of reading, but the output should be framed as a flag for a person with the authority to decide, never as a verdict. Extracted dates and amounts still need validation against the source text, ideally with a quoted passage and page reference attached to every answer so a reviewer can check in seconds.

    If you are weighing a build against a packaged tool, our comparison of a custom AI agent against an off-the-shelf tool walks through the decision.

    How Klevere approaches document processing

    Klevere has deployed 500+ AI agents across 50+ projects and 12 industries, and document-heavy work comes up in almost all of them. Our approach starts with the process, not the model. Before choosing any service we look at where documents come from, who handles them today, what a wrong value costs and which system the data must reach.

    In practice we build a pipeline in the stages described above, using whichever extraction engine fits the documents and your data-residency needs, and wrapping it in validation rules written with the people who own the process. We design the review queue as a first-class screen, set routing thresholds from measured accuracy on your own sample rather than vendor claims, and keep an audit trail so finance, auditors and regulators can see what happened to each document.

    We also keep the scope honest. Some document types are not worth automating yet, and we say so. Our AI automation services page describes how engagements are structured, and the work typically begins with one document type and one downstream system so that you see measurable results early.

    Measuring the return

    Return on a document project is easy to overstate. Count only what you can measure. The usual components are staff hours per document before and after, the share of documents processed with no human touch, the error and rework rate, the time from receipt to posting and any early-payment discounts captured or late fees avoided. Compare them against the full cost of the service, hosting, integration and the reviewers who remain.

    Be careful with headline cost-per-invoice figures from vendor blogs, which use very different definitions of cost and rarely say what they include. Build the baseline from your own timesheets and error logs instead. Our guide to how to measure AI ROI sets out a method that holds up in front of a finance director.

    A practical way to start

    A cautious first project looks like this. Pick one document type with real volume, such as supplier invoices. Gather a few hundred recent examples that include the awkward ones. Run two or three extraction services over the sample and score them field by field. Write the validation rules with the person who handles the exceptions today. Run the pipeline in shadow mode, where it extracts but a person still keys the data, until its results match theirs for a sustained period. Then switch to review-by-exception for the cleanest suppliers first and widen from there.

    This sequence is slower than buying a licence and hoping, but it gives you evidence at every step and a clear stopping point if the numbers do not work. It also produces the labelled sample you need for any future model change, so switching engines later costs days instead of months.

    Frequently asked questions

    What is AI document processing?

    It is software that reads documents such as invoices, forms and contracts, extracts the information you need into structured fields, checks it against rules and sends exceptions to a person. It combines OCR, layout analysis and language models, and it replaces manual retyping rather than the judgement of the people who own the process.

    How accurate is AI invoice processing?

    It depends on the supplier, the field and the scan quality. Clean digital invoices usually extract far better than photographed receipts, and independent tests show wide variation by field. Test on a few hundred of your own documents, score each field, and set review thresholds from that evidence rather than from vendor claims.

    Do I need templates for each supplier?

    Not with current services. Amazon states that Textract works on almost any invoice or receipt without templates or configuration, and Google and Microsoft offer pre-trained invoice models too. You will still need rules for your own master data, tax treatment and approval limits, which is where setup effort goes.

    Can AI process contracts as well as invoices?

    Yes, but the task differs. With contracts the value lies in finding clauses, comparing them with your standard terms and flagging deviations, not filling fixed fields. Treat outputs as flags for a qualified person, attach the quoted passage and page to every answer, and keep a human decision-maker in the loop.

    Is it safe to send invoices to an AI service under UK GDPR?

    It can be, provided you have a lawful basis, a processor agreement, sensible retention and access controls, and meaningful human review where decisions affect people. Invoices often contain personal data. Settle data location, retention and subprocessors before go-live, and document accuracy tolerances as the ICO recommends.

    How long does a document processing project take?

    A single document type with one downstream system can often reach shadow mode within weeks, but timing depends on integration access, document variety and how quickly the validation rules are agreed. Rolling out further suppliers and document types follows once the first is stable and measured.

    If you would like to know which of your document workflows are worth automating first, and what a sensible pilot would look like, book a free AI audit and we will map it with you.

    Ready to implement AI in your business?

    Let's discuss how AI agents can transform your operations and reduce costs.