An AI agent that clears supplier invoices, and knows when to ask a person
A manufacturing group with six plants receives about 40,000 supplier invoices a month in every format imaginable. Twenty-two people spend most of their day typing them into the ERP. This design puts an AI agent in front of that work: it reads, checks and posts the routine invoices, and hands anything unusual to a person with the reasons laid out.
Skilled people doing data entry, and still paying late
Invoices arrive by email, through a supplier portal and on paper at plant offices. About 3,000 active suppliers use their own layouts. The accounts payable team keys each one into the ERP, then matches it against the purchase order and the goods receipt before it can be paid.
On average an invoice takes nine days from arrival to posting. At month end the backlog doubles. Early-payment discounts are missed, suppliers chase payment by phone, and an internal audit found duplicate payments in about 0.4% of a sample.
The finance director had seen generic invoice-scanning tools fail on the variety of layouts. The question was whether an agent that can reason about an invoice, look things up and follow the company’s own rules could do better, without sending pricing data outside the company.
What could not be compromised
- Invoice and pricing data must not leave the company. No public AI APIs.
- The ERP stays the system of record. The agent cannot post anything the rules do not allow.
- Every action must be explainable to statutory auditors months later.
- Payments above set limits, new suppliers and bank-detail changes always need a person.
- GST checks on every invoice: valid GSTIN, IRN for e-invoices, correct tax split.
Why an agent, and not just better OCR
Three approaches were compared on a sample of 2,000 historical invoices, scored on correct postings and on how many still needed a person.
The agent can only do what the tool gateway allows
The language model never talks to the ERP directly. The agent plans each step, but every action goes through a tool gateway that exposes a small set of approved operations, each with its own limits. Guardrails check every prompt and answer, and every step lands in an audit trail.
| Building block | Why it is there |
|---|---|
| 1 Intake and extraction | OCR and layout analysis turn any PDF, image or e-invoice into fields with a confidence score for each one. Low-confidence fields are flagged, not guessed. |
| 2 AP agent | Plans the steps for each invoice: identify the supplier, find the PO, check quantities and prices, check tax, then decide to post or escalate. It writes down its reasoning for every decision. |
| 3 Self-hosted LLM | A 30B-class open model on four GPUs in the company’s own data centre, served with batching. A smaller model handles simple, well-known suppliers. |
| 4 Tool gateway (MCP) | The only way the agent touches business systems. It exposes about a dozen operations such as “find PO”, “check goods receipt” and “post invoice under limit”, each with value limits, rate limits and its own service account. |
| 5 Guardrails and policy | Masks bank details and personal data in prompts, strips instructions hidden inside invoice text, and enforces the allow-list of tools. |
| 6 Contract and PO index | A searchable store of purchase orders, contract rates and payment terms, so the agent can check a price against what was agreed. |
| 7 Human approval queue | Exceptions arrive with the invoice, the extracted fields, what did not match and a suggested action. One click to approve, edit or reject. |
| 8 Audit trail and quality monitoring | Every input, tool call and decision is stored. A weekly sample is re-checked by the AP team and scored, and drift raises an alert. |
GPU capacity sized from invoice volume and month-end peaks
Volume is about 1,800 invoices on a normal working day and about 5,500 on the busiest month-end day. Each invoice needs on average six model calls.
| Item | Figure | Basis |
|---|---|---|
| Peak daily invoices | ~5,500 | Month end, 3x a normal day |
| Model calls per invoice | ~6 | Extraction check, matching, tax, decision, summary |
| Peak model calls | ~33,000 a day, about 1 per second over 10 hours | Comfortable for one replica, two for availability |
| Tokens per call | ~3,000 in, ~400 out | Invoice text plus PO lines and rules |
| GPUs | 4 x 48 GB (2 replicas x 2 GPUs) | 30B-class model at 8-bit, batching enabled |
| Simple-invoice route | Small model on 1 GPU | Handles about 40% of volume from known suppliers |
| Storage | ~6 TB a year | Invoice images, extracted data and audit trail, kept 8 years |
Before buying GPUs, the same workload is benchmarked on a short-term rented server with the chosen model, so the hardware order is based on measured throughput.
From shadow mode to straight-through, in four steps
The agent earns trust gradually. It starts by watching, then suggesting, and only then posting, and only within limits.
Baseline
Weeks 1 to 3
2,000 historical invoices labelled with the right answer. Tool gateway built against an ERP test system.
Gate: Field accuracy measured on the labelled set.
Shadow mode
Weeks 4 to 7
The agent processes live invoices in parallel but posts nothing. Its decisions are compared with the team’s.
Gate: 98.5% field accuracy, zero wrong supplier matches.
Assisted
Weeks 8 to 11
The agent prepares every posting and a person clicks approve. Approval time and edits are measured.
Gate: Fewer than 5% of suggestions edited.
Straight-through
Weeks 12 to 16
Routine PO-matched invoices under the limit post automatically. Everything else goes to people.
Gate: Monthly audit sample clean for two months.
The risks that matter with an agent, and how each is contained
| Risk | What could happen | How the design handles it |
|---|---|---|
| Wrong values | The model misreads an amount or invents a field | Every value is checked against the PO, the goods receipt and GST data. Mismatches go to a person, never to posting. |
| Instructions hidden in invoices | Text in an invoice tries to change the agent’s behaviour | Invoice content is treated as data only. Guardrails strip instructions, and the tool allow-list means there is nothing dangerous to call. |
| Bank-detail fraud | A fake invoice carries new bank details | The agent can never change vendor bank details. Any mismatch with the vendor master is an automatic exception. |
| Model drift | Accuracy slips as suppliers change layouts | Weekly sampled checks, alerts on falling confidence, and re-evaluation before any model update. |
| ERP outage | Postings fail part-way | The gateway queues operations and retries safely. Each posting has an idempotency key, so nothing posts twice. |
What was optimised
Routed to a small model
Known suppliers with clean layouts go to a small, fast model. The large model handles the hard cases.
Right-sized serving
Batching and quantisation fit peak volume on four GPUs instead of eight.
Tools, not free access
A dozen narrow operations are easier to test, audit and secure than broad ERP access.
Faster human review
Exceptions arrive with the reason and a suggested action, so a person decides in seconds.
Data sent outside
Model, index and logs all run on the company’s own hardware.
Measured quality
Accuracy is a number on a dashboard, not an impression.
What the design is built to deliver
| Measure | Before | Design target |
|---|---|---|
| Invoices posted without manual keying | 0% | 60 to 70% of PO-backed invoices |
| Arrival to posting | ~9 days | Under 2 days, same day for straight-through |
| AP effort on data entry | Most of the day | Down about 70%; the team moves to exceptions and supplier queries |
| Duplicate payments | ~0.4% in audit sample | Caught before posting |
| Early-payment discounts | Often missed | Captured when terms allow |
| Audit evidence | Manual sampling | Full trail for every invoice |
Targets are confirmed in the baseline and shadow phases on your own invoices. Straight-through rates depend heavily on how many invoices have a clean PO and goods receipt.
What a team needs to deliver this
Agent design
Breaking a business process into steps, tools and decision points an agent can follow safely.
Tool gateway and MCP
Narrow, testable operations with limits, service accounts and idempotent retries.
LLM serving on GPUs
Model selection, quantisation, batching and capacity planning from real volumes.
Document AI
OCR and layout extraction across thousands of supplier formats, with confidence scoring.
AI security
Prompt-injection defences, PII masking and allow-lists that hold up in an audit.
ERP integration
Three-way matching, posting rules and safe integration with finance systems.
Have a process where people spend the day copying data between systems?
Tell us the process, the volumes and the systems involved. We will come back with a plain view of what an agent could safely take on, what should stay with people, and how to prove it before going live.