One pair of brackets, four hundred thousand pounds
Everyone examines how clever the agents are. Almost nobody checks whether what went into the tank was petrol.

“There is an entire industry discussing model capability, and the thing that will actually cost you money is a bracket.”
A financial statement shows a figure in brackets. In accounting, that means negative, and every finance professional reads it without thinking. The document goes through an extraction pipeline. The brackets do not survive.
Two hundred thousand becomes two hundred thousand the other way round, and the swing between the two is four hundred thousand pounds. Nothing hallucinated. No model invented anything. A typographic convention was not understood by something that had never needed to understand it, and everything downstream inherited the result as fact.
I like this failure because it is so unglamorous. There is an entire industry discussing model capability, and the thing that will actually cost you money is a bracket.
Your PDF is a photograph#
The first misconception worth clearing is that a PDF is a document. Often it is not. It is an image of a document, a photograph in a costume, no more readable to a machine than a snapshot of your lunch. Somebody scanned it, or generated it from a system that flattened everything into pixels, and the text you can see with your eyes is not text as far as anything downstream is concerned.
So before any of the impressive processing happens, something has to look at those pixels and guess what characters they represent. That guess is where your accuracy is determined, and it happens in a step that nobody demos, budgets for, or asks about in the procurement meeting.
The failure modes are specific and mostly boring. A one becomes a seven. A comma becomes a full stop, which in a currency figure moves a decimal point. A column boundary is misread and two values merge. A footnote marker attaches itself to a number. Brackets vanish. None of these look like errors afterwards, because the result is always a perfectly plausible number.
Nothing downstream will catch it#
In a chain of automated steps, the output of extraction becomes established fact for everything that follows. The second stage does not interrogate the first, because interrogating it was never anyone's design intention, and it has no independent basis on which to disagree.
Take a supply contract where fifteen hundred units is read as fifteen thousand. The procurement step calculates against that. The finance step projects cash flow from the procurement figure. The logistics step plans warehouse capacity. Every one of those calculations is performed correctly. The arithmetic is impeccable. You end up with a tenfold error expressed to two decimal places and presented with total composure, and the only person who could have caught it was the one who no longer sees the source document because the whole point was to stop them having to.
What makes this worse rather than better is how capable these systems have become. A crude system would parrot the wrong number and look odd doing it. A good one contextualises. It notices the figures imply an unusual pattern and writes you a thoughtful paragraph explaining the apparent turnaround, drawing on plausible commercial reasoning. The analysis is genuinely intelligent. It is intelligence applied to a premise that was never true, which is the most expensive kind.
The arithmetic of stacking#
Here is a calculation that ought to be done more often than it is.
A single processing step at 95% accuracy sounds respectable. Most people would sign that off. Stack three such steps in sequence and overall reliability lands around 86%, because the errors compound rather than cancel. Nobody in the approval chain ever saw the number 86, because each step was assessed on its own and each looked acceptable.
And accuracy is not distributed evenly. The material most likely to be misread is dense, numeric, tabular, footnoted content, which is precisely the material that matters most. Prose survives extraction well. A table of figures with alignment carrying meaning does not. Your error rate concentrates exactly where your risk does, which is a rather unfriendly coincidence.
Paranoid validation, applied selectively#
The answer is not to validate everything to the same standard, which is how you make automation cost more than the manual process it replaced.
It is to match the checking to what being wrong actually costs. Five questions get you most of the way there, and they take an afternoon.
What is the worst realistic consequence if this figure is wrong? A misread name in customer correspondence is embarrassing. A misread number in a regulatory filing is a different category of problem entirely.
How long before we would notice? A daily dashboard surfaces oddities within hours. A strategic input might not reveal its error for a year, by which time decisions have been taken and money committed on the strength of it.
What does extra checking cost against the cost of being wrong? Running two independent extraction engines and comparing outputs roughly doubles processing cost. For most correspondence that is absurd. For anything ending up in a board pack it is trivially worth it.
How structurally complex is this document? Plain text needs little. Scanned financial statements with tables, footnotes and mixed formatting need several layers and a low threshold for escalation.
Is there anyone available to adjudicate when confidence is low? If your expert is unavailable for weeks, your thresholds have to be more conservative, because the escalation path is theoretical.
Answer those honestly and your validation budget allocates itself.
What good looks like#
A well-built pipeline does not claim to be right. It tells you how confident it is about specific data points, not just the final answer, and it does so in a way that has been calibrated rather than asserted. It runs a second extraction where the stakes justify it and flags disagreement rather than silently picking a winner. It applies sanity checks against what the numbers should look like, so that a balance sheet which does not balance stops the process rather than proceeding elegantly. Below a threshold, it routes to a person, and someone senior has decided in advance what that threshold is.
None of that is exotic. It is the sort of thing any careful engineer builds when they have been burned once, which is rather the point.
There is nothing quite so dangerous as a very intelligent system that is confidently wrong about the facts, and the facts usually go wrong long before anything intelligent gets involved.
Bring us the problem.
A short, no-obligation call. If we are not the right fit, we will say so and point you somewhere better.