Everything in this series so far has been about arrangements you run yourself; this piece turns outward, to the numbers that arrive from somewhere else — a vendor’s deck, a provider’s methodology summary, a pilot report from another business unit, a case study in a procurement pack.

The failures people expect from an LLM form a list that’ll likely be familiar to you: a fabricated figure, a citation to something that doesn’t exist, arithmetic that doesn’t hold. Those are tractable, because they announce themselves to anyone who checks. The harder problems are upstream of the model entirely, in a gray space that neither prompt quality nor model capability touches.

Consider what happens when you upload a third-party valuation report and ask for a summary. The model reads it correctly. It summarizes it faithfully. Every figure in the output traces to a figure in the document — and nothing in that exchange establishes that the figures in the document mean what they appear to mean, or that the basis on which they were produced is one you’d have accepted had anyone walked you through it. The output is accurate and the conclusion can still be wrong, because the accuracy is about the document rather than about the world.

That is a gap sheltering a great deal of the risk in AI-assisted finance work, and it doesn’t automatically close with model improvements; it closes by asking where the thing came from.

Below, we’ll walk through four such questions, ranked in rough order of how much impact each will ultimately have on how you move forward.

Who or what served as the ground truth?

This matters most, and it’s often left unstated.

A great many accuracy and precision figures in our AI in finance space are produced along the same lines: an agent processes a set of items, a human reviewer checks the agent’s work, and the reported number is the proportion the reviewer agreed with. That figure is real, it’s reproducible, and it is a measurement of concordance between the agent and the reviewer rather than a measurement of correctness.

Two arguments from earlier in this series make that sharper than it first looks. The approval before action piece described how the content of a review thins under load, from independent re-derivation, to checking for obvious error, to accepting a well-formed output that nothing flags. A reviewer somewhere on the lower rungs of that ladder is a weak answer key, and the resulting number will be high. And the manual capability piece made the point that a check is only worth something if it doesn’t share failure modes with the thing it’s checking. A reviewer who sees the model’s output before forming a view shares a failure mode with it by construction, which is ordinary anchoring and doesn’t necessarily require carelessness on anyone’s part.

So the follow-up questions are: did the reviewer see the model’s output before deciding? Was there any independent source of truth — a confirmed outcome, a downstream correction, a differently-built check? And how were disagreements resolved, since a process that sends disputes back to the model’s own output for adjudication isn’t measuring much.

None of this means concordance figures are worthless, of course, as agreement between a model and a competent independent reviewer is genuinely informative. It just isn’t accuracy, and the gap between the two is one important place where optimism in this category resides. Whether that optimism is misguided or not depends, in the final analysis, on what your additional questions reveal.

What is the denominator?

A rate is a fraction, and the numerator is usually the part that gets reported.

The exception-only escalation piece ran into this directly: an escalation rate tells you how many items went to a person but cannot, on its own, tell you how many should have. The same structure shows up in almost every performance claim. A detection rate needs a population of things there were to detect. A false-positive reduction needs to say what happened to the true positives. A straight-through processing rate needs to say what was excluded from the population before the counting started.

The question to ask is plain: what was the full set of items, and what was removed from it before the number was calculated?

Is this deployed, piloted, or projected — and compared against what?

These two travel together, because the answers tend to come from the same slide.

Deployment status changes what a figure means more than almost anything else. A result from a controlled pilot, on curated data, with the team that built the system paying attention, is a measurement of a best case. That’s a legitimate thing to measure, but it is not a straightforward forecast of production, where the population is messier and nobody is watching that closely. Published material frequently mixes deployed results, active pilots, and explicit future states under a single heading, and the headings rarely announce which is which.

The comparison group is the other half. A performance gain is a difference between two things, and the second thing is often unstated or borrowed. The useful version compares against the prior process at the same firm, on the same population, over a comparable period. The less useful version compares against a different organization’s published case study — sometimes involving different technology entirely — or against a hypothetical manual baseline nobody ever ran. A percentage change calculated on a percentage is worth watching for as well; moving from 33% to 45% is twelve points, and expressing that as a “36% increase” is arithmetically fine, but can be rhetorically misleading if you’re not careful.

QuestionWhat a usable answer sounds likeWhat the answer changes
What was the ground truth?A confirmed outcome, or an independent check built differentlyWhether the number is correctness or agreement
What is the denominator?The full population, with exclusions namedWhether the rate means anything
Deployed, piloted, or projected?A stated production period and volumeWhether it forecasts your experience
Compared against what?The prior process, same firm, same populationWhether the delta is attributable

What do you do when the answers aren’t available?

Usually nothing dramatic – you discount the number and move on.

Most of the time these questions don’t have sinister answers; they have no answers, because nobody was asked. A figure gets produced for one purpose, travels into a deck, and acquires a halo of precision it never had. The person presenting it often doesn’t know how it was derived either, and asking gives them a reason to find out.

It’s worth noticing that this is itself a capability of the kind the previous piece was about. Evaluating a measurement is a skill that atrophies if nobody exercises it, and it atrophies fastest in organizations where the numbers have always arrived pre-packaged. The cost of maintaining it is roughly one conversation per claim.

Hence the diagnostic, which should take all of ten minutes. Take the last AI performance figure anyone showed you — vendor, internal, either — and work out what served as its ground truth. If the answer is a human reviewer looking at the model’s output, you know what the number measures, and it isn’t what the slide says. Tell me in the M.A.S.E. Discord, or get in touch directly. We’re researching how these arrangements behave in production, and what comes back shapes what we write next.

Frequently asked questions

Prove AI is building solutions to power more correct, explainable and auditable AI outcomes.

We’re always interested in learning about AI management challenges.

Get in Touch