Because finance is a domain in which the stakes tend to be very high, and because financial teams have adopted AI with such alacrity, there is a strong motivation to have a human being present at some point in the workflow to monitor progress, check for inaccuracies, nudge agents out of unproductive states, and so on. This is commonly known as having a “human in the loop,” but, across this series, we’ve been arguing that at least four distinct arrangements ship under that single label:
- approval before action;
- exception-only escalation;
- spot-check sampling;
- post-hoc confirmation.
Last time we looked at spot-check sampling; this is an arrangement that is honest about its coverage, but honesty, as it turns out, is not the same thing as understanding. Here we’ll finish the set with post-hoc confirmation, in which the output gets used and only checked afterward. Of the four, this is the one whose name does the most work to obscure what’s happening: the phrase “in the loop” suggests a person sits inside the process, but the arrangement we’re designating as “post-hoc confirmation” has the person sitting downstream of it, retrospectively examining something that has already gone out.
Third-party valuation flowing into financial reporting is a good place to illustrate this pattern. Imagine a firm whipping up a report, but bumping up against an instrument such as an illiquid credit position or a private holding which it is unable to price on its own. In such a case, the firm might reach out to a third-party pricing provider or valuation specialist for the missing data. Given the technological changes that have swept the world over the past few years, we can be more or less certain that some part of that provider’s process is AI-assisted. So, the pricing information (the ‘mark’) arrives, gets posted, and moves. Somebody reviews the methodology and challenges the assumptions later, on a cycle that was set for operational reasons and rarely aligns with when the number started mattering.
Below we’ll work through what confirmation actually costs, why its depth tends to fall exactly as the stakes rise, what the coverage ratio needs in order to mean anything under this arrangement, and why the evidence question here opens onto a larger problem we’ll take up separately.
What does post-hoc confirmation cost to run?
It costs whatever reconstruction costs, and reconstruction gets more expensive with elapsed time and with every system the number has passed through.
This stems from the fact that confirming a valuation requires more than simply reading it; it means rebuilding the state of the world at the moment the mark was struck: which inputs were used, which comparables were in the set and why, what market data was current on the valuation date, which version of the provider’s model produced it, and what assumptions were operative at the time. Each of those shards of context is retrievable early and becomes progressively harder to retrieve as time passes. To take but a small sample of possibilities, consider that the snapshots of market data may not be retained in the form they were consumed, the provider may have updated its methodology, or the analyst who could explain the comparable selection may have moved on.
So, this is one of the two main drivers of confirmation costs under a post-hoc confirmation arrangement. Another is the number of systems the mark has traversed, and this one compounds steadily. Though the details obviously vary from one firm to the next, under a standard setup, a mark may arrive with a methodology memo attached, then post to the relevant subledger as a value with a source code, then be aggregated into a net asset value, then appear in a schedule, then a footnote, then an investor report and a fee calculation. Every hop preserves the number while letting more of the contextual penumbra you would use to evaluate it evaporate. By the fourth artifact it is a figure with a provenance that exists somewhere, in principle, in a system somebody else administers.
This helps to explain a finding that otherwise looks like ordinary inefficiency. KPMG’s 2026 global survey of finance leaders reports that 42% describe themselves as strongly assurance-ready — able to produce AI-related audit evidence efficiently and without disruption. The phrase doing the work there is without disruption, and reconstruction would surely qualify as a noteworthy disruption.
Why can confirmation depth fall as the number becomes more important?
Because the cost of acting on an adverse finding rises faster than the cost of performing the confirmation, and the person confirming knows it.
There may or may not be cases in which someone explicitly instructs a reviewer to go easy on a mark that’s already in a published NAV; what’s important for our purposes is that, in many cases, it isn’t necessary for anyone to do so, because there’s already an incentive in place pushing in that direction. Consider what changing the number costs at four points in its life:
| Where the number has reached | What confirming it costs | What changing it costs | What the confirmer is actually deciding |
|---|---|---|---|
| Still in the provider’s file | Low; inputs are current and retrievable | Low; nothing downstream depends on it | Whether the methodology is sound |
| Posted to the subledger | Moderate; some context already stripped | Moderate; a journal entry and an explanation | Whether the methodology is sound |
| Rolled into a reported NAV | High; reconstruction across systems | High; a restatement and a set of conversations | Whether it is wrong enough to unwind |
| Published, and used for fees or covenants | Highest; provenance is distributed | Highest; external parties are affected | Whether it is wrong enough to unwind |
The right-hand column is critical to the whole argument. Somewhere between rows two and three the question the confirmer is answering subtly changes. Early on, they’re assessing whether the number was derived properly. Later, they’re assessing whether it’s wrong enough to justify the steps that correcting it will set in motion — and that’s a different question with a different threshold, asked by the same person, under the same procedure, producing the same artifact.
Two further things push depth down as the number ages. The reconstruction cost above means a thorough confirmation of an old mark takes materially more work than a thorough confirmation of a fresh one, so the same budgeted hours buy less scrutiny. By then, the mark has been implicitly ratified, as it survived the close before subsequently appearing in a report that went unchallenged. Epistemologically, none of this is evidence about the veracity of the valuation, but observe that all of it raises the bar for overturning it later, and that’s going to have an impact on whether anyone wants to expend the effort of doing so down the line.
What does the coverage ratio measure here?
The same thing it measures everywhere — the proportion of items a person reviewed — but under this arrangement the figure needs a clock attached before it says anything useful.
For the other three arrangements, the review and the use are simultaneous, or the review precedes the use:
- An approval happens before the item lands;
- An escalation is worked before the reconciliation closes;
- A sampled invoice is checked as part of the processing cycle.
In each case, a coverage ratio is a complete statement, because there’s no meaningful gap between reviewing an item and relying on it.
Post-hoc confirmation separates those two events in time, which means a bare ratio conflates situations that aren’t alike. Two firms can each confirm 100% of third-party marks and be running materially different controls: one confirms before the mark enters the reporting chain, the other confirms in the quarter after it was published and used to strike fees. Though the two counts may be identical, their values are starkly different, as they are describing antipodal approaches.
So the useful version of the number here is not simply “raw confirmations completed,” it is “confirmations completed before the mark became load-bearing somewhere else.”
What evidence survives, and where does it live?
The confirmation generally produces a decent record. The problem is that the record and the number end up in different places.
Here the arrangement runs into something larger than itself. A serious confirmation generates real evidence — who reviewed the methodology, against what, what they found, when. But the mark lives in the firm’s general ledger and reporting stack, while the evidence about how it was derived lives in the provider’s platform, and the evidence about how it was confirmed may live in a working paper, a spreadsheet, or an email thread. The system that is answerable for the number is not the system that holds the reasoning behind it.
KPMG’s report includes a case about a bank whose financial reporting relied on AI-enabled valuations from a third-party provider, operating without a defined assurance path over how those valuations were produced. This is framed as assurance readiness moving from “an internal discipline” into “a commercial requirement,” something counterparties ask about rather than something a firm simply decides to maintain on its own initiative.
This isn’t specific to valuation. Any time an AI-assisted output crosses an organizational boundary, a set of fields (what produced it, on what inputs, under what version, reviewed by whom, and when relative to use) either travels with it or doesn’t. Post-hoc confirmation exposes the seam more sharply than the other three arrangements because the time gap gives the context every opportunity to fall away in transit. What a receiving system would need to know about an input it didn’t produce is a question big enough to take on its own, and it’s one we’ll address in a future installment.
For now, the question I’d put to anyone relying on third-party marks (or doing something structurally similar) is: for last quarter’s valuations, what was the median gap between the date a number was first used downstream and the date somebody confirmed it — and on the second date, what would changing it have cost? Tell me in the M.A.S.E. Discord, or get in touch directly. We’re researching how these arrangements behave in production, and what comes back shapes what we write next.
Frequently asked questions
Prove AI is building solutions to power more correct, explainable and auditable AI outcomes.
We’re always interested in learning about AI management challenges.
Get in Touch


