In plenty of finance workflows, an AI agent is now responsible for producing the first pass of a forecast, a budget, or a scenario, often in just a few minutes. Rather than disappearing altogether, however, the hours that used to go into building have relocated into a much-expanded review step that got harder while nobody was measuring it. The errors that matter here are rarely the actual arithmetic; the trouble sits in the spaces around the formulas and their outputs, such as a vendor export where the AI model read an annual column as monthly, an FX rate pulled from the wrong date, a headcount file whose “active” flag means something different from what the prompt assumed, and similar such issues.
None of these are a prompt problem or a math problem. Each one is a question about whether the data going in is what it claims to be, and answering it to everyone’s satisfaction seems to take roughly as much time as building used to.
What got compressed, and what didn’t?
By and large, what got compressed was ‘structural work,’ things like constructing a formula, setting up basic scaffolding for different finance scenarios, and writing the first draft of variance commentary. Work that used to fill an afternoon in Excel can now be a prompt and a short wait.
What didn’t compress is knowing whether any of it is right. Nothing speeds up checking that a SUMIFS range stops at the last row rather than one short, that a hardcoded growth rate was meant to be hardcoded, or that the lookup against the chart of accounts matched on the code and not the description. That work runs at the same speed it always did, and it can run slower, for reasons we’ll elucidate shortly.
There’s also another, more subtle expansion. When scenarios are cheap to build, more of them get requested. A planning cycle that used to carry a base case and a downside can now carry five cases, because asking for three more costs almost nothing. Having said that, each one still has to be defensible to whoever eventually has questions about it. The building cost fell and the review cost multiplied, so the hours were conserved, just in a different column.
Why did the reviewer’s cheapest step disappear?
Reviewing work you didn’t do is a different job from reviewing work you did, and the difference comes down to one step.
In the traditional loop, the cheapest move a reviewer has is asking the builder. Why is this cell hardcoded? Where did the churn assumption come from? Why does your analysis omit the month of March? A thirty-second answer (“the client gave us that number on a call,” “March had the one-time credit, so I pulled it out”) replaces twenty minutes of tracing precedents. That answer is cheap because the builder made a decision and can recall it.
When an AI model built the workbook, that person doesn’t exist. You can still ask the question of the model, and you’ll get a fluent answer. But the answer is a fresh generation rather than a recollection. It’s a plausible reconstruction of why someone might have made that choice, and it’s one more output that needs checking. The cheapest step in review has become another thing taking up space in the review queue.
This hits hardest where the review loop has already collapsed, where the person who prepares the numbers is also the person who signs off on them. Self-review used to carry a built-in advantage: you were the builder, so you knew where you’d been careful and where you’d guessed. When the AI model does the building, that advantage goes away even for someone working alone.
What happens under pressure, and why can’t throughput tell the difference?
When review time runs out before the deadline does, there seem to be two broad strategies for absorbing the squeeze.
The first is accepting more. You review at the summary level, spot-check the totals, confirm that the balance sheet balances, then give everything else a “this seems right” pass. The second is delegating less. You take the hard parts back, rebuild the tabs where the logic matters, and keep the AI model on formatting and first-draft commentary. Both are reasonable, and experienced people choose each of them for good reasons.
From the outside, they might look identical because, either way, the package goes out on time, and the board deck has the same number of slides. Throughput measures what left the building; it can’t see what was checked before it left, or how thorough those checks were. A week spent accepting more and a week spent delegating less can produce the same count of shipped deliverables, and nothing in the output records which approach was taken.
In practice, I suspect most people don’t choose once. They do both, tab by tab, without deciding, and the mix drifts with the calendar. Close week pushes toward accepting more; a quiet week in the middle of the month pushes back. That’s how “the AI model saves me time” and “I’m working more hours than I used to” can both be true of the same person in the same month, and it’s why the feeling of saving time may be a poor way to measure it.
There is some mixed evidence of this in the domain of software (though that evidence is not without its problems), and I’m here inferring that the same gap between “felt time” and “measured time” carries over into finance work: the building gets faster in ways you notice, and the review expands in ways you don’t. We’ll have to wait for more rigorous studies to be sure, but this broadly accords with my personal experience, and with testimonies offered during Prove AI’s interviews with various financial professionals.
(If you’re one of these professionals and you’re open to speaking with us, please reach out here; we’d love to speak with you!)
What would you need to see to know which one you’re doing?
Take the last workbook you sent to someone who was going to question it. Could you say which numbers you traced back to a source and which ones you accepted because they tied? Could you say where your review hours went: which tabs got scrutiny and which got a glance? And could you say whether that mix looks different this quarter from last?
Most of us could answer from memory, if at all, and memory is exactly what the missing builder took with them. The question from the lender, the auditor, or the board tends to arrive weeks after the review happened, when the details of what got checked have faded.
We don’t think anyone has arrived at a satisfactory solution for this state of affairs, but we do see the problem, we’re asking the same questions you are, and we’re building against it. What I’d like to know is whether the picture above matches your week. If you’ve found a way to see which response you’re actually running, I’d like to hear how. And if you think the time went somewhere I haven’t looked, I’d like to hear that even more. You can reach out via the Prove AI contact page, of course, but you’re also welcome to join us in the Multi-Agent Systems Engineering (M.A.S.E.) Discord server, where we gather first-rate practitioners across many different domains to compare notes and stay abreast of new AI developments.
See you there!
Frequently asked questions
Prove AI is building solutions to power more correct, explainable and auditable AI outcomes.
We’re always interested in learning about AI management challenges.
Get in Touch


