Fourteen changes from one pass. Which would you check first?Pick a row — you can change your mind.

TIME
CELL
CHANGE
EFFECT ON VALUATIONNOT IN THE LOG
09:14:02
B14
Opex growth rate 2.4%→2.6%YOUR PICK
$2.4M
09:14:02
F31
Headcount Q3 142→145YOUR PICK
$310K
09:14:03
C08
Rent escalator 3.0%→3.1%YOUR PICK
$140K
09:14:03
D52
Bad debt provision 0.8%→0.9%YOUR PICK
$420K
09:14:04
H19
Payroll tax rate 7.65%→7.80%YOUR PICK
$95K
09:14:04
B27
Rounding adj. $148→$150YOUR PICK
$2
09:14:05
G44
Terminal growth 2.0%→3.5%YOUR PICK
$104M
09:14:05
E12
COGS ratio 61.2%→61.4%YOUR PICK
$1.9M
09:14:06
J07
FX rate EUR 1.08→1.09YOUR PICK
$260K
09:14:06
C33
Churn assumption 1.8%→1.9%YOUR PICK
$3.1M
09:14:07
D18
Capex timing Q2→Q3YOUR PICK
$180K
09:14:07
F60
Deferred rev. $4.1M→$4.3MYOUR PICK
$540K
09:14:08
B41
Discount rate 9.0%→9.1%YOUR PICK
$6.8M
09:14:08
H52
Inventory days 34→35YOUR PICK
$90K

Over the past few months, we’ve spoken with a number of financial professionals who’ve made AI a core part of their toolkit, and we’ve noticed that many of them are working in a similar loop. Using a tool like Claude, they go back and forth on a project many times, creating different versions of the final deliverable. Even when a particular version seems fine (i.e., all the numbers reconcile as expected), they usually check it again anyway because it’s hard to tell which changes have already been cleared and which need to be investigated further. As the loop continues and refinements pile up, they gradually lose track of all the little choices made along the way.

When they decide they want to compare two versions, they drag a number onto the active Excel spreadsheet and manually fiddle with it as a way to test their assumptions and find any underlying problems. The hours they saved on producing the actual model have, therefore, gone to vetting it — a process made all the more frustrating by AI’s tendency to hardcode values, rather than offering a formula analysts can test more easily.

AI has done a lot to cut the effort involved in building a financial model, but the time needed to review an AI agent’s output has proven harder to reduce. This dynamic is playing out in many fields where AI adoption has been strong, and there are some early attempts to quantify it. A recent Google survey of more than 600 scientists, for example, found that among respondents who reported saving time with AI, 46% said checking the output consumed more than a quarter of the hours saved.

Why doesn’t a complete record help?

Why not just make a better log; record every change, note what made it, keep all of it, and review when needed? This isn’t as helpful as it might seem at first, and most people really just need to be told where to look.

A list of changes alone doesn’t give you the context you need to allocate your scarce attention effectively. A cell that moved by two dollars and an assumption that moved a valuation by a hundred million sit in the same font, one above the other, and nothing in the list says which deserves your next hour. A hundred unranked changes leave you with two choices: read everything, or eyeball it and make what you hope is an educated guess.

How do practitioners decide where to look?

There are a number of different strategies for making reviews more targeted. On one team, it might be standard practice to look for the assumptions where being wrong by 1% moves the outcome by 50%; on another, the better approach could be to check the Excel figures against the company’s own history, noting when a particular number is, say, only a third of its usual value.

Whatever the case, I want to point out that each approach is a ranking of which issues to examine first and is built by hand by someone with intimate knowledge of the business. This is not nearly as scalable as it’ll need to be to keep up with the increasing volumes processed by AI-powered teams.

What would a ranking need to know?

A useful ranking has to know which figures matter in this model, for this business, this month. Most of that knowledge lives with the reviewer rather than in the file. Can it be written down before the review starts, or is knowing where to look just a part of the job that simply can’t be handed off easily?

As it happens, we’re actively investigating these issues and are iterating toward a solution. If you have any insight, any questions, or just want to see what we’ve been up to, we’d love to hear from you. Reach out via our Contact page, or head straight to our dedicated Discord server to compare notes with our growing community of AI-native builders.

Frequently asked questions

Prove AI is building solutions to power more correct, explainable and auditable AI outcomes.

We’re always interested in learning about AI management challenges.

Get in Touch