Fourteen changes from one pass. Which would you check first?Pick a row — you can change your mind.
Over the past few months, we’ve spoken with a number of financial professionals who’ve made AI a core part of their toolkit, and we’ve noticed that many of them are working in a similar loop. Using a tool like Claude, they go back and forth on a project many times, creating different versions of the final deliverable. Even when a particular version seems fine (i.e., all the numbers reconcile as expected), they usually check it again anyway because it’s hard to tell which changes have already been cleared and which need to be investigated further. As the loop continues and refinements pile up, they gradually lose track of all the little choices made along the way.
When they decide they want to compare two versions, they drag a number onto the active Excel spreadsheet and manually fiddle with it as a way to test their assumptions and find any underlying problems. The hours they saved on producing the actual model have, therefore, gone to vetting it — a process made all the more frustrating by AI’s tendency to hardcode values, rather than offering a formula analysts can test more easily.
AI has done a lot to cut the effort involved in building a financial model, but the time needed to review an AI agent’s output has proven harder to reduce. This dynamic is playing out in many fields where AI adoption has been strong, and there are some early attempts to quantify it. A recent Google survey of more than 600 scientists, for example, found that among respondents who reported saving time with AI, 46% said checking the output consumed more than a quarter of the hours saved.
Why doesn’t a complete record help?
Why not just make a better log; record every change, note what made it, keep all of it, and review when needed? This isn’t as helpful as it might seem at first, and most people really just need to be told where to look.
A list of changes alone doesn’t give you the context you need to allocate your scarce attention effectively. A cell that moved by two dollars and an assumption that moved a valuation by a hundred million sit in the same font, one above the other, and nothing in the list says which deserves your next hour. A hundred unranked changes leave you with two choices: read everything, or eyeball it and make what you hope is an educated guess.
How do practitioners decide where to look?
There are a number of different strategies for making reviews more targeted. On one team, it might be standard practice to look for the assumptions where being wrong by 1% moves the outcome by 50%; on another, the better approach could be to check the Excel figures against the company’s own history, noting when a particular number is, say, only a third of its usual value.
Whatever the case, I want to point out that each approach is a ranking of which issues to examine first and is built by hand by someone with intimate knowledge of the business. This is not nearly as scalable as it’ll need to be to keep up with the increasing volumes processed by AI-powered teams.
What would a ranking need to know?
A useful ranking has to know which figures matter in this model, for this business, this month. Most of that knowledge lives with the reviewer rather than in the file. Can it be written down before the review starts, or is knowing where to look just a part of the job that simply can’t be handed off easily?
As it happens, we’re actively investigating these issues and are iterating toward a solution. If you have any insight, any questions, or just want to see what we’ve been up to, we’d love to hear from you. Reach out via our Contact page, or head straight to our dedicated Discord server to compare notes with our growing community of AI-native builders.
Frequently asked questions
Prove AI is building solutions to power more correct, explainable and auditable AI outcomes.
We’re always interested in learning about AI management challenges.
Get in Touch


