Financial departments are among those adopting AI with the most alacrity, which is why we at Prove AI have been looking closely at how their work actually gets done. Our verdict is that, for all of the (mostly justified) hype, it’s clear that financial professionals still cannot simply trust AI outputs. It may save you time in extraction, summarization, and first drafts over document-heavy work, but human intervention is more critical than it’s ever been to ensuring you’re avoiding potentially massive and costly mistakes. This is why everyone continues to hand-check just about every financial number proffered by an LLM. The motivation behind their expending this effort is that the two primary failure modes encountered when using AI in finance — both taken apart below — occur when an automated workflow produces output that looks, for all the world, to be correct.

Calculating interest coverage, and what it tells us about AI in finance

Consider the example of interest coverage, which is a one-line check on whether a company earns enough to service its debt and is calculated as earnings before interest and taxes (EBIT) divided by interest expense. If you ask an AI tool to pull the inputs from a filing and compute it, you may well end up with an output (4.2x, for instance) that is clean, formatted, plausible — and wrong, because the interest-expense figure required wasn’t actually in the document. In such a scenario, the tool will fill the gap with a completely fictitious number that reads like the real one. Worse, you can drop that ratio into an Excel model, and every downstream cell reconciles, nothing errors, the leverage view looks fine. This will sound familiar if you’ve read our piece on clean traces and wrong outputs on the Prove AI blog.

Catching such a problem requires you to either already know the answer or to make the long trek through the existing source material. This is why the relevant experts default to checking every number, every time, but it’s also an issue, inasmuch as the point of adopting modern tools — AI workflows, multi-agent systems, what have you — is to obviate the need for all this tedium.

Though a general solution to making AI immune to these failure modes is not forthcoming, the sections below will discuss the most common sorts of tasks tackled with AI in finance before finishing with a treatment of two of the main problems encountered by those engaged in such efforts.

What do finance teams actually use AI for today?

By and large, the answer is the same for finance as it is for many of the rest of us. AI is great at narrow, document-shaped work, e.g., pulling figures from filings or similar documents, summarizing earnings calls, drafting memo sections, matching transactions. This is the quotidian reality behind the stratospheric adoption numbers we all read about every day, and it’s also where the firsthand accounts turn positive — the tools earn their keep on volume and drudgery, less so on tasks requiring real judgment.

There are three terms worth pinning down, since everything below depends on them:

  • Extraction is pulling a specific figure or fact out of a document — the number was there, and the tool found it.
  • Grounding is tying the output to that source, so the model answers from the filing in front of it rather than from its training.
  • Hallucination is the failure that shadows both: the model producing something fluent and plausible that was never in the source and isn’t true.

In document-heavy finance work, extraction is the job, grounding is supposed to be the safeguard, and hallucination is what the safeguard is for.

Though there are many ways in which AI in finance is playing a part, the following six workflows provide a high-level illustration of a pattern of adoption for each of the main financial seats.

Financial statement analysis and valuation

Here, AI extracts and normalizes filings well (pulling figures, tagging line items, flagging restatements) and drafts passable first-pass ratio commentary. It’s less helpful for tasks which turn on competitive judgment and reading where the cycle is; it will hand you a plausible growth rate, and knowing it’s wrong is the analyst’s edge.

Budgeting and forecasting

Variance analysis and the first-draft narrative are a good use case, such as if you ask a model to “flag every line more than 10% off plan and draft the explanation”, as is baseline statistical forecasts; what a model can’t read is the political dynamics of a department or what the CFO happens to care about this quarter.

Client planning and advisory

The analysis and scenario side scales well (Monte Carlo runs, tax-loss-harvesting candidates, allocation-drift alerts), and drafting the plan document is fine; it shouldn’t be client-facing, because, there, the value is in personal relationships.

Month-end close and reconciliation

This is among the strongest fits on the list: matching transactions, flagging unreconciled items, and clustering anomalies is something machine learning does well, and much of the market already runs rules-based automation here; the accrual judgment and the audit trail still need a named human, because auditors want accountability.

Deal diligence and credit

Performing document review across a broad spectrum of data is a headline use, surfacing every change-of-control clause and flagging inconsistencies across hundreds of contracts; what tools still miss is the kind of synthesis a skilled human excels at, such as noticing that three weak signals together indicate that a CFO is hiding a liquidity problem, and the go/no-go is a call someone made of Carbon still has to own.

Compliance and reporting

Monitoring and first-pass flagging at scale fits, and LLMs add real value reading the unstructured content (adverse-media screening) that older rules-based systems choke on; the pain is false-positive management, where more flags without more precision make the job worse, and a filing decision can’t be handed to a black box no regulator can hold accountable.

The consistent line across all six is that AI is strongest at extraction, matching, and first-draft synthesis over large document and data volumes, and weakest wherever judgment, accountability, or client trust become the deciding factors. The friction slowing adoption is rarely whether the model can do the task, turning instead on who’s accountable when it’s wrong, in contexts that are famously regulated, audited, and fiduciary by default. That quandary is brought into sharper relief when you consider that the two most common failures don’t look like failures at all — which is where we’ll head next.

Why is it so difficult to catch a number that an LLM got wrong?

A fabricated figure looks exactly like a retrieved one. There’s no marker at the point of use — no color, no flag, no dip in the model’s confidence — that separates the number the tool pulled from the filing from the number it just straight-up invented. Both arrive formatted to two decimals, both slot into the cell, both tie out. The reason nobody catches the wrong number is that nothing about it looks wrong.

This is a consistent theme in how finance professionals describe the risk, and it’s worth separating from the generic worry that “AI hallucinates,” which, in fairness, is a concern most of us are already aware of. The complaint isn’t that the tools are wrong sometimes; every tool is wrong sometimes. The complaint is where the wrongness crops up. When WallStreetPrep ran four tools — ChatGPT, Claude, Copilot, and a purpose-built Excel tool, Shortcut — through the assignment they give analyst trainees, a full three-statement model for Apple built from real filings and graded on the same rubric, even the strongest tool came in below a junior analyst. But the ranking wasn’t the alarming part. The two models that performed best also fabricated portions of the historical data, and the fabrications were subtle in a specific way: individual line items were slightly off, yet they still summed to the correct subtotals. Because the model reconciled everything, an auditor’s eye running down the totals would have found nothing obvious to stop on.

That’s the actual failure mode, and it’s why it earns its own section. Unlike a software bug, a hallucination doesn’t announce itself. A software bug fails loudly, which — the hue and cry of generations of software engineers notwithstanding — is a kind of mercy: the stack trace tells you where your investigation must commence. A fabricated number fails silently, and in a domain where a single figure can be material to a disclosure or a decision, ‘silent’ is the worst way to fail. A made-up interest-coverage ratio of 4.2x reads exactly like a real one. The problem financial professionals describe isn’t error in the abstract; it’s invisible error in a domain with precious little tolerance for it.

The mechanism is worth understanding, because it explains why the errors cluster where they do. An LLM generates text by predicting plausible continuations from patterns in its training data; left to its own devices, it isn’t looking a figure up in a system of record and copying it across. When the underlying data is thin — such as might be the case in a low-coverage private company, a historical period the model has seen little of, a filing it wasn’t actually given — the gap gets filled from the model’s sense of what a number like that usually looks like. The output is a plausible value, not a retrieved one.

And the risk compounds in derived metrics, because a ratio that combines several inputs can look entirely reasonable while resting on one invented figure three steps back; the arithmetic often actually is right, so the result inherits an authority to which the inputs have no claim. Precision, in other words, is a formatting property here, not an epistemic one. To a human reader, two decimal places signal care, but they mean nothing to the matrices of floating-point numbers churning inside an AI agent.

Why can’t an AI agent just get the right number?

The previous section assumed the model had the data and mangled it. The problem identified here is upstream of that: often the model never had the number to begin with, but instead of telling you that, it returns one anyway. Unless an automated workflow is instrumented very carefully, with all the proper AI observability, a retrieval miss and a retrieval hit come back looking identical — same confidence, same formatting, same absence of any note that the figure was reconstructed from memory rather than pulled from a source. Though this gets conflated with the fabrication problem, it’s a distinct failure with a distinct cause, and confusing the two leads teams to reach for the wrong fix.

Practitioners describe hitting a wall that has nothing to do with the model’s reasoning. On the Wall Street Oasis banking forum, the working consensus is that these tools save real time on repetitive work — research, drafting, summarizing, pulling data for decks and comps — but aren’t trusted for full models or CIMs; and the reflex that comes through most clearly is one of verification. Now, the habit of checking every number long predates AI, and characterizes users of many offerings. What’s new, however, is that we now have a tool that looks more authoritative than a data terminal while being less connected to a system of record than one.

This is more of a data-access problem than an intelligence problem. Better prompting doesn’t grant the model access to data it doesn’t have. When you ask a base model for a company’s valuation multiple and it isn’t connected to a financial database, it doesn’t look the figure up — it generates a plausible one from patterns in its training data, the same mechanism behind the fabrication in the previous section, now operating on a number that was never retrievable in the first place.

The instinct is to reach for a workaround, but a system-prompt instruction to “flag when you’re unsure” produces a caveat, not a retrieval; the number underneath is still generated. Web lookup is the workaround that sometimes provides a fix, but it can’t always sustain a query across multiple periods or screen across a universe of companies unless it is exceptionally well-engineered.

What changes when an agent reads the output instead of a person?

Both problems so far assume a person is reading the numbers. The whole defense — check every figure, catch the implausible accrual, verify against source — runs on human eyes at the point of use.

The thing is, when you’re managing agents rather than building the workflow yourself, the wrong number isn’t merely something you read. It’s an input to the next agent’s step. The fabricated interest-coverage figure doesn’t helpfully loiter on a screen waiting for someone to squint at and give the ’ol thumbs up or down; it flows into a covenant check, a screen, a summary that feeds the next call. The checking layer that catches these failures today resides where the concerned human sits — and in an agent-run pipeline, that seat may be empty at the moment the number is used.

Whether errors compound across agent steps, or get caught by some other mechanism, or turn out to matter less than they look is not a question we have sufficient evidence to settle here. The point is narrower: the failure modes don’t change in an agent workflow, but the thing that catches them might not be there.

Where this leaves us

The discipline that makes AI usable in finance today is human verification — someone looks at every number, because both dominant failure modes produce output that looks correct and which can only be reliably caught by a skilled and motivated human. That works, albeit slowly, and it assumes a reader.

What we don’t quite know yet is how this plays out in the context of automated workflows powered by AI agents. If you’re running finance workflows with AI in the loop, where’s the checking actually happening right now — and have you started handing any of it to the tools themselves? We’d love to hear about your experience in this exciting, fast-evolving domain, so please reach out to chat with us!

Frequently asked questions

Prove AI is building solutions to power more correct, explainable and auditable AI outcomes.

We’re always interested in learning about AI management challenges.

Get in Touch