Every cost we mapped in part two eventually comes to rest somewhere specific. Verification and documentation, for example, land on a person whose name goes on the output, and inside systems that were built to hold financial data with no obvious place to put anything else.

Importantly, little about this procedure improves when the AI model does. That’s the last thing this arc has to say about why capable tools stall in finance workflows:

  1. The tool produces;
  2. A person attests;
  3. The systems remember, or they don’t.

On its own, higher output does nothing for the second and third steps; if anything, it probably makes them more challenging, because the faster the drafting runs, the more there is to attest to.

We’ve gotten a lot of mileage out of our credit memo example, so let’s return to it once more. Ninety seconds to draft, hours to verify, and now it lands on the desk of the credit officer whose signature turns it from a document into a decision. What, exactly, is she signing?

Who actually signs off on AI in finance work?

A person does. It’s always a person signing, and that’s a structural fact underlying everything else in this piece.

It’s an easy fact to lose sight of, because much of the conversation about AI in finance is framed around whether a workflow is compliant — as though compliance were a property of a process, something you could certify once and then run at volume. It usually isn’t. Audit documentation standards require the file to show who performed the work and who reviewed it, by name and by date, and (barring obvious improprieties) they are entirely uninterested in how the work got produced. What’s more, these standards also generally stipulate that you must submit additional documentation when uploading a report, verifying that it’s been audited by a trusted source. The same basic shape holds elsewhere in finance, and I’m not personally aware of a frame in scope here that contains a provision varying with the production method.

A signature today, therefore, means roughly what it has always meant, whether the draft took ninety seconds or three days:

  • The inputs are what they’re represented to be;
  • The conclusions follow from them;
  • I checked, and I stand behind it.

Which is why human-in-the-loop is still a requirement. In practice, ‘human-in-the-loop’ often describes an arrangement where a person is present in the workflow: the draft arrives, they read it, nothing jumps out, they click approve, and the item moves. There are many circumstances in which that is satisfactory to everyone concerned, but that’s much less the case in finance, because the obligation is attestation, which, ultimately, is a claim about work that someone actually did.

There’s an uncomfortable corollary here from part two. The checking a signature attests to is precisely the activity the AI model has been observed to work against: push back on a synthesized conclusion and what tends to come back is a longer, better-organized version of the same conclusion rather than a disclosure of the weakness. The signer is attesting to having successfully completed something the tool was, in a small and entirely unmalicious way, making harder. The reviewer who walks away more convinced rather than better informed has still signed.

If your organization has experience with the challenges of navigating this kind of agent-powered, human-centric work, I’d like to hear it — the M.A.S.E. Discord is the place to tell me, and I suspect a few people there would be as interested as I am.

What does signing actually require you to have done?

Less than some fear, and rather more than most workflows currently produce.

It doesn’t require understanding the AI model. To my knowledge, no one is ever asked to explain the internal behavior of a calculation engine, and the officer signing a credit memo has never been expected to account for how a spreadsheet’s solver arrived at a rate. The standards ask for something much more ordinary — the evidentiary substance of professional work, which is roughly what we covered above, i.e., that the inputs were what they claim to be, the conclusions follow from those inputs, that someone with the standing to judge actually judged, and that all of the above is recoverable later by a person starting from a blank slate.

The credit memo reviewer’s hours produced every one of those — each is a defensible act of professional judgment, and any of them would look perfectly at home in a workpaper.

The problem, as part two argued, is that it won’t survive in a form anyone else can see unless special pains are taken to capture everything. It is to a fuller discussion of this dynamic that we turn in the next section.

Where does the record of AI-assisted work actually live?

Usually nowhere, and the reason is less about integration than is commonly assumed. It’s tempting to think of this as a matter of plumbing — the tool doesn’t talk to the system of record, so someone should roll up their sleeves to write a connector. That’s rarely the real obstacle. The obstacle is that the system of record has no field shaped like the thing you’re trying to store. This matters today, but it’s only going to get more important as more rigorous AI standards proliferate (which is already happening in the EU and can be expected to occur elsewhere). Teams leveraging AI in finance should start thinking now about how to handle this issue as they roll out more powerful automated workflows.

To illustrate, take one of the most regulated corners of the credit workflow: the adverse action notice. When a lender declines an application for a loan, the borrower has to be told the principal reasons, and Regulation B provides sample forms (like this one) carrying a fixed checklist of them. A loan origination system implements that checklist the obvious way, as a dropdown with items like ‘insufficient income,’ ‘length of employment,’ ‘delinquent past or present credit obligations,’ etc. The analyst picks one, the letter generates, the file closes.

Then the underwriting starts leaning on AI models nobody can summarize in a phrase, and the checklist stops fitting. The CFPB directly addressed the question of whether creditors using artificial intelligence or complex credit models may fall back on the sample checklist when those listed reasons don’t accurately capture what actually drove the decision. The answer is no: the listed reasons don’t satisfy ECOA if they fail to specifically and accurately indicate the principal reasons for the adverse action. A creditor whose real reason isn’t on the list has to modify the form or check “other” and write the explanation in, and the Bureau was explicit that selecting the closest available factor, when it isn’t the actual one, doesn’t work.

It’s worth pausing on this for a moment, because it generalizes well beyond lending; consider that:

  • The field exists;
  • The field is an enumeration;
  • The true value isn’t in the enumeration;
  • The escape hatch is a free-text box.

While there is somewhere to put the real reason a loan was declined or a person’s credit was negatively impacted, it’s an unstructured text field, at the end of a form, filled in by whoever is closing the file. The regulator has correctly insisted that the specific reasoning be recorded, and the system has correctly provided a place to record it, and the result is that the most important sentence in the file sits in a field unlikely to be queried.

Now, adverse action codes are a particularly vivid case, but I’d argue the pattern applies more broadly across financial systems, and for an entirely sensible historical reason: the relevant platforms were designed to record outcomes, but they’re less well-equipped to record derivation.

AI-assisted work makes that arrangement considerably more complicated. It is, of course, possible to mechanically capture which documents were in scope, which passages the output rested on, which version of the tool and which instructions produced it, what the reviewer changed, etc., because all of that exists at the moment the work happens. The problem is that almost none of it has a standard place to go.

And evidence that lives outside the system of record doesn’t function as evidence. It might be perfectly true that your team ran a careful review; if that review exists in a chat transcript, a Slack thread, and a reviewer’s memory, then six months later, under examination, it exists nowhere the examiner can reach. Part two argued the checking mostly doesn’t get recorded. What we’re talking about here is the more specific and more stubborn version of that problem: even where a team wants to record it, the schema tends not to have a column for it.

What happens when the AI agents do more of the work?

Everything above assumes a person producing a draft with an AI agent or a similar tool. That assumption is getting shakier, and it’s worth saying what does and doesn’t change when more of the production runs as part of an autonomous workflow.

One of the biggest changes is throughput, inasmuch as there will be much, much more of it. What doesn’t change, at least as the rules are currently written, is the signature. Every obligation this piece has touched assigns responsibility to a person or a named function: the auditor who performed the work, the reviewer who reviewed it, the creditor who states the principal reasons.

So the arithmetic works out roughly along these lines:

  • Production cost falls as more of the drafting is automated;
  • Verification cost stays per-item, because each item rests on different documents and different judgments;
  • The attestation requirement stays per-item as well, because someone still has to be able to say they checked;
  • The volume arriving at that person’s desk goes up.

Which gives a direction without giving a magnitude. I suspect that the gap between what an organization can produce and what it can take responsibility for widens rather than closes, particularly in those parts of the work where the tools are working best.

What this series has actually been about

We’ve covered a bevy of problems across these three posts, but we can think of them as being one singular problem whose facets are observed from different angles.

We have output that’s fabricated but well-formatted, output that quietly missed a retrieval, output that’s expensive to check, checking that leaves no trace, a signature that means more than a click, and systems requiring you to produce reasons for a decision with no good way to record the evidence. In every case, the underlying commonality is that the tool is faster than the organization’s ability to absorb what it produced — not just faster than the analyst, but faster than the accountability. The bottleneck, then, is all that must occur before someone can be answerable for the result.

That’s also why the optimistic version of this isn’t wrong, just early.

These are logging problems, schema problems, and workflow-design problems, which is a much better class of problem to have than, say, an interpretability problem. They’re simply not solved yet in most of the places where AI in finance is already in production.

This arc was one pass through the workflow — where these tools earn their place, where they don’t, and what stands athwart our attempts to close the gap. Obviously, there are open questions in every direction: whether the checking discipline holds as volume rises, where the evidence ends up living, what happens at the seams, when an AI-assisted output from one system becomes an input to another that has no idea where it came from, and so on. These and related issues are ones we’ll be writing about in the future.

So, a question, and I’m asking it because the answers are genuinely useful to us: when your team verifies AI-assisted work, where does that verification actually live — and if an examiner asked next year, what would you show them? Come tell me in the M.A.S.E. Discord, or reach out directly to let the team know. I’d rather hear how you’re handling it than guess.

Frequently asked questions

Prove AI is building solutions to power more correct, explainable and auditable AI outcomes.

We’re always interested in learning about AI management challenges.

Get in Touch