In our earlier piece on four different human-in-the-loop arrangements, we made the case that the “human-in-the-loop” label is rather more expansive than it initially appears, subsuming at least four different patterns of activity:

  • approval before action;
  • exception-only escalation;
  • spot-check sampling;
  • post-hoc confirmation.

Here, we’ll undertake a more sustained study of “approval before action.” As you’ll recall, this is the human-in-the-loop arrangement wherein a person handles every item, with the AI agent relegated to drafting the narrative and assembling the context, and nothing closing until a person accepts it.

This most resembles how things operated prior to the rise of AI in finance — same queue, same reviewers, same signature at the end — which makes it the easiest to explain to an examiner and the easiest to hollow out without a single document changing.

What does approval-before-action actually cost to run?

Because the model compresses the assembly process but does little to compress the decision process, approval-before-action will usually cost much the same as it did before.

We’ll illustrate with our running anti-money laundering (AML) alert queue example, where the task for a human and their AI agents is to examine potential violations of AML laws. There are many ways in which one could partition the work involved in an alert investigation, but we’re going to make a rough distinction between two kinds.

First, there is assembly — pulling the customer’s history, the counterparty’s profile, the prior alerts on the same account, the transaction detail, and composing all of it into a narrative that will survive being read by someone else in two years. And then, there is judgment — deciding whether the pattern in front of you is the thing the scenario was written to identify. AI models are very good at the first and do not, under an approval-before-action arrangement, perform the second.

Throughput remains bounded by the number of reviewers, because every item still requires a review. Per-item cost falls, but those reductions come from the reduced friction associated with assembling context. If assembly was sixty percent of the analyst’s time, in other words, and the model takes the lion’s share of that work, the analyst can work more alerts in a day. Having said that, she cannot work an unbounded number because of the Three Issues Problem.

This fact has a few implications for how productivity appears, how the arrangement can break down, and more. The three sections that follow unpack them.

Why does the productivity gain show up in the wrong place?

A queue whose primary bottleneck is reviewer attention will probably see a model’s performance in the form of a smaller backlog rather than reduced headcount, and a firm that planned around the second might very well be inclined to read the first as a failure.

The business case for AI in an alert queue will likely be written around a staffing number, since staffing is the most visible cost of the queue. What arrives instead is a shorter queue, faster cycle times, and a smaller aging tail, attended to by roughly the same number of analysts. Everyone is working the same hours on more items. The savings are real, of course, but they are denominated in something the original case did not account for.

What the case forecastWhat the arrangement deliversHow it gets readWhat actually changed
Fewer reviewersSame reviewers, shorter backlogModel underperformedAssembly time per item
Lower cost per alertLower cost per alertConfirmedAssembly time per item
Faster clearanceFaster clearanceConfirmedAssembly time per item
Capacity headroomCapacity headroom, bounded by reviewer-hoursPartially confirmedThe ceiling moved; it did not lift

The two middle rows land, and the one that determines whether the project is judged a success (the first) does not.

Where does approval-before-action degrade?

Though it may not always be true, approval quality likely falls as queue depth rises. To see why, consider the differences between three approaches to review:

  • Independent re-derivation, where the reviewer works the alert and compares her conclusion to the draft;
  • Error-checking, where she reads the draft and looks for something wrong with it;
  • Acceptance, where she reads the draft, finds nothing obvious out of place, and moves on.

Regardless of how a particular analyst proceeds, all three will be captured in the same artifact. The third may or may not meet a technical definition of “oversight,” and there is nothing in the record itself that distinguishes it from the rigors of the first.

What makes this worse rather than better as the model improves is that the model’s output quality can work against the reviewer. A rough draft invites correction; a fluent, well-formatted, correctly cited disposition that reaches the right answer nine times in ten does not. The tenth, erroneous disposition will arrive in the same format as the first nine, and the reviewer’s job is to catch it after having just approved nine that were fine. This is the verify-expensive case from earlier in the series, and it has a particular sting under this arrangement: approval-before-action is only a meaningful (and highly scalable) control where verifying the item is cheap relative to producing it. Where verification is expensive — where confirming the disposition means redoing most of the analysis — the arrangement’s one-to-one coverage is nominal, and the reviewer will be incentivized to agree rather than check everything again, especially if deadlines are tight and there’s pressure to pick up the pace.

Note what has no artifact here. The reviewer’s disconfirming work — the 20 minutes she spent pulling the counterparty history to make sure the draft was right — leaves nothing behind when it confirms. Only the approval survives, and (this should sound familiar) the approval looks identical whether it took 20 minutes or 20 seconds.

What do you measure, if coverage is always 1.0?

Start with approvals per reviewer-hour, but assess this figure against the review depth specified by the control description.

The coverage ratio is necessary and not sufficient for the approval-before-action arrangement. It does the work it was proposed to do — it separates approval before action from sampling and post-hoc confirmation, which is the whole point of stating it — and then it stops, because within this arrangement it is a consistent 1.0.

Most queue systems provide metrics around approvals per reviewer-hour. Read on its own, it says nothing; read against the depth the control description implies, it says quite a lot. If the documentation stipulates that an analyst should independently evaluate each alert and the queue is clearing at a rate of ninety seconds per item, those two statements are unlikely to both be true. This isn’t necessarily a result of anyone’s dishonesty; it’s just that they were written by different people at different times for different readers, and no process brings them into contact by default.

What evidence survives?

An approval is a discrete event with a named person, a timestamp, and usually a comment field, which puts this arrangement close to satisfying a serious audit specification for free. Under approval-before-action, most of that exists as a byproduct of the workflow rather than as an additional burden, meaning that many of the schema deficiencies identified in part 3 of “How Financial Professionals Use AI” come close to disappearing.

The record establishes that an approval happened and who owns it, but it does less to capture the depth undergirding that approval. I would love to hear from anyone running a process similar to the approval-before-action arrangement. Talk to me in the M.A.S.E. Discord, or get in touch directly. We are researching how these arrangements behave in production, and what comes back will have an impact on what we write next.

Frequently asked questions

Prove AI is building solutions to power more correct, explainable and auditable AI outcomes.

We’re always interested in learning about AI management challenges.

Get in Touch