The purpose of a general control description is to tell a reader what happens to an item; the phrase “human in the loop” has ceased performing that function because it is now attached to at least four arrangements that differ in terms of:
- When a person touches the work;
- What fraction of the work they touch;
- Whether the action has already occurred by the time they touch it.
This matters more than a terminology complaint usually would, because the phrase is load-bearing. It appears in control descriptions, in model risk documentation, in procurement questionnaires, in the answer a firm gives when a regulator asks how AI-assisted decisions are overseen, and in any number of other places. When one term covers four arrangements with four different oversight profiles, the answer to that question stops being as informative as it would be if more rigorous terminology were on offer.
Take something like an anti-money laundering (AML) alert queue as the working example, and keep it in mind as we proceed. In this example, a transaction monitoring system is running in the background, checking scenarios against activity and generating an alert when some conditions are met. An analyst reviews each alert, gathers context, and decides what to do with it. This kind of workflow is high-volume, heavily documented, and among the first places in a bank where AI got put to work in earnest. It is also a place where all four arrangements are plausibly running right now, in different institutions, all described the same way.
What does “human in the loop” actually mean?
Human in the loop can mean at least four different things, and it’s often not clear from the documentation which meaning is operative.
The Cambridge Centre for Alternative Finance’s 2026 global survey of AI in financial services draws a distinction between:
- Human-in-the-loop: the AI acts, and requires human approval.
- Human-on-the-loop: the AI acts autonomously, and a human can intervene.
- Human-out-of-the-loop: fully autonomous.
- Human support: the AI recommends, and a human acts.
Note that one axis along which you can sort these four terms is authority, which tracks who takes the action, and who has the power to stop it. That is a sensible partitioning strategy, and it has the additional virtue of mapping well onto how a lawyer or a regulator thinks about the problem.
The arrangements this piece takes apart each raise a different set of questions once you look at them closely, and each gets its own treatment in the pieces subsequent to this one.
Which arrangements actually ship under the “human in the loop” label?
At least four, and they don’t sort on the same authority axis delineated above.
If you carefully examine what actually gets built, you find controls distinguished not by authority but by timing and coverage, i.e., when the human touches the item, and what proportion of items they touch. Keeping with our AML alert queue, consider the following schemas:
- Approval before action. Every alert reaches an analyst, who dispositions it. The AI drafts the narrative, assembles the context, proposes a disposition; nothing closes until a person accepts it. Coverage is one-to-one (one human covering each alert), which inevitably bounds throughput by the number of reviewers while raising a question about what happens to the quality of an approval as the depth of the queue increases.
- Exception-only escalation. The automated AI workflow handles alerts that fall below a configured confidence or risk threshold and routes the rest to a human analyst. The analyst sees the exceptions. Coverage is whatever fraction the threshold produces, which is a number someone set and can reset. Against this context, it’s worth pointing out that the definition of an exception will almost certainly drift, which is a problem that shows up most clearly in reconciliation work.
- Spot-check sampling. The system dispositions everything while a quality-control function periodically pulls a sample and reviews it. Under this setup coverage simply is the sampling rate, and the review happens after the disposition is recorded. Sampling assumes error arrives at random, which may or may not be the case.
- Post-hoc confirmation. The system dispositions everything and a person signs off on the batch, or on a summary of the batch, at the end of a period. Coverage is nominally complete and probably substantively unclear, because by the time confirmation happens the output has usually already moved — which is the shape of the third-party valuation problem, and the point at which the question stops being about oversight and starts being about what a receiving system needs to know about an input it didn’t produce.
Now, note carefully that none of the foregoing is a claim about the software (which may very well do exactly what it says), it’s a claim about the label: if particular materials (a report, say, or a document) designate four different control structures as “human in the loop,” a firm reading those materials cannot use the phrase to tell what it’s buying, and an examiner reading a control description cannot use it to tell what a firm is running.
Now put all of this side by side and consider what follows:
| Shipped arrangement | Who acts | Coverage | CCAF taxonomy |
|---|---|---|---|
| Approval before action | Human, on every item | 1:1 | Human-in-the-loop |
| Exception-only escalation | AI, except above threshold | Set by threshold | Human-on-the-loop |
| Spot-check sampling | AI, on every item | Sampling rate | Human-on-the-loop |
| Post-hoc confirmation | AI, on every item | Nominal, after the fact | Human-out-of-the-loop, with a review attached |
Only the first row meets the taxonomy’s standard for what qualifies as “human-in-the-loop”; the other three are something else, and the third and fourth are a considerable distance from it.
Why does the difference matter?
All of this is worth straightening out because the obligation attached to a signature doesn’t scale with the batch.
The previous arc landed on this: attestation is a claim about work that someone actually did, and the standards that govern it ask what was done, by whom, and when. They are largely uninterested in how the work was produced. That indifference is usually read as permissive, and that’s mostly correct. But, here, it cuts the other way. If the obligation is per-item and the arrangement is not, the arrangement has to make up the difference somehow, and three of the four rows in the table above don’t.
Under exception-only escalation, the analyst dispositions the exceptions and the threshold disposes of everything else. Whatever the analyst attests to, it is not the items she didn’t see. Under sampling, a clean result on a sample is inherited by a population — which is a perfectly respectable statistical practice, and a different claim from the one a per-item attestation makes. Under post-hoc confirmation, the action has already happened; the review can characterize it, and can catch a pattern, but it can’t be the judgment that authorized it.
None of these is illegitimate. Sampling is how quality control has worked in every operational function for decades, and thresholds are how any high-volume queue stays tractable. The problem is narrower: the number that would let a reader tell them apart isn’t stated anywhere the label appears.
I’d propose calling that number the coverage ratio, which I’d define as roughly the proportion of items processed in a period that a human being actually reviewed. This isn’t a standard metric, and I’m not aware of anyone reporting it, but I think it fills a useful niche inasmuch as the four arrangements above are indistinguishable without it and trivially distinguishable with it. A queue running at 1.0 and a queue running at 0.04 can carry the same control description, the same label, and the same signature, which is less than ideal.
If you think this reads as a distinction without a difference or is otherwise flawed, I’d like to hear the argument in the M.A.S.E. Discord.
What happens to the coverage ratio as volume rises?
The coverage ratio moves with volume, and nothing announces this change.
The threshold in an exception-based queue is a configurable value; raise it, and more items are automatically processed; the coverage ratio falls, the arrangement slides from something close to approval-before-action toward something close to sampling, and no document changes. The control description still says human-in-the-loop, because it still, in the taxonomy’s loosest sense, is. A change that would require a memo if it were a policy change requires nothing if it’s a settings change.
What does the label cost?
The tools work well enough on alert triage that all four arrangements are running somewhere right now. What a firm frequently cannot do is say — in terms an examiner, a board, or its own second line could act on — which of the four it is running, and a lack of a label is a key part of the reason why.
That’s a naming problem before it’s anything else, which as problems go is a considerably better one to have than an interpretability problem (though we still have plenty of those to go around). The arrangements differ in ways that are perfectly observable; the information that would separate them exists at the moment the work happens. Whether it survives to anywhere useful is a different question, and one we’re not going to settle here.
So here’s what I’d like to know from anyone running one of these queues. If you had to state your coverage ratio — the actual fraction of items a person laid eyes on last month — could you produce the number, and would it match what your control description implies? Tell me in the Discord, or reach out directly. We’re researching this, and the answers shape what we write next.
Frequently asked questions
Prove AI is building solutions to power more correct, explainable and auditable AI outcomes.
We’re always interested in learning about AI management challenges.
Get in Touch


