← All articles

GenAI

Human in the Loop Is a Placement Decision, Not a Checkbox

Picture two security guards watching the same door.

Guard A stops five people a day and checks their badges properly — looks at the photo, reads the name, asks a question if something’s off.

Guard B stops five hundred. By the three-hundredth person, he’s waving everyone through on reflex. He isn’t lazy. He’s a human being, and no one can scrutinise five hundred badges a day with the same care they gave the first five.

Now a stranger walks in with a forged badge. Which guard catches him?

That’s the entire problem with “human in the loop.” Both guards are humans. Both are in the loop. Only one is actually a safeguard.

Two guards, one door — 500 approvals a day vs 5

The phrase that hides the decision

“Keep a human in the loop” has become a reflex in agent design. Add a confirmation step, call it oversight, ship it. It feels responsible.

But it skips the only question that matters: where does the human go?

An agent workflow isn’t one action. It’s a chain — read the request, look up some data, draft a response, take an action, notify someone. The human can sit at any link in that chain, or several, or all of them. Each placement has a cost. Most teams never make the choice consciously; they either gate everything out of caution, or gate nothing out of impatience. Both are mistakes, and they fail in opposite directions.

The two ways it fails

Too little review. The agent can send the email, move the money, delete the records, change the permissions — and nobody looks until after. An agent with no gate on a consequential action carries the full blast radius of its worst possible decision, every single time it runs. Most of the time it’s fine. The one time it isn’t, there was no one there.

Too much review. This is the failure almost everyone misses, because it looks like diligence. Gate every step, and the reviewer drowns. Fifty approvals a day, then two hundred, then five hundred. Attention is a finite resource, and every routine “approve this harmless thing” spends a little of it. Past a point, the reviewer stops reading and starts clicking. You’ve kept the ceremony of oversight and lost the substance — which is worse than no gate at all, because now everyone believes the system is supervised.

It even has a name: approval fatigue. And it has a counterintuitive consequence that’s worth sitting with.

The hump: why more oversight can mean less safety

Here’s the finding that changed how I think about this.

Researchers modelled exactly the two-guards scenario — an oversight policy that escalates a handful of actions a day versus one that escalates hundreds. The intuitive assumption is that more escalation is always safer: at worst, it’s wasteful. The modelling says otherwise. Because the reviewer fatigues, a more-escalating policy can be the less safe one. The safety-optimal escalation rate sits well below “escalate everything.”

That’s not a straight-line trade-off where you pick your point. It’s a hump. Safety climbs as you add gates to the dangerous actions, peaks, and then falls as you keep adding gates to the harmless ones — because each harmless approval is quietly eroding the attention you need for the dangerous one.

Rather than describe the shape, let me let you find it. Below is a small workflow. Tick which steps get a human gate, set how many tasks the agent handles in a day, and watch what happens to effective safety.

A five-step agent workflow. Tick where the human gate goes, set the daily volume, and watch effective safety — not just the number of approvals.

Approvals per day

—

interruptions for the reviewer

Reviewer attention

—

how carefully each one is read

Risk left unguarded

—

consequential steps with no gate

Effective safety

—

risk covered × attention paid

Try gating everything. Then try gating only the one irreversible step. Notice which one is actually safer — and how many fewer interruptions it costs.

Where the gate goes

So if “everywhere” is wrong and “nowhere” is wrong, where’s right?

An action deserves a human checkpoint when it is irreversible, costly, regulated, or high-blast-radius — and especially when it’s more than one of those at once. Wiring money is all four. Deleting production data is at least three. Reading a calendar is none of them.

The most useful rule of thumb in agent design right now is to split every action into a read or a write:

  • Reads — looking things up, searching, fetching, summarising. Let the agent run. Nothing is changed in the world; a wrong read is recoverable.
  • Writes — anything that sends, pays, commits, deletes, or changes a record of record. These are where consequences live. Gate the writes that carry weight.

Walk your workflow, find the expensive irreversible step, and put the human there. Usually there are one or two such steps in a chain of ten. That’s your checkpoint — not the other eight.

Not every gate has to block

A second refinement most teams skip: a checkpoint doesn’t have to stop the agent cold.

Approve-before-act is for the genuinely irreversible — the money, the deletion, the external send. The agent proposes, waits, and does nothing until a human says yes. This is the right gate for a small number of actions, and it’s slow by design.

Review-after-act is for everything with consequence but reversibility. The agent does the thing, logs it, and a human reviews the log — hourly, daily, in a batch. You keep accountability and auditability without paying the latency on every action. A draft saved to a folder, a ticket created, a label applied: review these after, don’t block on them.

Reserve the blocking gate for the actions that actually justify the latency. Everything else can run and be checked.

Design the gate, not just the veto

Even a well-placed checkpoint can be a bad one. Three things separate a gate that works from one that gets rubber-stamped:

Show the reasoning next to the action. A gate that says “Agent wants to send this. Approve?” with nothing else is asking a human to judge blind. A good gate shows the draft next to the thread it’s answering, the invoice next to the match it found, the proposed change next to the evidence for it. If approving an action means opening three other systems to check it, the gate is in the wrong place or missing the wrong information — and your reviewers will start clicking approve without reading.

Give every pending approval a deadline. A checkpoint that can block forever is its own failure mode. What happens to the task if nobody approves it by 5pm? If the answer is “it sits there indefinitely,” you’ve built a workflow that quietly dies. Every gate needs a timeout and a fallback — escalate, expire, or default to the safe option.

Measure the override rate. Every checkpoint should produce data: how often does the reviewer actually change or reject what the agent proposed? If reviewers are correcting outputs frequently, the gate is earning its place. If they approve 99% unchanged, one of two things is true — the agent has got good enough that the gate can relax, or the reviewer has stopped looking. Either way, the number tells you something a feeling of safety never will.

The honest limits

Three things a well-placed gate still doesn’t give you.

Approval is not security. A human approving an action is only as good as the match between what they approved and what actually executes. If the payload can change between approval and execution — a config that updates, a command that hides arguments behind what was shown — the gate was theatre. Approval has to sit alongside permissions, sandboxing, and verification, not replace them.

Some domains don’t get to optimise. Credit decisions, benefits, clinical steps, anything a regulator has an opinion about — these warrant a gate regardless of how reliable the agent has been or how fatigued the reviewer is. In those cases the fix for fatigue isn’t fewer gates; it’s more reviewers.

The agent shouldn’t decide when to ask. It’s tempting to put the approval logic in the prompt — “ask for confirmation before anything risky.” Don’t. The agent’s judgement of “risky” is exactly the thing you’re trying to oversee. Approval rules belong in the workflow, enforced outside the model, where the model can’t talk itself out of them.

Oversight has a capacity

The mistake underneath all of this is treating human attention as free and infinite. It’s neither. A reviewer has a budget — a number of decisions a day they can make well — and every gate you add spends from it.

“Human in the loop” sounds like a safety property. It’s actually a resource-allocation problem. You have a limited amount of careful human attention. The design question isn’t whether to use it; it’s where.

Spend it on the one action that can’t be undone. Let the rest run.