Human-in-the-loop design: where to put the approval gate
"A human approves it before it happens" sounds like a safety net until the human is asked to approve the hundredth low-stakes action of the day and starts clicking approve without reading, at which point the safety net is a formality that happens to still be in the audit log.
1. The gate isn't the control -- the reviewer's attention is
An approval step is often described as if inserting it is itself the safety measure: "a human approves every deployment" sounds like a control regardless of what that approval actually involves. It isn't. The control is whatever judgment the human applies in the moment they're asked to approve, and that judgment is a finite, exhaustible resource that degrades with volume and with how little information the approval screen actually gives them to evaluate. A gate that presents a reviewer with "approve this action, yes or no" and nothing else is asking them to either trust the system completely or block everything out of caution -- neither of which is the judgment the gate was supposed to introduce.
This reframes the actual design problem. It's not "where do we add a human" -- it's "what does the approver need in front of them to make a real decision, and how do we keep the volume of requests low enough, and the stakes of each one high enough, that reading it is worth their attention every time." A gate that fails either half of that -- insufficient information, or too much volume -- will degrade into a rubber stamp, usually within weeks, regardless of how carefully it was justified when it was designed.
2. What belongs in front of the reviewer
The single most common design mistake is asking a reviewer to approve an intention rather than an effect. "Agent wants to deploy to production" is an intention; "this diff, these three files, this is what changes in the running system if you approve it" is an effect, and only the second gives a reviewer something they can actually evaluate. The same distinction applies everywhere a gate shows up: not "agent wants to email the customer," but the drafted email itself; not "agent wants to modify the pricing table," but the specific rows and the specific before/after values.
Implementation note: where the action is reversible with a clear undo (a draft that hasn't sent, a change that hasn't deployed), show the effect and let approval trigger the irreversible step. Where the action is not cleanly reversible (an email that's already sent the moment it's approved, a payment that's already settled), the approval screen is the last point a mistake can be caught, which raises the bar for how complete and accurate that screen's representation of the real effect needs to be -- a summary that drops a detail here is categorically worse than the same gap on a reversible action.
Beyond the diff or the draft itself, a reviewer making a real decision usually needs: what triggered this action (the specific request or event, not a generic "the agent decided to"), what the agent considered and rejected if there was a real alternative, and a rollback path if approval turns out to be wrong. Omitting the third item is common and expensive -- a gate that can approve forward but has no documented path to undo turns every approval into a one-way bet, which pushes reviewers toward caution that has nothing to do with the actual risk of the specific action in front of them.
3. Placement: not every risky action needs the same gate
Not every consequential action deserves a synchronous, blocking approval, and treating them all the same is exactly what produces reviewer fatigue. A useful split is by two independent axes: how reversible the action is, and how much it costs to delay it until a human is available. A low-cost, reversible action (a draft reply, a non-production config change) can go through an asynchronous review -- it executes into a staging state and a human reviews on their own schedule, with the system clearly tracking what's pending. A high-cost, irreversible action (an external payment, a customer-facing message that can't be recalled) needs synchronous, blocking approval even if that means the agent waits. The mistake to avoid is defaulting everything to synchronous blocking approval because it feels safer uniformly -- it isn't uniformly safer, it's uniformly slower, and the slowness is what eventually gets the gate quietly routed around or the threshold for "needs approval" quietly raised.
A second placement question: approval at the step level, or at the outcome level? Approving every individual tool call in a multi-step task is thorough and also the fastest route to reviewer fatigue, because most individual steps in a well-functioning task are unremarkable. Approving only the final outcome is faster for the reviewer but means a problem introduced at step three is invisible until the entire chain has already run. The usual resolution is to gate the steps that are irreversible or externally visible, and let the reversible, internal steps run without a stop -- which is the same reversibility axis from above, applied inside a single task rather than across a system.
4. The failure pattern: habituation
Common mistake: a gate that worked well at low volume gets routed to the same reviewer as volume grows, with no change to the screen or the stakes per request. The reviewer's actual behavior shifts from reading to pattern-matching ("this looks like the last fifty I approved") well before anyone notices, because nothing in the system distinguishes a careful approval from a reflexive one -- both produce an identical "approved" event in the log.
The only reliable countermeasures are structural, not motivational -- telling a fatigued reviewer to "please read more carefully" doesn't survive contact with volume. Cap the rate: if a queue is growing faster than a human can give each item real attention, that's a signal to raise the bar for what reaches the gate at all, not a signal to ask the human to go faster. Vary the presentation of genuinely different-risk items so a reviewer's pattern-matching doesn't collapse everything into one visual shape. And periodically route a known-bad or known-ambiguous case into the queue deliberately, specifically to check whether the gate is still catching things -- a calibration practice borrowed from manual QA sampling, and one of the few direct measurements of whether a human-in-the-loop control is still a control.
5. A minimal checklist
- The screen shows the effect, not the intention — the actual diff, draft, or change, not a one-line description of what the agent wants to do.
- Context is attached: what triggered this, what alternative the agent considered, what the rollback path is.
- Placement matches reversibility and cost of delay — not every action gets the same gate, and not every action gets a gate at all.
- Volume is capped to what a human can actually read, not expanded to match whatever the system produces.
- Calibration checks run periodically — a known problem case is routed through the gate on purpose, to confirm it still gets caught.
- Every approval is logged with enough context that a later audit can tell what the reviewer actually saw, not just that they clicked approve.
6. Where this fits
Approval gates are one of the control points AI harness engineering names as part of the surrounding system, and the same reversibility-based placement logic shows up in the risk framing of prompt injection and agent security — an injected instruction that reaches an irreversible action with no gate in front of it is the specific failure a well-placed approval step exists to prevent.