Human in the Loop Approval Workflow for AI Agents

A human in the loop approval workflow for AI agents is more than a Slack button. How to pick gates, model interrupts, and keep long runs durable and auditable.
Approval Is a System, Not a Button
Most teams start building a human in the loop approval workflow for AI agents the week after they needed one. The agent sent the wrong invoice, or refunded a customer twice, or pushed a config change at 2am that nobody asked for. The reflex fix is a Slack message with an Approve and a Reject button, shipped in an afternoon.
Two weeks later that channel has 400 unread approvals, an engineer is clicking Approve without reading, and the control you added has become a rubber stamp with an audit trail. You now have the worst of both worlds: the latency of a human and the reliability of nobody.
Real oversight is a piece of system design, not a UI widget. It has to decide which actions stop, how the run survives the wait, what the reviewer is actually shown, and what happens when the answer never comes. At Kuaray we build these gates into agentic systems before they go near a production credential, and the design decisions below are the ones that matter.
What a Gate Is Actually For
A gate is not there because the model is untrustworthy. It's there because some actions are irreversible, and irreversibility is the only property that matters when you're deciding where to spend human attention.
An agent that reads a database, drafts a summary, or proposes a plan can be wrong a hundred times a day at nearly zero cost — you read the output and discard it. An agent that issues a refund, deletes a bucket, emails a customer, or merges to main has spent something you cannot get back. The first category needs evaluation. The second needs a gate.
This reframing kills the most common design mistake, which is gating by confidence. Teams wire up "if the model's confidence is below 0.8, ask a human." Self-reported confidence is not a calibrated probability, it's a token distribution over words like "certain," and it correlates poorly with correctness on exactly the long-horizon tasks where you need it most. Gate by consequence, which you can enumerate, not by confidence, which you cannot trust.
If you want the wider picture of what goes wrong when those controls are missing, we covered the failure taxonomy in why AI agents fail in production. This piece is about building the gate itself.
Where a Human in the Loop Approval Workflow for AI Agents Puts Its Gates
Every tool the agent can call gets classified once, and the classification determines the control. Four tiers is usually enough.
| Tier | Example actions | Reversible? | Control |
|---|---|---|---|
| Read | Query DB, fetch ticket, search docs | Yes | No gate — budget + rate limit |
| Internal write | Draft doc, create branch, stage file | Yes, cheaply | Post-hoc review, async |
| External write | Send email, post to CRM, open PR | Awkwardly | Approval gate, batched |
| Irreversible / financial | Refund, payment, delete, prod deploy | No | Approval gate, per-action, named approver |
The point of the table is that most tools should have no gate at all. If everything is gated, nothing is reviewed. A workflow where 90% of approvals are trivially fine trains the reviewer to click through the 10% that aren't — this is approval fatigue, and it is the defining failure of naive human-in-the-loop design.
There's a second axis worth adding once the first is stable: magnitude. A $12 refund and a $12,000 refund are the same tool call with wildly different consequences. Threshold rules inside a tier ("auto-approve under $50, gate above, two approvers above $5,000") remove far more noise than any prompt engineering will.
The Mechanics: Interrupt, Persist, Resume
Here is where most implementations quietly break. An agent run is a stateful process — a message history, a tool call in flight, partial work already done. A human approval takes minutes at best and days at worst. Something has to hold that state across the gap, and it cannot be a Python process waiting on a blocking HTTP call.
The three approaches you'll actually choose between:
- Blocking call in-process. The agent calls a tool that polls for a decision. Trivial to write, and it dies the moment the process restarts, the container is rescheduled, or the reviewer goes to lunch past your timeout. Fine for a demo, never for production.
- Interrupt and checkpoint. The agent framework serialises its state to a store when it hits a gate, returns, and is rehydrated from that checkpoint when the decision arrives. This is the pattern behind LangGraph's
interruptprimitive and equivalent checkpointer designs, and it's the right default for most teams. - Durable execution. The whole run is a workflow in a durable engine — Temporal, Restate, or similar — where the approval is a signal the workflow waits on. Execution history is replayed deterministically after any crash. Heavier to adopt, and the correct answer when runs span days, touch money, or must survive a full region failover.
Whichever you pick, three properties are non-negotiable. The gate must be idempotent — a reviewer clicking Approve twice, or your queue redelivering the message, must not issue two refunds. The pending action must be immutable once proposed: the agent does not get to revise what it's doing between the request and the decision, because the reviewer approved a specific payload, not an intention. And the gate needs a default on timeout, declared per tier. For irreversible actions the default is always deny and escalate; silent expiry into an approval is how you build an unattended system that looks supervised.
What the Reviewer Sees
The approval payload is the product. A reviewer given "Agent wants to send an email — Approve/Reject" has no basis for a decision and will approve out of politeness. The request has to carry:
- The exact action, rendered as it will execute — the actual recipient, amount, SQL, diff, or API body. Not a paraphrase.
- The reasoning trace, compressed to why this action follows from the task.
- The provenance: which request or ticket started this run, and who or what authorised it.
- A diff or preview where one exists. For a deploy, the changeset. For a DB write, the before/after rows.
- A one-click alternative to binary approval: edit-and-approve. Most rejections in practice are "right idea, wrong value," and forcing a full rerun to fix a date wastes the human's time and the model's tokens.
Get the payload right and review takes eight seconds. Get it wrong and every approval becomes an investigation, which is the real reason these systems get abandoned.
Escalation, Authority and Audit
Approval implies an approver, and "whoever is in the channel" is not an authorisation model. In any regulated or enterprise context — and increasingly under the EU AI Act for higher-risk systems — you need to answer who was entitled to approve that action, and prove it later.
Practically: map gates to roles, not people. Resolve the on-call approver from the same identity provider that governs everything else, so an offboarded employee stops being a valid approver the same hour they lose their laptop. Require a second approver above a threshold, and forbid the requester from being the approver when the run was human-initiated. Then log the decision as a first-class record — approver identity, timestamp, the exact payload approved, and the decision — in append-only storage alongside the agent's own run journal. An audit trail assembled from Slack history is not an audit trail.
Two more things worth designing deliberately. Rejections are training data: every denied action is a labelled example of your policy, and a rejection reason field turns a control into a feedback loop that shrinks the gate over time. And autonomy should be earned: start a new agent fully gated, measure the approval rate per tool, and promote tools to auto-approve once they've cleared a few hundred clean decisions. That's how the number of interrupts goes down without anyone deciding to "trust it now."
What a Good Implementation Looks Like
Done well, a human in the loop approval workflow for AI agents disappears into the system. Gates sit in the tool layer, not the prompt — a tool that structurally cannot execute without a decision record is a control; an instruction asking the model to check first is a hope. State lives in a checkpointer or a durable workflow, so a pending approval survives a deploy. The reviewer gets a complete, editable payload in the surface they already live in. Every decision lands in an immutable log. And the set of gated tools shrinks quarter over quarter as evidence accumulates, rather than growing until everyone ignores it.
That's the difference between human oversight of autonomous agents and the theatre of it. If you're building agentic automation that touches money, customers or production infrastructure, this is the layer worth over-engineering. We do this work as part of AI and agent engineering, wired into the same identity and logging stack as the rest of your platform via our DevOps practice — and if you're building the first version, our MVP approach puts the gate in before the agent gets a credential.
Talk to Kuaray about designing approval gates for your agents — we audit what your agents can already do unattended, and build the oversight layer that lets you widen their autonomy without guessing.