I have sat through a lot of AI demos this year, and they all end the same way. The agent reads the email, drafts the reply, updates the CRM, and the room claps. Then the pilot starts, and three weeks later nobody is using it.
It is rarely because the model was wrong. It is because nobody designed the moment after the demo: the Tuesday morning when a clinic manager, a leasing agent or a restaurant GM opens their phone and has to decide whether to let the AI send the thing. That moment has a screen. On most projects, that screen is an afterthought, a raw JSON log, or an email that says "Action required."
The model is the engine. The approval screen is the steering wheel. People do not buy engines.
I spent years designing for a bank, where the rule is simple: checking your balance takes one tap, moving money gets a review screen. Nobody thinks of that review screen as friction. It is why people trust the app with their paycheque. AI systems need exactly the same thing, and most teams are shipping them without it.
What a good approval queue feels like
Here is the version we design toward. A dental clinic in Vancouver has an AI handling its front-desk inbox overnight. By 8 a.m. the manager has a short queue on her phone. She clears it before her first coffee is cold.
Eight items, ninety seconds
Watch the order of decisions. Nothing leaves the building without a person, but the person only does the part that needs judgment.
- The AI drafted a reschedule reply. She checks the evidence it used.
- One word is off. She fixes it; the edit is saved as feedback.
- Approve. It sends, and the send is logged.
- A $480 refund is over her limit, so it goes to the owner.
- Six routine reminders: one tap approves the batch.
Refunds over $250 need Dr. Anand. The AI attached the invoice, the cancellation note and the card on file.
Notice what she did not do. She did not open the calendar, find the policy doc, retype a message, or forward the refund with a "can you look at this?" The AI did the legwork and brought the evidence with it. She made five decisions. That is the whole job of the screen: turn AI output into decisions a person can make in seconds.
Three lanes, not one on/off switch
The most common mistake I see is treating autonomy as a single switch for the whole system: either "the AI runs everything" or "a human checks everything." The first one breaks trust the first time it emails the wrong patient. The second one creates a queue so long that people start approving without reading, which is worse than no approval at all.
Autonomy should be decided per action type, and the deciding question is not "how smart is the model?" It is "how bad is it if this is wrong, and can we undo it?"
Reversible, low stakes, proven.
Customer-facing or hard to unsend.
Money, legal, or over a limit.
| Lane | Typical actions | How approval works |
|---|---|---|
| Runs on its own | Tagging emails, internal notes, reminders from a fixed template, updating a status field | No approval. Shown in a daily digest so someone can spot drift. |
| Needs a quick yes | Replies to customers, quotes inside a price band, booking changes | One tap per item or per batch, with the evidence on the card and a short undo window. |
| Needs the owner | Refunds, discounts over a limit, contract changes, anything legal or medical | Routed to a named person with the full context attached. Never auto-approved. |
This is also where most of the cost savings actually come from. In a typical service business, the "runs on its own" lane is the biggest by volume and the "needs the owner" lane is the smallest. You get the speed on the boring majority and keep a human on the few decisions that can hurt you.
The anatomy of an approval card
If the lanes are the strategy, the card is where it succeeds or fails. We sketch this card before we write a single prompt, because it forces the right questions: what does the AI need to show for a person to say yes in two seconds?
Around the card, a few rules do most of the work:
- Evidence beats confidence scores. "92% confident" means nothing to a clinic manager. "Thursday 3:30 is open with Dr. Lee" is something she can verify instantly. Show sources, not probabilities.
- Friction should match irreversibility. A reminder needs no confirmation. A refund needs a named approver. Borrow the banking rule: the harder something is to undo, the more deliberate the screen.
- Edits are data. When someone changes "Thursday" to "Thursday at 3:30", capture the difference. Those edits are the best training signal you will ever get, and they tell you which action types are not ready for autonomy.
- Batch the boring. Six identical reminders are one decision, not six. Group items that share a template and a risk level.
- Every yes leaves a receipt. Who approved, when, what they saw, and what changed. When a customer disputes something in March, you want that trail, not a guess.
Autonomy is earned, per action
The approval lane is not meant to be permanent. It is a probation period. Each action type builds a track record, and the numbers tell you when it is ready to run on its own.
Four numbers are enough to run this:
- Edit rate: the share of drafts a person changed before approving.
- Reversal rate: approved actions that later had to be undone.
- Time in queue: how long items wait. If it climbs, the queue is too long or in the wrong place.
- Rubber-stamp rate: approvals made in under a second. A high number means people have stopped reading, which is a design problem, not a discipline problem.
Our default rule: four straight weeks under 5% edits with zero reversals earns auto mode for that action type. One incident sends it back to the approval lane. The owner can override either way, and the override is logged too.
Put the queue where people already are
The last mistake is building a beautiful approval dashboard that lives at a URL nobody visits. Approvals have to show up where the approver already spends their day.
- For an owner-operator, that is usually their phone, with a notification that opens straight into the card.
- For teams in India, it is very often WhatsApp. An approval that arrives as a WhatsApp message with two buttons gets answered; a login link does not.
- For office teams in Canada and the US, it is Slack or Teams, with the full card one tap away.
The better agent frameworks have started shipping "approval gates" as a built-in step, which is good news. But a gate in the code is not the same as a screen a person trusts. The framework can pause the agent. Only design can make the pause useful.
How we build this at LoopSuit
On every AI system we ship, the order is the same. First we list every action the AI could take and sort it into the three lanes with the owner. Then we design the approval card and the queue before any prompt is written. Then we build the automation underneath, instrument the four numbers, and hand over a system the client owns, including the rules for when each action graduates.
It is less exciting than a demo where the agent does everything. It is also the version still running six months later. You can see the same approach in our VeritasGuard compliance work, where every AI finding carries its evidence and an audit trail, and in why AI systems need product design.
Questions people ask us
What does human-in-the-loop mean for a small business?
It means the AI drafts or prepares the work and a person approves it before anything irreversible happens: a message goes out, money moves, a record changes. The person stays accountable; the AI removes the typing, lookup and copy-paste.
Won't approving everything slow us down more than doing it ourselves?
Only if every action sits in the same lane. Low-risk, reversible actions (reminders, internal notes) should run on their own or be approved in batches. Only actions that are hard to undo need a one-by-one decision, and those are usually a small share of the volume.
How do we decide when an AI action can run without approval?
Track it. When people approve an action type for four straight weeks with very few edits and no reversals, it has earned auto mode. If it ever causes an incident, it drops back to the approval lane. Autonomy is granted per action type, never for the whole system at once.
Where should the approval queue live?
Where the approver already is. For a clinic owner that might be their phone; for an India-based team it is often WhatsApp; for an office team it is Slack or Teams. A new dashboard nobody opens is how AI pilots quietly die.
Is a confidence score enough for people to trust AI output?
No. A percentage means nothing to an operator. Show the evidence the AI used (the calendar slot, the policy line, the previous message) so the person can check it in two seconds. That is what builds trust.