
Everyone talks about human-in-the-loop — or HITL, if tech acronyms are your love language — with AI agents and agentic workflows. But just as not all AI models are equally capable, and not all tasks get the same token budget, not all human-in-the-loop steps are the same either.
Something that gets flattened by the phrase “human-in-the-loop” is that it treats the human almost like a Boolean: human involved = yes/no. But the human step can vary by orders of magnitude in quality, cost, and value:
How much experience and expertise do they bring to that step?
How much time do they spend thinking in that step?
A summer intern spending 15 seconds checking a box and a 20-year domain expert spending 45 minutes wrestling with an ambiguous decision are both technically human-in-the-loop. But they’re certainly not equivalent.
Because 2×2 frameworks (and Venn diagrams) are my love language, I mapped those two dimensions — expertise and effort — in the illustration at the top of this post. It has four quadrants that categorize the kind of human-in-the-loop intervention we expect:
Verification (low expertise/low effort): Did it work?
Consideration (low expertise/high effort): Does this make sense?
Discernment (high expertise/low effort): Is this good?
Wisdom (high expertise/high effort): What should we do?
Verification is for things where there is a knowable correct answer. The human isn’t being asked for taste or strategy. They’re checking execution against intent.
Consideration involves investigation rather than mere approval. The person doesn’t necessarily have an immediate answer, but spending time reviewing evidence and reasoning through the task adds value.
Discernment leverages compressed expertise. The key is that the expert can often judge something quickly because they already understand the brand, customer, market, or account.
Wisdom involves wrangling a real trade-off that calls for both expertise and deliberation. Something like a pricing exception, a sensitive customer issue, or a campaign go/no-go decision.
The point isn’t that wisdom is always better than verification. The goal is to allocate the right kind — and amount — of human judgment to the task.
(There’s a little Thinking, Fast and Slow hiding in this matrix too. Discernment is often “fast” thinking — expertise compressed into intuition and pattern recognition. Consideration and wisdom are more “slow” thinking — deliberately working through evidence and alternatives.)

Since there’s a lot of discussion right now about model routing and token spend with AI, I propose that, in my 2×2, expertise is roughly analogous to model capability, while effort is roughly analogous to inference-time compute. (Admittedly, you have to squint.)
A stronger human “model” can often reach a good judgment with fewer “tokens.” But more tokens can’t always compensate for a weaker model — and even the strongest model can perform poorly if you starve it of thinking time.
Human-in-the-loop is, in part, a compute allocation problem: which human, with how much attention?
Human-in-the-loop isn’t always in the loop
There’s another way the phrase “human-in-the-loop” flattens reality: it implies there’s only one place for the human to be. But there are several different ways humans can be positioned in relation to agentic workflows:
Human in the loop: the agent pauses and waits for a human decision before it can continue. Think approval gates, required reviews, or sign-off on a proposed action.
Human on the loop: the agent keeps running, while a human monitors what it’s doing and can intervene when needed.
Human at the edge: the agent handles routine cases autonomously, while exceptions, anomalies, or sensitive situations get escalated to a human.
Human before the loop: people establish objectives, policies, examples, guardrails, and escalation criteria that shape how the agent behaves.
Human after the loop: people audit the agent’s outcomes, look for patterns of failure or opportunity, and improve how the agent operates over time.
Where the human sits changes what their judgment is doing and what it costs — approving an action, monitoring behavior, handling exceptions, shaping the system, or learning from its outcomes.
Imagine an AI agent dynamically generating customer messaging across campaigns. One human-in-the-loop design would require a marketer to review every new message, offer, or audience variation before the agent can deploy it.
That may be appropriate for particularly sensitive messages, but it can quickly turn scarce expertise into an approval queue.
The same marketer might create far more leverage by spending an hour before the loop defining what good messaging looks like, which claims and offers are allowed, what kinds of personalization cross the line, and which situations should escalate to a human. Then, after the loop, they can review a sample of outcomes, spot recurring problems or opportunities, and improve the workflow.
The highest-leverage use of human expertise can be shaping the decisions the system makes, not reviewing them one by one.
Building with rules, reasoning, and responsibility
When I was a kid — back in another century — basic education was sometimes referred to as the three R’s: reading, ‘riting, and ‘rithmetic. (Yes, I know, the literacy jokes ‘rite themselves. But I appreciate the stretch for alliteration.)
Without much stretching (or aphaeresis), the basic building blocks of agentic workflows are also three R’s: rules, reasoning, and responsibility.
Rules are the deterministic steps in a workflow — the if-this-then-that logic that classic marketing automation was built on, and that we still rely on for things that must happen predictably, consistently, or within hard guardrails.
Reasoning is where we leverage an LLM (or SLM) for inference — to interpret ambiguity, synthesize signals, and generalize across situations too varied to encode explicitly.
Responsibility is where we put the human-in-the-loop — in, on, before, after, or at the edge — when we need accountability, judgment, discretion, or taste.
The art of agentic architecture is in how we combine them.
An agentic workflow can use rules to enforce eligibility criteria, spending limits, or other hard constraints; reasoning to interpret the situation and choose an appropriate next action; and responsibility when a human should verify, investigate, or make a judgment call.
One useful way to decide how much responsibility to introduce — and where — is to consider uncertainty, consequence, and reversibility.
Uncertainty calls for deliberation. The less certain the answer, the stronger the case for spending more effort examining evidence and alternatives.
Consequence calls for expertise. The larger the potential blast radius, the stronger the case for routing the decision to someone with deeper experience and judgment.
Reversibility affords distance. The easier a mistake is to detect and undo, the farther the human can sit from the individual action.
These three factors interact. A highly uncertain decision may not deserve much human attention if the stakes are trivial and the action is easily reversed. A relatively straightforward decision may still warrant expert involvement if getting it wrong would have major consequences.
This maps back to the 2×2 and the different positions around the loop. Uncertainty pushes us right toward more effort. Consequence pushes us up toward more expertise. Reversibility determines where the human needs to sit.
More uncertainty → more effort
More consequence → more expertise
More reversibility → more distance
Don’t spend a frontier human on a checkbox.
Closing this loop,
Scott
P.S. Headed to Dreamforce this year? Come join me and Doug Tallmadge, the CEO of Gradial, on Tuesday, Sep 15 for a candid fireside chat about The End of Martech: Marketing's New Operating Model. Yes, me, Mr. Martech, talking about a post-martech world — that is, the industry-wide shift from stacks of applications to more foundational and fluid infrastructure — with the founder of one of the hottest AI-native marketing platforms native to that new world.
It’ll get a little spicy, but it’s a good heat.
Our session is at 1pm in Theater 2 of the Moscone North LL Campground. After our chat on stage, we’ll be camping out at Gradial’s booth to meet folks and talk one-on-one from 2pm-3:30pm. Click the banner below to reserve a spot. Would love to see you there!


