Every AI sales demo I've sat through follows roughly the same script.
Someone shows you a dashboard. They enter a company name, click a button, and thirty seconds later a personalized email sequence appears on screen, drafted, formatted, ready to send. The audience nods appreciatively. The presenter says something like "and then it just goes out automatically."
And that's where I always want to raise my hand.
Automatically is doing a lot of work in that sentence. It's glossing over an entire category of decisions that weren't made by the system, they were assumed away. Someone decided the tone was right. Someone decided the timing was right. Someone decided this particular prospect should get this particular message on this particular day. They just didn't show you any of that happening, because showing you a review step isn't as impressive as showing you a one-click send.
The problem is that most teams don't realize they've bought the demo version of the workflow until something goes wrong. And by then, the message has gone to a hundred people.
The Steps That Should Be Mechanical
There's a category of outbound work that is genuinely suited to automation, not as a compromise, not as a shortcut, but because it's repeatable, rules-based, and doesn't benefit from human judgment on a per-contact basis.
Account matching is one of them. Given a defined ICP, the right company size, industry, growth signals, tech stack, matching accounts from a database is a lookup problem. The rule is the judgment. Running it for ten accounts and running it for ten thousand takes the same quality of logic.
Enrichment is another. Pulling firmographic data, finding the right contact at an account, checking whether any qualifying signals are active, this is research work, and it's the kind of research work that benefits from being fast and consistent rather than human. When a human does it, the quality varies by who's doing it and how tired they are. When a system does it, the quality varies by how well the rules were written, which you can fix once and it stays fixed.
Signal detection sits here too. If a VP of Sales just joined a company, that's a fact about the world. The system can observe it. What the system can't do, without help, is decide what that fact means in the context of this specific persona, at this specific account, given what you know about the conversation you're hoping to start. That's the interpretation step, and it requires something the matching and enrichment steps don't.
Drafting, done right, also belongs in this category. If you've encoded persona-specific messaging, defined what a good opening line for this signal looks like, and established the business hypothesis you're trying to open with, the draft is assembly. It's putting the right pieces in the right order. The quality of the draft depends on the quality of the inputs, not on whether a human typed it.
The thing all of these steps share: they can be checked. You can run them against a test set, inspect the outputs, compare them to what a skilled human would have done, and adjust the logic until they're producing results you'd be proud to put your name on. That's what makes them safe to mechanize.
The Steps That Stay Human on Purpose
Then there are steps that look like they could be automated but shouldn't be. The distinction isn't capability, it's consequence and context.
Messaging approval. Before a draft goes out, a human sees it. This is the checkpoint that matters most, and it's the one most automation demos skip because it slows down the story.
The reason it exists isn't because the system drafts poorly. It's because the standard isn't something the system can hold on its own. The draft can be good, matching the persona, referencing the right signal, leading with the right hypothesis and still be wrong for this specific prospect on this specific day. The commercial relationship has context that lives outside the data. The system doesn't know that this account is already in conversation with a competitor. It doesn't know that a senior rep just had coffee with the champion last week. It doesn't know that the company announced layoffs this morning and this probably isn't the moment.
A human review step isn't a sign that the automation is untrustworthy. It's a sign that you've designed a system that knows what it doesn't know.
Commercial judgment. Pricing discussions, contract negotiations, concessions, deal structure, none of this should be automated, and not just for legal reasons. These moments depend on reading a situation in real time, making a judgment call about what this particular buyer values, and being present enough to respond to what's actually happening in the conversation rather than what was predicted in the playbook.
A system can tell you which accounts are most likely to close based on pattern matching. It can't tell you whether to offer a discount to save a deal. That's a relationship call, and it requires someone with authority and context and the ability to be wrong and recover.
Sensitive sends. Champion departures. At-risk renewals. Founder-to-founder outreach. Any message where the wrong tone isn't just ineffective but actually damaging. These go through a human every time, regardless of how good the draft is, because the cost of getting them wrong is asymmetric.
When a champion leaves, the way you handle the next thirty days can determine whether you keep the account. That's not a templating problem. That's a relationship problem, and it needs a person.
The Checkpoint Pattern
What this looks like in practice is a workflow that separates generation from delivery and puts a review queue in between them.
The system handles the mechanical steps: it matches accounts, enriches contacts, detects signals, generates drafts. Those drafts go into a review queue. A human opens the queue, sees the draft, can edit it or approve it or kill it. Once approved, the send is logged, to the CRM, to the delivery system, to whatever shared record keeps the team aligned on what's been said to whom.
account_match() → queue
↓
contact_enrich() → queue
↓
signal_detect() → queue
↓
draft_generate() → REVIEW QUEUE
↓
[human reviews]
↓
approve / edit / reject
↓
send() → log()
The review queue is where human judgment lives. It's not a bottleneck, it's a gate that's there on purpose, because the gate is the point.
One decision you have to make upfront: what does the reviewer need to see to make a good decision? If they just see the draft and the recipient name, they're going to approve things they shouldn't and reject things they should've kept. The review interface needs to surface the signal that triggered the draft, the persona rationale, the campaign context. It needs to show them enough to actually review it, not just sign off on it.
A review step that surfaces no context isn't a safety measure. It's a checkbox that makes you feel like you've kept a human in the loop while actually removing their ability to use judgment.
Why This Is a Design Choice, Not a Limitation
The framing matters here. A lot of teams think of the human review step as a temporary inconvenience, something they'll eventually automate away once the system is good enough.
That's the wrong model.
What stays human on purpose isn't a residual category of things automation hasn't gotten to yet. It's a deliberate design decision about which judgments should have a person accountable for them. Commercial calls, relationship-critical messages, and anything with asymmetric downside, these stay human because the accountability structure is right, not because the system can't draft something plausible.
The goal of a well-designed outbound system isn't to remove people from the process. It's to remove manual, repeatable work from their week so they're present for the moments that actually need them. The rep who's not spending two hours a day on research and drafting has two hours to spend on the calls that matter. That's the compound.
When you build the system with that mental model, mechanical steps automated, judgment steps deliberate, approval gates explicit, you get something the demo version never shows you: an outbound motion that the team actually trusts. Not because it's fast. Because they know exactly what's automated and what isn't, they know where their judgment is expected, and they're not waiting for a surprise to tell them where the gaps are.
The Approval Checklist
Here's the version we use internally as a first-pass gate. Copy it, adapt it, use it as a starting point for defining your own review criteria.
OUTBOUND APPROVAL CHECKLIST
─────────────────────────────────────────────
ACCOUNT MATCH
□ Company fits current ICP criteria
□ No active deal or recent conversation in CRM
□ Not on suppression list (competitor, partner, do-not-contact)
CONTACT SELECTION
□ Persona matches campaign targeting
□ Contact has not been reached in last 90 days
□ Contact is reachable (valid email, active at company)
SIGNAL RATIONALE
□ Signal is recent (within defined lookback window)
□ Signal is relevant to this persona's problem
□ Opening line correctly references signal, not generic
DRAFT REVIEW
□ Tone matches voice guide for this persona
□ Opening hypothesis is specific, not vague
□ No metric claims that can't be substantiated
□ No competitor mentions without explicit approval
□ CTA is singular and low-friction
SENSITIVE SEND FLAGS (escalate if any apply)
□ Champion at account has recently changed
□ Account is currently at-risk or in renewal negotiation
□ Send involves pricing, concession, or contract language
□ Founder-to-founder context requires personal sign-off
□ Any reason this shouldn't go out today (market event,
account news, relationship context outside the system)
─────────────────────────────────────────────
Approved by: ____________ Date: ____________
The last item, "any reason this shouldn't go out today", is intentionally unstructured. The structured items catch the predictable cases. That line is the catch-all for everything the system couldn't know to ask about.
Further Reading
For the full breakdown of where to draw the mechanical/relational line and what the decision-making framework looks like when you're building it for the first time, the source post on what stays human on purpose is the right place to start. The section on the command centre as the review interface is worth reading alongside it.
How does your team handle the review step? Curious whether anyone has built a review interface that surfaces enough context to actually be useful and what made the difference between a gate people trusted versus one they rubber-stamped. Drop it in the comments.
Top comments (0)