DEV Community

Cover image for I Let an LLM Write Cold Emails but Never Decide Who Gets Them: Lessons From 678 Real Sends
mhk sameera
mhk sameera

Posted on

I Let an LLM Write Cold Emails but Never Decide Who Gets Them: Lessons From 678 Real Sends

Connecting a contact database, an LLM and a company mail server is the easy part. A weekend gets you a pipeline that pulls a thousand leads, has GPT write to each one, and sends the result unattended.

Then it does what unattended pipelines do. It emails three people at the same company. It attaches a stranger's job title to a lead. It invents a product the sender doesn't sell. It keeps chasing someone who already replied. It counts an out-of-office message as interest.

Every one of those failures lands in a real person's inbox, under a real employee's name.

I built a system for a food exporter that recruits international distributors, and ran it for nine weeks with six sales reps. This post is the architecture, the specific hazards, and the numbers, including the parts that didn't work.

(I also wrote this up as a full paper with references and the complete evaluation: read it on SSRN.)

The stack

  • Next.js (App Router, TypeScript) + MongoDB, one service on a VPS
  • Apollo.io for people search and enrichment
  • An OpenAI-compatible LLM (GPT-4o class) for language tasks only
  • Microsoft Graph to send and read mail from each rep's own mailbox

57 API routes, 17 pages, about 31,000 lines of TypeScript.

The one rule that shaped everything

The model writes prose and picks from lists. It never decides whether, when, or to whom a message is sent.

Everything with commercial or reputational consequence is a small, pure, unit-tested function with no database or network access: identity matching, deliverability, ownership, approval, pacing, stopping. The LLM sits in a box on the side.

Apollo API    LLM endpoint    Microsoft Graph
     │             │                │
     └──── adapter layer ───────────┘
                   │
   deterministic policy core (pure, tested)
   identity guard · approval gate · ownership
   pacing · reply classifier · follow-up rules
                   │
   interactive routes  +  scheduled jobs
        (both call the SAME policy modules)
Enter fullscreen mode Exit fullscreen mode

That last line matters. An early version had the scheduled sender keep its own copy of the queue conditions. Safety rules added later applied everywhere except the one path that runs unattended. One shared filter removed that whole class of bug.

Hazard 1: the LLM invents identifiers

Apollo identifies industries by internal 24-character IDs, and one invalid ID fails the whole search with an unhelpful error. So the model is never allowed to produce IDs.

A rep types: "owners and purchasing directors at confectionery importers in Mexico, 10 to 200 employees." The model turns that into a filter, but it can only choose industry names from a fixed list of 82 passed in the prompt. The server resolves names to IDs from a catalogue harvested from Apollo. Seniority and company-size values are checked against closed sets, and anything outside them is silently dropped.

The pattern: the model chooses from a list the program owns, and the program translates the choice into anything with operational consequence.

Even so, it failed in a quiet way. Asked for confectionery distributors, the model reliably picked import/export, wholesale and logistics, and left out "food & beverages". Industry filters are OR-combined, so omission throws no error. It just silently shrinks your reach. Errors of omission are harder to catch than errors of commission, so I added an operator-owned override that pins the industry list.

Hazard 2: fuzzy enrichment returns a plausible stranger

Apollo's person-match is fuzzy. When it doesn't hold the address you queried, it can return someone with the same first name at a similarly named company. Attach that record and you've put another person's job title into a personalised email.

The identity guard:

function acceptEnrichment(query: Lead, match: ApolloPerson): boolean {
  // exact address returned: accept
  if (match.email && sameAddress(match.email, query.email)) return true;

  // a different confirmed address: it's another person
  if (match.email && !sameAddress(match.email, query.email)) return false;

  // address hidden: require BOTH company domain and surname to match,
  // compared after Unicode normalisation and accent stripping
  return (
    sameDomain(match.companyDomain, domainOf(query.email)) &&
    sameSurname(match.lastName, query.lastName)
  );
}
Enter fullscreen mode Exit fullscreen mode

LinkedIn URLs get the same suspicion: kept only if the slug contains the person's name, and both first and last name are required when known, because a common first name alone admits the wrong profile.

Hazard 3: the model makes things up about the sender

Give an LLM only a company name and it will confidently describe products, prices and credentials you don't have.

So the sender's own approved template is the source of truth, and the prompt says the model may re-say what the template offers but must not introduce any product, service, claim, price, statistic, timeline or credential that isn't in it. The recipient's website text (first 4,000 characters) supplies the personal detail.

Then deterministic post-processing cleans up after the model: strip any sign-off it added (the real signature is appended later) and restore the presentation link if it dropped it. A missing call to action is invisible until nobody replies.

Translation is a separate structured call at low temperature, with rules to keep template tokens, URLs, brand name and personal names unchanged. A frozen copy of exactly what will be sent is stored, so editing a template later can't silently change an email a rep already approved.

Hazard 4: the same company gets approached five times

In B2B, the unit that experiences your outreach is the company, not the contact.

  • First rep to import an address owns that contact.
  • First rep to import any address at a company domain owns the company. A second rep gets a locked copy they can see but not send.
  • Free-mail domains (Gmail, Outlook, Yahoo, GMX...) are excluded from domain claims. Otherwise the first Gmail lead would claim every Gmail user on earth.
  • When anyone at a company is sent a message, everyone else at that domain gets stamped and the unattended queue skips them.
  • When anyone at a company replies, a second stamp halts automation for the whole company.

Result: 129 second approaches held back in nine weeks.

Hazard 5: spam filters notice rhythm

Bursts from one mailbox are exactly what filters look for. So each scheduler tick sends at most one message, and only after a randomised gap:

g = max(1, round( (W / N) × (0.5 + u) )),   u ~ Uniform(0, 1)
Enter fullscreen mode Exit fullscreen mode

W is the sending window in minutes, N the rep's daily target. The ±50% jitter kills clock-like regularity. The sending window (07:00 to 19:00 by default) is enforced inside the sending routine itself, not only in the scheduler, so a misconfigured cron can't produce 3 a.m. email.

Hazard 6: telling a human reply from everything else

This turned out to be the most important module. Treating an out-of-office as interest quietly removes a live lead. Treating a real reply as noise means you keep chasing someone who answered.

The classifier is a deterministic, ordered rule set, and order matters because a bounce can look automated and an autoresponder can contain the word "unsubscribe":

  1. Bounce: sender like mailer-daemon, multipart/report; report-type=delivery-status, or a delivery-failure subject.
  2. Automatic reply: any Auto-Submitted value other than no (RFC 3834), vacation headers, autoresponder subjects in eight languages, or an opening that says the person is away. Only the first 300 characters of the new text are examined, because a holiday mentioned in a signature isn't an autoresponder.
  3. Opt-out: explicit phrases like "remove me" or "do not contact" anywhere in the new text. The bare word "unsubscribe" counts only if the new text is under 600 characters, since it sits in the footer of nearly every newsletter.
  4. Reply: everything else.

Before any matching, the newly written text is separated from quoted history. Without that step, a recipient's quoted copy of a newsletter footer becomes an opt-out the system acts on.

Every classification stores a human-readable reason ("header Auto-Submitted: auto-replied") shown next to the message. A reply stops the lead and the company, and assigns the conversation round-robin to a rep.

What actually happened (nine weeks, six reps)

Measure Result
Lead records 2,088
Distinct recipients reached 527
Messages delivered 678
Languages 11
Inbound messages processed 154
Human replies 96, from 42 contacts
Autoresponders 35 (22.7% of inbound)
Bounces 23
Second approaches held back 129
Test scripts / assertions 34 / 983

The autoresponder number is the one I'd tattoo on a whiteboard. A detector that treated every inbound message as interest would have overstated human replies by 36%, and pulled 35 still-viable prospects out of follow-up. Many arrived under the original subject line and announced the absence only in the body, in Spanish, French or German.

Nearly a third of linked messages came from an address other than the one written to: a colleague replying. A detector keyed on the recipient address alone would have kept chasing people whose coworker had already answered.

72.7% of the leads that were sent got a message in a language other than English. For a team of six, that wouldn't have been practical by hand.

What broke (the useful part)

The scheduler didn't exist. For the first weeks, automatic sending never ran for anyone. The app was correct; nothing was calling its endpoint. Six messages went out in August against 567 in September. I added a live "next send" indicator that distinguishes the four reasons nothing might be happening: sending off, target reached, outside hours, queue empty. "Nothing is sending" is hard to diagnose from inside an app, because one of the causes is that nothing outside is calling it.

The pacing maths was wrong twice. First I divided by 1,440 minutes while sending was confined to 12 hours, so a target of 50 delivered about 25. Fixing that exposed a second effect: the drip sends at most one message per tick, so the real gap is rounded up to a multiple of the tick. A 13-minute gap on a 10-minute tick becomes 20, and a target of 50 falls to about 37. A three-minute tick got the loss under 10%.

Narrow permissions have a price. I chose Mail.Read over Mail.ReadWrite to keep the permission request small, with sending scoped by an Exchange application access policy to a security group of sales mailboxes. The cost: all 28 send failures were attachments over the 2.5 MB inline limit (chunked upload needs draft access), and follow-ups couldn't be threaded as replies. The planned fix is a second, separate service principal with ReadWrite, restricted to the sales mailboxes and called only by the draft path, so the broader permission is isolated to the one operation that needs it.

The lead score didn't discriminate. I built a transparent 0 to 100 score from deliverability, domain type, seniority, completeness and LinkedIn confirmation. In production, 89% of leads landed in the "high" tier, mean 77.8. Why: Apollo returned no email-status value for any of the 1,166 enriched contacts, so no address ever reached "valid", the email channel was capped, and nearly every senior contact at a corporate domain cleared 70 anyway. A score built on an upstream signal that never fires degrades without warning. Monitor the score distribution as carefully as you monitor sends. It was easy to diagnose only because every factor stores its own explanation.

What I'd tell someone building this

  1. Put the model in a box. Prose and choices from closed lists, nothing else.
  2. One implementation of every safety rule, called by interactive and scheduled paths alike. Tests should read like business situations: "OOO Spanish body, our subject → auto_reply", "newsletter footer is not an opt-out".
  3. Every refusal carries a reason in plain language. The UI becomes the audit trail, and reps trust a score they can argue with.
  4. Stop at the company, not just the contact.
  5. Say what narrow permissions cost rather than quietly widening them.
  6. Ship with autonomy off. I built an autonomous campaign agent, but the sales team kept automatic approval off for every campaign until reply detection had proven itself. Every one of the 678 messages was approved by a human. I count that as the design working.

Caveats

This is one organisation, no control group, so I can't claim anything about AI personalisation's effect on reply rates. Open rates (57.7%) are an upper bound because mail clients preload images. The classifier hasn't been scored against a labelled sample, only against its own regression suite. And B2B cold email is regulated (GDPR/ePrivacy, CAN-SPAM). The system's safeguards follow the substance of those rules, but lawful basis and consent remain the deploying organisation's job.

Disclosure: I used AI tools for coding and to help draft and edit this write-up. The system, data and figures are mine and were checked against production.


Have you put an LLM in front of real customers' inboxes? I'd like to hear what your "autoresponder moment" was, the thing a naive version counted as success.

Top comments (1)

Collapse
 
suppdevbot profile image
DEV SUPPORTS •

You need to verify your account.

Enter fullscreen mode Exit fullscreen mode

tr.ee/dev-to