Nakodo sends outreach email for a brand and reads the replies. Replies arrive through a provider's inbound webhook, which hands us a JSON object with, among other things, text and html. The old line that decided what somebody had said was this:
return { email, text: email.text ?? (email.html ? textFromHtml(email.html) : "") };
Read it out loud and it sounds right. Use the plain text part; if there is not one, render the HTML down to text; if there is neither, empty string. The comment above it even said so.
Then a reply came in whose stored text began <!doctype html>.
What the sender sent
Plenty of mail tools build HTML email from a template framework and produce a multipart message where the text alternative is not a plain-text version of the message. It is the same HTML, byte for byte. Some builders do it by accident, some do it because a template has no text version and the sender's client fills the slot with whatever it has.
Here is the shape, cut down from a real one. Note the MJML scaffolding and the Outlook conditional comments:
<!doctype html>
<html xmlns="http://www.w3.org/1999/xhtml">
<head>
<title></title>
<!--[if !mso]><!--><meta http-equiv="X-UA-Compatible" content="IE=edge"><!--<![endif]-->
<style type="text/css">p { display:block;margin:13px 0; }</style><!--[if mso]>
<xml><o:OfficeDocumentSettings><o:PixelsPerInch>96</o:PixelsPerInch></o:OfficeDocumentSettings></xml>
<![endif]-->
</head>
<body style="word-spacing:normal;">
<!--[if mso | IE]><table align="center"><tr><td><![endif]-->
<div><p>Hi,</p><p>Looks like you’re not on Mosaic yet.</p></div>
</body>
</html>
email.text was 594 bytes of that. ?? only falls through on null and undefined, and this was a non-empty string, so the HTML branch never ran. We stored the markup as the reply, showed it in the app, and quoted it into the notification email that tells the brand their reply needs an answer.
The person had written nine words. The customer got a screenful of OfficeDocumentSettings.
The fix is a sniff, not a parser
We do not need to know whether a string is valid HTML. We need to know whether treating it as text would be a mistake:
// Some senders' mail tools put the whole HTML document in the text part as
// well, so it can't be trusted to be text.
const HTML_DOCUMENT = /^\s*<(!doctype|html|head|body|meta|table|div|p)\b/i;
// What an email says: its text part, or its HTML read as text when it has no
// text part or the text part is really HTML.
export function emailText(text: string | null | undefined, html: string | null | undefined): string {
if (text?.trim() && !HTML_DOCUMENT.test(text)) return text;
const source = html?.trim() ? html : (text ?? "");
return textFromHtml(source);
}
Three things that regex is doing on purpose.
It is anchored at the start, after optional whitespace. A message that merely mentions HTML, or includes a code sample halfway down, is still a text message and is returned untouched.
It requires a word boundary after a known element name. This is the test case that stops the fix from being worse than the bug:
assert.equal(emailText("<3 love it", null), "<3 love it");
<3 is not a tag. A naive "starts with <" check would have mangled every reply from anybody who types affection with a keyboard.
And when the text part is HTML, the HTML part is still preferred as the source, falling back to the text part only if the HTML one is missing. If a sender puts markup in both, the html field is the one more likely to be complete.
The hard part was keeping the line breaks
Reducing HTML to text is where it gets interesting, because the output feeds something that reads it line by line.
Replies carry quoted history. Before a reply is stored or quoted back to the customer, it goes through stripQuotedReply, which cuts at unambiguous markers: a line matching On ... wrote:, a ---- Original Message ---- separator, a run of > lines that reaches the bottom. Every one of those tests is ^...$ against a trimmed line.
So if your HTML-to-text converter flattens everything into one long line, the quote trimmer silently stops working. Not crashes: stops. Every reply keeps its entire thread history, which then gets quoted into the next notification, which grows every round.
That is why the converter inserts breaks before it strips tags, and in this order:
export function textFromHtml(html: string): string {
const text = html
.replace(/<!--[\s\S]*?-->/g, "")
.replace(/<(head|script|style|title)\b[\s\S]*?<\/\1\s*>/gi, " ")
.replace(/\s+/g, " ")
.replace(/<br\b[^>]*>/gi, "\n")
.replace(/<\/(p|h[1-6]|table|blockquote)\s*>/gi, "\n\n")
.replace(/<\/(div|tr|li)\s*>/gi, "\n")
.replace(/<[^>]+>/g, " ");
return decodeHtml(text)
.split("\n")
.map((line) => line.replace(/\s+/g, " ").trim())
.join("\n")
.replace(/\n{3,}/g, "\n\n")
.trim();
}
Comments go first. Outlook-only markup lives inside <!--[if mso]> blocks, and a converter that strips tags without removing comments leaks OfficeDocumentSettings and bare <xml> content into the body. The head, scripts, styles and title go next, because CSS rules are text too and nobody wants display:block;margin:13px 0; in a quoted reply.
Then the surprising one: whitespace is collapsed before the breaks are inserted. The newlines in the source are formatting of the markup, not of the message, and keeping them produces a ragged mess with paragraph structure that does not match what the recipient saw. Collapse first, then let <br> make a line and </p> make a paragraph. The result is the structure the reader saw, not the structure the generator emitted.
Here is what that gets you, straight from the test file:
assert.equal(
textFromHtml('<div dir="ltr">Hi there,\n sounds good<div><br></div><div>Thanks,<br>Sam</div></div>'),
"Hi there, sounds good\n\nThanks,\nSam",
);
assert.equal(emailText(TEMPLATE, TEMPLATE), "Hi,\n\nLooks like you’re not on Mosaic yet.");
594 bytes in, 41 characters out, and the apostrophe comes back as a real one because entity decoding happens after the tags are gone.
The quote trimmer then works on HTML-only replies the same way it works on plain ones:
const html =
'<div dir="ltr">Sounds good, what\'s the fee?</div><br><div class="gmail_quote"><div dir="ltr" class="gmail_attr">' +
"On Mon, 5 Oct 2026 at 09:00, Acme <acme-abc@intros.nakodo.app> wrote:<br></div>" +
"<blockquote>Hi Sam, we'd love to send you our kettle.</blockquote></div>";
assert.equal(stripQuotedReply(emailText(null, html)), "Sounds good, what's the fee?");
That test is the one worth having. It proves the two functions compose, which is the property that broke.
Why not an html-to-text library
A dependency would handle more cases than nine regexes: nested lists, tables with meaningful structure, links turned into footnotes. We did not take one, for a reason specific to this code path rather than a general principle.
This function runs inside a webhook handler on every inbound email, including the pitches a published contact address attracts, and its output has exactly two consumers: a human reading a quoted reply, and a classifier deciding whether an answer is a yes. Neither needs a faithful rendering. Both need the line structure preserved and the markup gone. Nine regexes and 13 unit tests cover that, and they are 40 lines we can read in a minute when a new sender shows up with new nonsense.
If the requirement ever becomes "render this email properly", the library is the right call and these regexes are the wrong one. They are sized for the job they have.
The transferable bit
a ?? b asks whether a exists. Most of the time the question you actually have is whether a is usable. Any time you are pulling a field out of somebody else's payload and the field has a stated type, it is worth asking what you would do if the value were the right type and the wrong kind of thing. Inbound email is the richest source of those you will ever integrate: every field is attacker-controlled or client-controlled, and the clients are 30 years of mail software.
A couple of live pieces of the same machinery, if you want to see where this ends up. The domain's MX record is a catch-all into the provider, which you can check yourself:
$ dig +short MX nakodo.app
10 inbound-smtp.us-east-1.amazonaws.com.
And every outreach email carries an opt-out link to a token page that asks before it acts, because mail security scanners open links in email and a one-click unsubscribe would be triggered by a robot: nakodo.app/o/abcdefghjkmn is a live one with a made-up token. The reply handling itself is described without the code on how it works.
Top comments (0)