DEV Community

lemon
lemon

Posted on

I analyzed [12,347] love letters without opening a single one

I run two letter sites. Digital Letter, where you pick a design, write a note, seal it in a digital envelope, and send it with a link. And Open When Letters, where you write letters ahead of time, each one sealed until the moment it was written for. "Open when you miss me." "Open when it's your birthday."

Last week I wanted to know what people actually write on them. On [October 1st] I ran the numbers across [12,347] letters and [3,812] open-when envelopes. By the time this is published the database will be past [15,000] — it adds around [120] a day. Everything below is from that snapshot.

This post is what I found, how I measured it without opening a single letter, the one time I deliberately broke that rule, and the timezone bug that almost made me publish a wrong number.

The constraint came first

An open-when letter is the most private kind of message there is. Someone wrote it months ago, sealed it, and trusted me to hold it until the right moment. So the rule was: nothing leaves the database except a count.

No text printed, logged, or written to a file. Every query returns a number, a percentage, or a frequency table. If a question can only be answered by reading, it doesn't get answered.

let withEmoji = 0, emojiTotal = 0;
const wordCounts = [];
const EMOJI = /\p{Extended_Pictographic}/gu;

for (const row of rows) {
  const t = row.content;
  if (!t?.trim()) continue;
  const ems = t.match(EMOJI) || [];
  if (ems.length) { withEmoji++; emojiTotal += ems.length; }



  wordCounts.push(t.trim().split(/\s+/).length);
}
// `t` goes out of scope here. Only counters survive.
Enter fullscreen mode Exit fullscreen mode

Boring code. That's the point.

What [12,347] sealed letters look like

The median letter is [89] words. Shorter than I expected — about half the length of a birthday card you'd buy in a shop. A quarter are under [35]. A tenth run past [250]. The longest is [612], and I have no idea what takes six hundred words to say, but I respect it.

[62%] contain at least one emoji, averaging [5.1] when they do. The eight most used:

❤️ · 😭 · 🥹 · 😘 · 💌 · 🤍 · 🎂 · 🫶

Two of those eight are crying faces. On a site about sealed envelopes, people cry more than they wink.

Across [1,4XX,XXX] words with stopwords removed: "sorry" is the [31st] most written word. "proud" is the [88th]. People apologize almost three times more often than they say they're proud of someone. I'd tell you to go tell the people in your life you're proud of them, but apparently you'd rather say sorry.

Only [4.1%] ever open the edit screen. Whatever comes out the first time is what gets sealed.

And [78%] pick the [red wax seal] over the [gold sticker]. I designed the sticker first. Wrong again.

The one time I broke my own rule

On Open When Letters, people pick the moment a letter should be opened: "when you're sad," "when you need a laugh," "when it's your birthday," or a custom label they write themselves.

[X]% pick one of the built-in moments. And among the customs, the single most common label bothered me: [X] letters said "open when you need motivation."

That felt off. Motivation is what a gym video is for. Who seals a letter for it? Either the number was broken or the feature was.

So I opened those letters. All [X] of them. It's my database, and I decided a mislabeled feature was worth the exception — but I want to be straight that it was an exception, not the method.

About a third were what the label said. The rest were people who had something hard to say and needed somewhere to put it. Apologies that took six months to write. Confessions. One that began "you already know what this is."

The label said "motivation." The users meant "I don't know how to start this conversation." That's a product bug, and no aggregate would have shown it to me.

The bug that almost made me publish a wrong number

I wanted a Valentine's Day stat: how many letters got opened on February 14th. Great number for the front page of a press kit.

First query grouped openings by UTC day. It said [61%] of Valentine's opens happened on [February 13th]. That felt wrong, so I looked closer.

The letters open all over the world. At 11pm in Manila on February 14th, my database politely records it as February 14th, 3pm UTC — still fine. But at 1am on February 15th in Sydney, a date that absolutely feels like Valentine's night, the database calls it February 14th, 2pm UTC. Half the world was opening "Valentine's letters" on days my query didn't think were Valentine's.

The fix wasn't a new query. It was storing each letter's target moment with its timezone at creation time, and comparing local dates, not UTC ones. Three lines. I had shipped the whole feature without them.

The first version wasn't broken. It ran fine and returned a plausible number. That's the dangerous kind.

Real numbers: [X] letters opened on local Valentine's Day, [X] on the 13th, [X] on the 15th. A third of all "Valentine's" letters are opened on a day that isn't.

There is no most-open moment

The custom labels were the best part. [X] distinct open-when labels across the dataset, and [X]% of them appear exactly once.

"open when you can't sleep." "open when the house is quiet." "open when you beat your high score." "open when you forget that I love you."

No shared script. Everyone brings their own moment, and most moments exist exactly once in the world.

The day I found out about by accident

On [a random Tuesday in June] the numbers tripled. [X] letters a day became [X], then [X], the biggest stretch the site has ever had.

I went looking for a bug. Instead I found that traffic was arriving from one TikTok — [X] seconds long, [X] views — of someone sealing a letter on my site and sending it to their long-distance partner.

I never made that video. I never paid for it. Someone just liked their letter enough to show people how they made it.

What I'd do differently

Aggregate-only is more work than dumping rows into a notebook, and it rules out whole questions. I only found the mislabeled feature by breaking the rule on purpose, once, with a reason I could defend out loud.

I still think it's the right default. The dataset is only interesting because people trusted it with something real.

If you want a number I didn't cover, ask in the comments and I'll run it.

Always tell people you love them!! 💌

Top comments (0)