DEV Community

Daniel Pertu
Daniel Pertu

Posted on

Our activation code has no O, no 0, no I and no 1, and 32 is what makes the maths safe

Notifio is a desktop app that watches rental search pages and tells you when a new listing appears. It is a one-time £20 purchase and it has no accounts, no sign-up, no password. After checkout you get an email with an activation token, you type it into the app, and that is the whole of our identity system.

Which means the token is not only a credential. It is a piece of user interface that somebody has to read off a screen, in a monospace font, and retype into a desktop app, sometimes months later on a new laptop after a reinstall.

Here is the whole generator:

import crypto from "crypto";

const ALPHABET = "ABCDEFGHJKLMNPQRSTUVWXYZ23456789";

function randomSegment(length: number): string {
  const bytes = crypto.randomBytes(length);
  return Array.from(bytes)
    .map((b) => ALPHABET[b % ALPHABET.length])
    .join("");
}

export function generateToken(): string {
  return `NTFIO-${randomSegment(4)}-${randomSegment(4)}-${randomSegment(4)}`;
}
Enter fullscreen mode Exit fullscreen mode

Twenty lines, and three separate decisions are hiding in them.

The characters that are not there

Count the alphabet: 32 characters. The full uppercase set plus digits would be 36. Four are missing: I, O, 0, 1.

Those four are the entire transcription error budget. In most fonts a capital I and the digit 1 are a serif apart, and O and 0 differ by a slash that plenty of mail clients do not render. If a token can contain them, a support email that says "it says invalid token" is unanswerable, because neither of you can tell which character was mistyped.

The rest of the ambiguous pairs survive because they do not actually collide once case is fixed. We uppercase everything, so there is no l-versus-1 problem to have.

32 is also the number that makes the maths correct

b % ALPHABET.length is the obvious way to turn a random byte into a character, and it is usually the wrong way, because it skews the distribution. A byte is uniform over 0 to 255. Taking it modulo n is only uniform when n divides 256.

32 divides 256 exactly, eight times over. Every character in the alphabet is produced by exactly eight of the 256 byte values, so the output is uniform and the shortcut is safe.

Had we kept the letter O and gone to 33 characters, 256 bytes split as 7 × 33 + 25: twenty-five characters would come out eight times each and the other eight only seven times. That is a 14% relative skew in favour of most of the alphabet.

Does a 14% skew break a 60-bit token? No. Twelve characters at five bits each is 60 bits of entropy, and the licence table has a unique constraint on the column, so the failure mode of the skew is theoretical. But the fix costs nothing and the bug is not self-announcing, which is exactly the kind of thing to get right by construction instead of by argument. If you ever need a non-power-of-two alphabet, reject the bytes above the largest multiple of n and draw again. Do not reach for % and hope.

Both ends normalise, in the same two calls

The app's activation screen does this before it posts:

const trimmedEmail = email.trim();
const trimmedToken = token.trim().toUpperCase();
Enter fullscreen mode Exit fullscreen mode

And the API route does it again:

email = (body.email ?? "").trim().toLowerCase();
token = (body.token ?? "").trim().toUpperCase();
Enter fullscreen mode Exit fullscreen mode

The duplication is deliberate. The client normalises because the user should never see "invalid token" for a trailing space they cannot see. The server normalises because the client is not the only thing that can POST to a public endpoint, and because a lowercase token from some future paste path must not create a licence lookup miss that looks like a revoked licence.

The input also tells the user the shape before they type anything:

placeholder="NTFIO-XXXX-XXXX-XXXX"
Enter fullscreen mode Exit fullscreen mode

A credential with a visible shape is a credential people can proofread. That is worth more than it sounds when the alternative is a bare UUID.

The token is generated once and then never again

The other half of treating a token as a UI object is accepting that it outlives everything around it. People keep the email. People reinstall. So a token is issued exactly once per licence and every other path reuses it.

In the Stripe webhook, three different situations all resolve to an existing token rather than a new one:

if (bySession) {
  // This exact session was already processed, so this is a retry.
  token = bySession.token;
} else {
  const [byEmail] = await db.select().from(licenses)
    .where(eq(licenses.email, email)).limit(1);

  if (byEmail) {
    // Duplicate purchase with a *different* session id. Reuse the existing
    // license instead of inserting (which would violate UNIQUE(email)).
    token = byEmail.token;
  } else {
    token = generateToken();
    try {
      await db.insert(licenses).values({ /* ... */ });
    } catch (err) {
      // Concurrent webhook won the race and inserted first. Fall back to
      // the row that won so we still deliver a valid token.
    }
  }
}
Enter fullscreen mode Exit fullscreen mode

A webhook retry, a second purchase from the same address, and two webhook deliveries racing each other all end with the customer being emailed the token they already have. The only place a token is born is the one branch where no row exists.

This is the same reasoning as our script for granting free licences, which I wrote about in Our free-licence script reuses your token, because the token is not ours to regenerate. What this post adds is where the token comes from in the first place, and why its alphabet is shaped the way it is.

What the token cannot tell an attacker

The token is a bearer credential paired with the purchase email, and the lookup refuses to be an oracle:

if (!license || license.email !== email || !license.active) return null;
Enter fullscreen mode Exit fullscreen mode

One null for three different failures. Wrong token, right token with the wrong email, and a deactivated licence are indistinguishable from outside, so nobody can use the validate endpoint to confirm that an email address bought the app. The endpoint is also rate limited, and a 429 is deliberately not treated by the app as "your licence is bad", which I went through in A 429 is not a no.

The rule worth taking away

If a human has to transcribe your secret, its format is a product decision, not a cryptography decision. Entropy is the easy part. Which characters are in it, who normalises the input, and whether the value is stable across a reinstall are the parts your support inbox will measure you on.

See it working

The purchase side is on notifio.app/pricing, and the app you type the token into is on notifio.app/download for Mac and Windows. The activation steps, including what to do if the email never arrives, are on notifio.app/help.

If you want the product context for why there is no account system to begin with, notifio.app/about is the short version, and the per-site pages under notifio.app/alerts show what the app is actually doing with that licence once it is activated.

Top comments (0)