DEV Community

Mr.Rain
Mr.Rain

Posted on

Validate LLM Output: Catching AI-Invented Emoji

emoji tokens flowing left to right through a glowing validation gate, flat illustration

If an LLM generates structured output for your product, any string it emits can be fiction, and "looks right" is not validation. On my site emoji-combos.net, a language model drafts decorative emoji combos for bio text, and it kept producing emoji that don't exist in the Unicode standard: valid-looking sequences that no platform actually supports. The fix is deterministic and cheap: a Set of every real emoji sequence, Intl.Segmenter to split strings into user-perceived characters, and a pipeline that fails closed when validation rejects too much. This post walks through all three, with the code.

We'll cover:

  • Why LLMs hallucinate emoji (and why a better model is not the answer)
  • Why [...str] and str.split('') break emoji, and the correct segmentation
  • Building the ground-truth dictionary from Unicode's official data
  • The full validator: skin tones, ZWJ sequences, invisible characters
  • Pipeline policy: batch minimums, dedupe, fail-closed behavior
  • Two gotchas that cost me an evening each

The stack is TypeScript, but nothing here is stack-specific. The concepts port to Python, Go, or wherever your AI feature lives.

The feature and the lie

The pipeline behind this is deliberately boring. When a visitor searches a "vibe" the site hasn't curated, the server calls an OpenAI-compatible chat endpoint (DeepSeek in production, any /chat/completions provider works) with JSON mode on:

response = await fetch(`${base}/chat/completions`, {
  method: 'POST',
  headers: { authorization: `Bearer ${key}` },
  body: JSON.stringify({
    model: 'deepseek-chat',
    messages: [
      { role: 'system', content: SYSTEM_PROMPT },
      { role: 'user', content: `Generate 10 emoji combos for the vibe ${JSON.stringify(query)}.` },
    ],
    response_format: { type: 'json_object' },
    max_tokens: 4000,
    temperature: 1.2,
  }),
  signal: AbortSignal.timeout(20_000),
});
Enter fullscreen mode Exit fullscreen mode

The system prompt asks for {slug, name, description, combos[]}: ten short decorative strings for social bios. Temperature runs at 1.2 because variety is the product.

And the model lies. Not always, and less with each model generation, but reliably enough that shipping on vibes alone is malpractice. The lies take a specific shape: emoji-shaped sequences that aren't in the Unicode standard. A "family" with zero-width joiners in the wrong order. A skin-tone modifier glued to a base that doesn't accept one. Two emoji fused into something no consortium ever approved.

This isn't a malfunction. An LLM predicts plausible token sequences; a plausible sequence that's fake is a successful prediction by its own lights. Emoji make the problem unusually crisp because validity has a public ground truth: Unicode's emoji-test data defines exactly which sequences are real. That turns "is this AI output acceptable?" from a judgment call into a lookup.

That's the position you want to be in with any AI feature. A registry exists for your domain (Unicode, IANA, ISO codes, your product database), so the model proposes and the registry disposes.

Step 1: Split strings like a human sees them

Before validating anything, you need the right unit. This is where nearly everyone starts wrong, including me.

Emoji are multi-code-point. A family emoji like ๐Ÿ‘จโ€๐Ÿ‘ฉโ€๐Ÿ‘งโ€๐Ÿ‘ฆ is man + ZWJ + woman + ZWJ + girl + ZWJ + boy, where ZWJ (U+200D) is an invisible joiner. The two classic JavaScript string-splitting idioms both operate below that level:

// Splits by UTF-16 code units: destroys ANY emoji with a surrogate pair
const byCodeUnit = '๐ŸŽ€๐Ÿซงโœจ'.split('');

// Splits by code points: keeps solo emoji, SHREDS joined sequences
const byCodePoint = [...'๐Ÿ‘จโ€๐Ÿ‘ฉโ€๐Ÿ‘งโ€๐Ÿ‘ฆ'];
// ['๐Ÿ‘จ', '\u200D', '๐Ÿ‘ฉ', '\u200D', '๐Ÿ‘ง', '\u200D', '๐Ÿ‘ฆ'] โ€” 7 fragments
Enter fullscreen mode Exit fullscreen mode

Code-point splitting sounds like the fix and demos fine on ๐ŸŽ€. Then it quietly dismantles every ZWJ sequence into components that are individually valid emoji but collectively wrong, and your validator flags good output as garbage (or worse, passes fragments through as "known" and reassembles nonsense).

The correct tool is Intl.Segmenter, which has been baseline in Node and browsers for a while now and asks the platform's own ICU for user-perceived characters:

const segmenter = new Intl.Segmenter('en', { granularity: 'grapheme' });

export function splitGraphemes(content: string): string[] {
  return [...segmenter.segment(content)].map((s) => s.segment);
}

splitGraphemes('๐ŸŽ€๐Ÿ‘จโ€๐Ÿ‘ฉโ€๐Ÿ‘งโ€๐Ÿ‘ฆโœจ');
// ['๐ŸŽ€', '๐Ÿ‘จโ€๐Ÿ‘ฉโ€๐Ÿ‘งโ€๐Ÿ‘ฆ', 'โœจ'] โ€” 3 units, family intact
Enter fullscreen mode Exit fullscreen mode

One API call, and the unit of validation matches the unit of meaning. Everything downstream now operates on graphemes, never on spread results.

If you're in Python, the equivalent is grapheme or a regex with \X. The language doesn't matter; the principle does: validate at the granularity your user perceives, or your checks measure the wrong thing.

Step 2: Build the dictionary from the source of truth

Unicode publishes the complete list of valid emoji sequences in the emoji-test.txt file. Every line describes one sequence the standard considers fully-qualified emoji. That file is your registry.

I wrote a small import script that parses it into a JSON array of sequence strings, checked into the repo as src/data/emoji/sequences.json, 3,564 of them at the time of import. Regenerating is one command when Unicode ships a new revision.

At runtime it becomes a Set, because membership checks are O(1) and the whole thing fits in memory comfortably:

import sequences from '@/data/emoji/sequences.json';

const KNOWN_SEQUENCES = new Set<string>(sequences);
Enter fullscreen mode Exit fullscreen mode

Three properties of this approach deserve emphasis:

  1. It's complete by construction. Whatever the model invents, if it isn't in the file, it isn't an emoji. No fuzzy matching, no embedding similarity, no second model scoring the first.
  2. It's auditable. When validation rejects something, "why" has a one-line answer: not in Unicode's list. Try getting that clarity from a moderation model's score.
  3. It's swappable. New Unicode revision โ†’ re-run the import. The validation contract never changes.

Note what we did not do: ask the LLM to "only use valid emoji" harder. Prompt constraints are suggestions; a Set.has() is a fact. Use prompts to steer quality (variety, style, tone) and code to enforce correctness. Conflating the two is how teams end up with prompt paragraphs doing a schema's job.

Step 3: The validator

Here's the whole thing, lightly trimmed from production:

// Skin tone modifiers are valid components, but only glued to a person emoji.
// Matching the tone-stripped form covers every mixed-tone combination
// (๐Ÿง‘๐Ÿปโ€๐Ÿคโ€๐Ÿง‘๐Ÿฟ) without enumerating hundreds of variants.
const TONE_MODIFIERS = /[\u{1F3FB}-\u{1F3FF}]/gu;
const EXTENDED_PICTOGRAPHIC = /\p{Extended_Pictographic}/u;

// Zero-width and bidi control characters are never legitimate in a combo:
// platforms use them for mention/tag injection tricks.
const INVISIBLE = /[\u{200B}-\u{200F}\u{202A}-\u{202E}\u{2060}-\u{206F}\u{FEFF}]/u;

const segmenter = new Intl.Segmenter('en', { granularity: 'grapheme' });

export interface ComboValidationResult {
  /** Graphemes that are emoji-like but not in the official dictionary. */
  invalidEmoji: string[];
  /** Zero-width / bidi control characters are present. */
  hasInvisible: boolean;
}

export function isKnownEmoji(grapheme: string): boolean {
  return KNOWN_SEQUENCES.has(grapheme)
    || KNOWN_SEQUENCES.has(grapheme.replaceAll(TONE_MODIFIERS, ''));
}

export function splitGraphemes(content: string): string[] {
  return [...segmenter.segment(content)].map((s) => s.segment);
}

export function validateComboContent(content: string): ComboValidationResult {
  const invalidEmoji = new Set<string>();
  let hasInvisible = false;

  for (const grapheme of splitGraphemes(content)) {
    if (INVISIBLE.test(grapheme)) {
      hasInvisible = true;
    }

    // A cluster built only from tones is junk: presentation needs a base.
    const toneStripped = grapheme.replaceAll(TONE_MODIFIERS, '');
    const toneOnlyCluster = toneStripped !== grapheme && toneStripped.trim() === '';
    const looksLikeEmoji = EXTENDED_PICTOGRAPHIC.test(grapheme) || toneOnlyCluster;
    if (looksLikeEmoji && !isKnownEmoji(grapheme)) {
      invalidEmoji.add(grapheme);
    }
  }

  return { invalidEmoji: [...invalidEmoji], hasInvisible };
}
Enter fullscreen mode Exit fullscreen mode

Four things are happening, and each one earns its keep:

Skin-tone stripping. A handshake between people with different skin tones is a single grapheme, but Unicode doesn't enumerate every tone mix in emoji-test. Stripping modifiers before the lookup collapses all variants onto a base sequence that is in the dictionary.

Extended_Pictographic as the "looks like emoji" test. The Unicode property \p{Extended_Pictographic} matches emoji-ish characters broadly, broader than the official emoji set. That asymmetry is exactly the hallucination signature: emoji-shaped, but not in the registry. Plain text (letters, digits, punctuation) fails the property and passes through for length/content rules elsewhere; fake emoji trip it and get checked against the Set.

The tone-only cluster guard. After segmentation, a cluster of only tone modifiers (sometimes glued to a stray space) is garbage with no base character. It doesn't match Extended_Pictographic on its own, so the explicit check catches it.

Invisible character rejection. Zero-width joiners, word joiners, and bidi controls are legitimate in some writing systems, which is precisely why they're a documented trick for splitting @mentions in pasted bios. For a site whose output gets pasted into profile fields on other platforms, this is a real attack surface, not paranoia. The regex rejects the lot.

Per combo, all of this runs in microseconds. The LLM call it guards costs five orders of magnitude more. Guard rails are cheap; use them everywhere.

Step 4: Pipeline policy, the part that makes it a system

A validator is a function. A pipeline is a policy about what happens to the function's answers. Ours, in production:

Request 10, require 4. Each generation asks for ten combos. If fewer than four survive validation, the entire batch fails. No "serve whatever made it." A 60%+ rejection rate means something upstream is wrong (bad prompt drift, a model regression), and shipping the survivors means shipping a biased sample of a broken process.

const valid = filterValidCombos(generated.combos);
if (valid.length < MIN_VALID_COMBOS) {
  throw new AiGenerationError('generation produced too few valid combos');
}
Enter fullscreen mode Exit fullscreen mode

Dedupe at the database, not just in memory. The in-memory filter drops exact duplicates within a batch, but the model also regurgitates combos that already exist on a vibe. A unique constraint at the persistence layer is the last line of defense, and unique-violation errors during insert are tolerated as "already have it," not propagated as failures:

try {
  await repository.createWithTags({ combo: { /* ... */ }, tagIds: [tag.id] });
} catch (error) {
  // A combo that already exists must not sink the batch.
  if (!isUniqueViolation(error)) throw error;
}
Enter fullscreen mode Exit fullscreen mode

Fail closed on the provider. No API key configured? The route returns 503 and the frontend shows its normal empty state. Provider down or timing out? Same. The failure mode of an AI feature should be "honestly unavailable," never "confidently degraded."

Cache past the model. Once a vibe has six approved combos stored, generation is skipped entirely for that query. The model is the cold-start path, not the hot path. (This also means validation made the feature cheaper: trustworthy stored output is worth serving, and serving it skips the expensive, hallucination-prone step.)

Human gate on permanence. Generated vibes are created as drafts, excluded from public listings and the sitemap until an admin promotes them from a review queue. The code validates correctness; the human validates taste (variety, on-theme-ness, non-cringe descriptions). Visitors who triggered a generation see their results immediately; only permanence is gated.

Abstract pipeline: three connected stations with emoji stickers traveling along a line, one station carrying a shield

Testing the gate

The repo uses fast-check for property tests, and the validator is a property-test target from heaven: for any string, validation should never accept an emoji-ish grapheme absent from the dictionary. Sketch:

import fc from 'fast-check';

test('rejects every non-dictionary emoji-shaped grapheme', () => {
  fc.assert(
    fc.property(
      fc.array(fc.constantFrom('๐ŸŽ€', 'โœจ', '๐ŸŽƒ', '\u200D', '\u{1F3FF}', 'x'), {
        minLength: 1,
        maxLength: 20,
      }),
      (parts) => {
        const input = parts.join('');
        const result = validateComboContent(input);
        // Any emoji-looking grapheme in the result must be a known sequence.
        for (const g of splitGraphemes(input)) {
          if (/\p{Extended_Pictographic}/u.test(g) && !KNOWN_SEQUENCES.has(g)) {
            expect(result.invalidEmoji).toContain(g);
          }
        }
      },
    ),
  );
});
Enter fullscreen mode Exit fullscreen mode

Generative testing shines here because the input space (all strings) is infinite and the property is crisp. It caught the tone-only cluster case during development, which I had not thought of manually.

Two gotchas worth your evening

JSON mode is not a fidelity guarantee. Even with response_format: { type: 'json_object' }, models occasionally wrap the payload in markdown fences or lead with a chatty sentence. Parse tolerantly: try JSON.parse on the trimmed body; on failure, slice from the first { to the last } and parse that. And validate the shape (string slug, string array of combos) before touching values, because payload.combos being null is not an exception until you make it one.

Reasoning tokens eat your completion budget. On some providers, hidden reasoning happens inside the same max_tokens budget as the visible JSON. A ceiling that's generous for ten emoji strings can be too small for ten emoji strings plus the model's private scratchpad, and the symptom is not an error but an empty completion, intermittently, which looks exactly like a network flake. If your structured calls come back empty on a reasoning-flavored model, raise the ceiling before you blame the network. (The system prompt in my pipeline also asks the model to reuse an existing vibe slug when the query is a synonym, "cute" vs "ๅฏ็ˆฑ", which cuts duplicate entities; canonicalization logic and JSON tolerance together mean the pipeline assumes the model is helpfully sloppy, never precise.)

Takeaways

  • An LLM's job is plausible output; your job is a truth source. Where a registry exists (and it exists more often than you'd think), validation is a lookup, not a model.
  • Segment with Intl.Segmenter before validating anything emoji-adjacent. [...str] silently destroys exactly the sequences most likely to be hallucinated.
  • Reject the shape mismatch: \p{Extended_Pictographic} matches but the dictionary misses. That's your hallucination signature.
  • Policy beats function: batch minimums, DB-level dedupe, fail-closed provider handling, cache past the model, human gate on permanence.

The site is emoji-combos.net; the AI-assisted search is worth a try, and the curated field guide is the human-picked counterpart. Questions, corrections, and stories of your own model's inventions: comments are open.

FAQ

How much latency does validation add?
Microseconds per combo. The expensive step remains the LLM call itself; the dictionary is a static Set, the segmenter is a platform API, and the regexes are linear scans over a few dozen graphemes.

What if Unicode publishes new emoji?
Re-run the import against the new emoji-test file and redeploy. The validation contract is unchanged; the dictionary is data, not code. Ship the data with the app (or load it at boot) so the runtime never depends on fetching it.

Does this work for non-emoji LLM output?
The mechanism generalizes to any domain with a canonical registry: check output against the registry at the unit users perceive, and reject on mismatch. The emoji-specific parts are the segmenter settings and the tone-stripping rule; everything else is domain-neutral policy.

Why reject the whole batch below four valid combos?
Because survivors of a bad batch are a biased sample, and "ship whatever passed" hides upstream regressions. A hard floor turns the rejection rate into an observable signal instead of silent quality decay.

Why not use a second model to check the first?
A second model has the same failure mode as the first: it predicts plausibility, it doesn't consult truth. Chains of models checking each other add latency, cost, and new hallucination surface, while a Set.has() is exact. Use models for proposal; use registries for verdicts.

Top comments (0)