DEV Community

Cover image for Reversing Letter Order for QA: Building Test Fixtures From a Word Scrambler
Tea-sip for Lizely

Posted on

Reversing Letter Order for QA: Building Test Fixtures From a Word Scrambler

When you test language features, search ranking, or content-pipeline jobs, the input data is often more interesting than the code you run it on. A small scrambler that jumbles the inside letters of a token — while keeping the first and last characters fixed — is one of the cheapest ways to build a fixture corpus that proves a tokenizer, a similarity function, or a regex is doing the right thing. This piece is a walkthrough of the underlying rule, the envelope back on edge cases, and how to wire the tool into a test harness without spending marketing budget.

What "scramble the middle" actually means

The operation most casual users describe as "rule of typoglycemia" is a well-defined transformation. Given a token of length n, you leave token[0] and token[n-1] in place and randomly permute the substring between them. Lengths below four usually pass through unchanged, because the inner substring is empty or a single letter. Lengths of four keep two letters and swap the middle two. Lengths of five and above shuffle the interior freely.

The reference table below captures the envelope of cases worth pinning down in tests. I keep it next to the fixture generator so reviewers can see, at a glance, which behaviors are deliberate and which are leftovers from earlier drafts.

Input Inner length Possible outputs
aaaaa (n=5) 3 aaaax, aaxaa, axaaa (3 permutations)
abcab (n=5) 3 abacb, abcab (stays), abcba
abcab (n=5) 3 aacbb stays out — first and last move
aaaaaa (n=6) 4 24 permutations of inner
abcdefgh (n=8) 6 720 permutations

Two practical consequences fall out of that table. First, a four-letter token can give at most two outcomes, so any test that expects many distinct results from short inputs is testing the wrong thing. Second, a token with all-identical interior letters cannot distinguish a correct implementation from one that accidentally copies the original. For five-letter tokens, aaaaa only exposes three distinct outputs.

A second envelope worth naming is the "first and last fixed" envelope. The constraint that the boundary letters stay put is what lets a reader recognize the original word. Drop the constraint and the word becomes unreadable; tighten it to "no movement" and the corpus collapses to the identity mapping. The intermediate zone is the only useful one.

A third envelope is the whitespace envelope. Most scramblers operate per whitespace-separated token, so "the quick brown fox" becomes four independent draws rather than one permutation of nineteen letters. Pipeline code that joins tokens and then re-splits will quietly break that assumption.

Building a regression suite around the rule

The reason I reach for a scrambler in QA work is that it gives me tokens that are simultaneously realistic and adversarial. A token like "tlhe" or "qciuk" is a perfectly normal lowercase English shape that any regex tuned on a clean corpus will misclassify.

Here is a checklist for setting up the regression case you actually need to crib against:

  1. Pick a seed for the permutation generator. Without a seed, two consecutive runs of the same input produce two different files and the diff is meaningless.
  2. Pick a minimum length below which the tokenizer passes tokens through. Common values: 3, 4, or 5.
  3. Decide whether punctuation is part of the token or a separator. If part of the token, the rule "keep first and last" can trap a stray period in the boundary slot.
  4. Decide whether digits are tokens. A DOI handler that runs against a fixture of phone numbers will return scrambled phone numbers, which is rarely what you want.
  5. Generate at least 200 tokens per length bucket. Below 200, the standard deviation of unique values is bigger than the bucket's count, and you cannot tell whether a missing value is missing because the scrambler never produces it or because your sample was unlucky.
  6. Save the seed alongside the fixture. The only thing worse than a flaky test is one that looks stable but cannot be regenerated.

When I write tests against this corpus, I almost never assert on exact scrambled outputs. I assert on the invariant the rule is supposed to enforce. The invariants are small enough to fit on a checklist.

  1. scrambled[0] === original[0] for every token with length ≥ 2.
  2. scrambled[len-1] === original[len-1] for every token with length ≥ 2.
  3. The multiset of inner letters is preserved.
  4. Length is preserved.
  5. For length ≤ 3, the token equals the original.

Any implementation that breaks one of those five is broken. Any implementation that satisfies all five is correct under the rule, even if the exact permutation differs from yours. That separation between rule and specific output is the whole point of property-based testing.

When the rule fails: cases worth covering

There are several families of input where the "keep first and last" rule produces an unreadable token even though the rule is followed. They are worth covering because they catch the mistake of assuming the rule alone is sufficient.

Hyphenated compounds: "well-known" splits into "well-known" as one token in many tokenizers but two tokens in others. If the rule keeps the hyphen as a boundary letter, the inner scramble runs on "ell-known" or "ell" alone, neither of which is what a human reader needs.

Tokens with repeated first and last letters: "cocoa" and "papa" have identical boundary letters. A test that asserts scrambled[0] === original[0] passes trivially, but a reader cannot recover the original without the inner shape. This is the case where the rule is mechanically followed and semantically broken.

Mixed-case tokens: "iPhone" keeps "i" and "e" fixed. The interior "Phon" reshuffles into "Phon", "Pohn", "hPon", "honP" and so on. Capital letters in the interior behave like any other character — the rule does not preserve casing patterns. If your downstream code is case-sensitive, this is where flakes appear.

Tokens containing only two distinct characters: "aaaaaa" or "ababab" have so few unique inner permutations that a faulty shuffler can pass tests by accidentally returning the original on every draw.

Non-ASCII letters: the rule treats Unicode code points, not grapheme clusters. A token like "👨‍👩‍👧" is one grapheme cluster but multiple code points, and a per-character scrambler will break the cluster. For the deeper walkthrough on how to handle these cases when arranging scrambled output, the practical guide on how to arrange scrambled words covers the constraints in more detail than fits here.

Wiring the tool into a test harness

A practical harness treats the scrambler as a fixture source, not as the system under test. Concretely, I expose it through a small builder with three knobs: minimum length, a separator regex, and a seed. The builder emits a list of strings, which the test framework then consumes.

import random
import re

def scramble_word(word: str, rng: random.Random) -> str:
    if len(word) < 4:
        return word
    middle = list(word[1:-1])
    rng.shuffle(middle)
    return word[0] + "".join(middle) + word[-1]

def build_corpus(text: str, seed: int, min_len: int = 4) -> list[str]:
    rng = random.Random(seed)
    return [scramble_word(tok, rng) for tok in re.split(r"\s+", text) if len(tok) >= min_len]
Enter fullscreen mode Exit fullscreen mode

Two notes from running this in CI. The seed must come from a versioned source, not from time.time(), otherwise the corpus is unreproducible across machines. And the separator regex must match the tokenizer you actually use downstream, otherwise the boundary behavior is silently wrong.

Property-based testing on top of the corpus

Once the fixture pipeline is wired, property-based testing is the natural next layer. The Hypothesis documentation for property-based testing is the standard entry point if you work in Python; the fast-check documentation is the equivalent for TypeScript. Both libraries ship strategies for strings, regexes, and arbitrary text, and both let you assert invariants without pinning specific shuffles.

The pattern is the same in both libraries. Define a strategy that produces tokens of varying shape. Run the operation under test. Assert that the invariants hold across hundreds of generated examples. The corpus from the previous section is the deterministic counterpart — a fixed seed, a fixed text, and a snapshot-able file that lives in version control.

The reason to combine property-based tests with a seeded corpus is that the two catch different bugs. The seeded corpus catches regressions in the rule itself. The property-based test catches regressions in the operation that consumes the corpus.

Frequently asked questions

How do I pick a minimum length for the scrambler?

Four is the most common choice because it leaves exactly two inner letters and one nontrivial permutation. Three is the lower bound at which the rule still does something, but the only nontrivial case is a single swap of two inner letters, which is so close to the identity that tests rarely distinguish them. Five or above gives the shuffler room to permute meaningfully but starts to produce output that is hard for human reviewers to verify by eye.

What do I do about tokens that contain punctuation?

Treat punctuation as a separator, not as part of the token. A rule that keeps the first and last characters of "well-known" fixed will trap the hyphen in the boundary slot and scramble the wrong substring. Splitting on a separator regex like [^\w]+ sidesteps the trick entirely.

Can I use the same fixture for multiple systems under test?

Yes, as long as the seed is preserved and the systems are checking the same invariant. If two systems are checking different invariants — for example, one checks the boundary rule and the other checks a downstream similarity score — they should use separate corpora, because the corpus must be adversarial to the rule it is testing against.

Why not just use a dictionary attack instead?

A dictionary attack exercises known tokens. The point of scrambled fixtures is to exercise unknown tokens that look like real tokens. A similarity function trained on clean text will score typewell and typedll very differently, and that asymmetry is what your tests should pin down.


This article was drafted with AI assistance and reviewed for technical accuracy before publishing.

Top comments (0)