DEV Community

Cover image for I built a 3KB alternative to replace zxcvbn (389KB) without detection loss
Fayaz F
Fayaz F

Posted on

I built a 3KB alternative to replace zxcvbn (389KB) without detection loss

zxcvbn is the most widely used password strength estimator with 1M npm downloads a week. It's also 389KB gzipped and hasn't shipped a commit since 2017. Most sign-up forms are hauling that around just to block password123.

Poor password UX is a real conversion problem. A strength meter that adds 389KB to your bundle delays page load — on mobile, measurably so. Users who hit a slow registration page don't wait. They leave. The irony is that most of that weight goes toward catching passwords nobody is actually using to register on your site.

So I built passcore - 3.0KB gzipped and 98.4% detection rate on real breach data - same as zxcvbn, benchmarked against a deduped list of passwords pulled live from RockYou, Adobe, HIBP, and other major leak lists.

zxcvbn takes ~9.7ms to load — it's parsing 389KB of dictionary into memory on every cold start. passcore loads in ~0.2ms. It evaluates a password in ~2,600 nanoseconds. For a registration form, it's effectively invisible — no jank, no layout shift, no contribution to your Core Web Vitals score. The strength meter shows up before the user finishes typing their first character.

How it works:

passcore runs five detection layers on every password:

  1. Dictionary - All entries sourced directly from breach data, not a generic word list
  2. Keyboard patterns - qwerty , asdf , 1234 , numpad walks
  3. Repeats - aaaa , ababab
  4. Sequences - abcdef , 123456
  5. L33t speak - decodes p@ssw0rd → password , m0nk3y → monkey , then dictionary lookup

The dictionary is small by design. Every entry was chosen because it appears in real breach data - not because it's a common English word. Password1! is caught not by a 40k word list but by stripping the suffix and checking if the core word is in the breach list. It is.

The scoring model:

passcore returns a score from 0 to 4 - same scale as zxcvbn.

The detection layers run first. A dictionary match, keyboard pattern, repeat, sequence, or l33t substitution scores 0 or 1 immediately - no further calculation. If none of those fire, scoring falls through to length and character variety: uppercase, lowercase, digits, symbols. A password that clears all five layers but is only 6 characters long still scores low.

There's also a length floor, aligned with NIST SP 800-63B: passwords 20+ characters score at least 3, passwords 30+ characters score 4, regardless of character variety. A passphrase like correct-horse-battery-staple is vastly harder to crack than P@ss1 - the scoring reflects that.

The research:

Getting to 98.4% detection required more than a dictionary lookup. A few problems that came up during development:

Word+affix patterns: Password1!, Admin123, Welcome1 - extremely common in breach data, none of them are in any dictionary as-is. The fix was a matchCommonRoot layer: strip leading and trailing non-alpha characters, check if what's left is a breach word. It is, every time, for this class of password.

L33t speak with separators: N0=Acc3ss decodes to no=access. A naive l33t decoder finds no dictionary match and passes it. The fix was to split the decoded string on non-alpha characters and check each segment independently. access is in the breach list. Caught.

Missing critical roots: Running against real breach lists exposed that admin, test, user, login, pass weren't in the dictionary - meaning Admin123, test1234, user2024 all slipped through. Added those five. Caught.

Switching looks like this:

// before
import zxcvbn from 'zxcvbn';
const { score } = zxcvbn(password);

// after
import { passcore } from 'passcorelib';
const { score } = passcore(password);
Enter fullscreen mode Exit fullscreen mode

One caveat: result.feedback.warning becomes result.warning, making it one level flatter.

zxcvbn zxcvbn-ts passcore
Bundle (gzipped) 389 KB 855 KB 3.0 KB
Speed 77,578 ns/op 839,991 ns/op 2,622 ns/op
Detection rate 98.4% 98.4% 98.4%
Maintained No Yes Yes

The tradeoff:

The tradeoff is dictionary size: 329 entries vs 40k+. But the passwords responsible for most credential stuffing aren't obscure literary references - they're Password1!, baseball123, keyboard walks, and l33t variants of the top breach list. passcore catches those.

So that's the bet passcore makes: that 329 targeted entries catch more of what actually matters than 40,000 words that cover everything, including passwords no one uses and/or no attacker is trying. The benchmark agrees — 98.4% detection rate across 370 real breach passwords, same as zxcvbn, at 130x less weight. For the 1% that need exhaustive coverage, use zxcvbn.

TL;DR — zxcvbn is 389KB and abandoned. passcore is 3KB, same detection rate, actively maintained. If bundle size matters to you, it's a near drop-in swap.

GitHub · npm

Top comments (5)

Collapse
 
itskondrat profile image
Mykola Kondratiuk •

389KB is wild for what's basically a password strength check. appreciate someone shipped the lighter version.

Collapse
 
thormeier profile image
Pascal Thormeier •

Great project, thank you for sharing!

A couple thoughts I had and a bit of feedback:

  • Removing 99.2% of the data used by the original lib to save space is a pretty clever tactic - but also a tad dangerous. It would be interesting to know where you made the cut, especially since the number 329 sounds arbitrary at first glance. Is this the top 94.8% of passwords and you chose that exact number to match the benchmarks you did on zxcvbn? What's your reasoning behind that exact number?
  • Perhaps giving the option to define additional dictionary entries would be a nice feature to have? On Github, you mentioned "What it won't catch is a user setting their password to an obscure literary reference, a foreign-language word, or a surname that happens to be in a dictionary but not in breach data." - perhaps having an API to add user-land dictionaries could help cover regional differences? I could imagine that some people would be highly interested in contributing region packages for passcore, too!
  • I read through the code a little and noticed that you're catching sequences from keyboard rows. Coming from a non-US country, I usually use a keyboard that has Y and Z swapped, which is actually rather common in central Europe. Are there any plans to cover these as well? QWERTY, to me, seems equally unsafe as QWERTZ, but the code is likely not catching that.

Keep up the great work! It's always refreshing to see someone questioning the status quo and coming up with more sensible solutions! :D

Collapse
 
fz_1357 profile image
Fayaz F • • Edited

Hi Pascal! Firstly thank you for taking the time to give such detailed feedback, I appreciate you.

On the 329 and whether it was chosen to match the benchmark:
No. The dictionary by itself directly matches 250 of the 370 benchmark passwords — not all of them. The remaining 114 that get caught are flagged by the pattern matchers (keyboard walks, repeats, sequences, l33t substitutions, word+affix detection). The 329 is simply how many breach password entries came out of curating the top passwords from HIBP and SecLists after deduping the compiled list. Structural patterns like pure number sequences and keyboard walks were left to the matchers rather than putting them in the dictionary. The 98.4% detection rate matching zxcvbn was observed after building both independently - not a number that was engineered or tuned to hit.
On user-land dictionaries and regional packages:
Yes, this idea is interesting and something I want to add. The API shape could have a second argument: passcore(password, { dictionary: [...] }) - merged at runtime, zero impact on bundle size for anyone who doesn't use it. Region packages as separate npm packages (passcore-de, passcore-fr etc.) is exactly the community model that would make this useful without bloating the core. Worth opening an issue for this one.
On QWERTZ:
The keyboard rows and patterns in the code are QWERTY-only, so the concern is valid. That said, the 60% adjacency threshold provides more incidental coverage than you'd expect - I tested it:
If 60% or more of consecutive character pairs are adjacent on the QWERTY keyboard, it's flagged as a keyboard pattern.
qwertz is caught (80% of consecutive pairs are adjacent even in the QWERTY map), qwertzuiop is caught (78%), yxcvbnm — the QWERTZ bottom row — is caught (83%). The real gap is short sequences in the t-z swap zone; something like tzui only hits 33% and slips through. Explicit QWERTZ support would be a small targeted fix. If you'd want to contribute it, very welcome.

Once again, thank you for your support and feedback, means a lot :)

Collapse
 
voltagegpu profile image
VoltageGPU •

Interesting approach to shrinking a widely used tool without sacrificing accuracy—good work! From an infrastructure standpoint, I'm curious how this would perform in a GPU-accelerated setup, especially if you're doing real-time validation at scale. I've seen VoltageGPU help with similar workloads by offloading repetitive checks.

Collapse
 
stacksmith profile image
stacksmith • • Edited

TLTR: zxcvbn-ts/core does the same but in a much more refined and way more secure way while keeping a low profile as the dictionaries are optional. passcorelib is a fancy password requirement tool but is not comparable to zxcvbn.

I was searching for a way to validate my registration password and found this post. So I dug into the rabbit hole of password validation and used my whole weekend for that.

This is such a weird post. You only compare against zxcvbn but have a benchmark against zxcvbn-ts as well.

For me, the whole benchmark seems rigged. You fetch like 200 entries from each password leak, which are mostly already in your dictionary, and then compare that with those other two. What's the difference between password 200 and password 201? Nothing, they are both used a million times. There are password lists out there that include occurrence counts, so they are sorted by how often each password was used. And in the original password dictionary it seems like this was such a list used to create the dictionary.

In your benchmark, you compare your deliberately tiny 329-entry breach dictionary against zxcvbn-ts with its common and English language dictionaries. That's a valid comparison of the default configurations, but it is not really a comparison of the algorithms using the same dictionaries.
The core package itself is only around 10 KB. The large dictionaries are separate packages and are optional. Of course, you normally want to use them because removing the dictionaries makes zxcvbn-ts much less effective. But comparing a 329-entry dictionary with the full zxcvbn-ts dictionaries and then comparing the bundle sizes makes it look like you are comparing equivalent functionality when you are not.

Additionally, there are a lot more algorithms in place to check a password.
passcore does not just look like zxcvbn-ts with fewer features. It implements some of the same types of password pattern detection, but the actual algorithms are much simpler and cover a much smaller subset of what zxcvbn-ts does.
For example, the keyboard layout support is extremely limited. There are only a few hard-coded patterns like QWERTY, ASDF and numeric sequences. zxcvbn checks an actual keyboard layout and can detect paths across the keyboard, including turns. zxcvbn-ts/core event supports different keyboard layouts.
The same applies to the other matchers. The dictionary matcher also supports things like reversed words, l33t substitutions, Diceware and optional Levenshtein matching.
Every matcher in passcore is basically a much simpler version of what zxcvbn does.

The benchmark seems to work largely because of how the scoring is designed. Once one of your detection layers finds something, the password is basically pushed into the same weak-score category. So if a password is detected as a dictionary password, keyboard pattern, repeat, sequence or l33t pattern, it can already end up with a score of 1. That makes the benchmark good at showing that your matchers can detect those particular passwords, but it doesn't tell you much about the quality of the actual strength estimation.
Additionally, something you don't mention is that for passwords scoring above 1, you fall back to the old-style password validation rules: uppercase, lowercase, number, symbol and a minimum length.
Those rules have been known for a long time to be a poor way of estimating password strength. Adding an uppercase letter, number and symbol doesn't necessarily make a password meaningfully harder to crack, especially when those characters are added in predictable places.
So you have essentially replaced the more complicated password-strength estimation with "does it contain one of these patterns, otherwise does it have enough character types and enough characters?"

The algorithms in zxcvbn are not just there to detect whether something is a bad pattern. They are used to estimate how many guesses an attacker would need. The different matchers produce different kinds of matches and those are used in the scoring.
This is also why getting 364 out of 370 passwords from your breach lists does not mean that passcore and zxcvbn provide the same password strength estimation. It only means that they both happen to classify most of those particular passwords as weak.

There is a difference between detecting a pattern and actually estimating the strength of the password.

For example, there is a difference between qwerty being caught by a dictionary matcher and being caught by the keyboard matcher. Those are different patterns and have different guess costs. The same applies when you combine multiple patterns, like a dictionary word followed by a year, a sequence or another pattern.

You published a security-related article that praises your work for having a smaller bundle size and faster execution while neglecting the whole security point of these libraries.
You don't even mention that your algorithms are a basic subset of what the original zxcvbn approach and zxcvbn-ts are doing, which makes the comparison even worse.

If the goal is to have a tiny and fast password checker that catches obviously bad passwords, that's fine. That's a reasonable trade-off.
But then you should compare it as a tiny and fast password checker, not make a benchmark showing that it catches the same passwords from a small breach sample and then imply that it provides comparable password-strength estimation.