<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Alamzeb Khan</title>
    <description>The latest articles on DEV Community by Alamzeb Khan (@alamzebkhan).</description>
    <link>https://dev.to/alamzebkhan</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4167391%2F53b2f86e-7608-4ab4-9fcc-ddc430a1dfb3.png</url>
      <title>DEV Community: Alamzeb Khan</title>
      <link>https://dev.to/alamzebkhan</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/alamzebkhan"/>
    <language>en</language>
    <item>
      <title>Building a Wordle Solver: Letter Frequency Analysis in JavaScript</title>
      <dc:creator>Alamzeb Khan</dc:creator>
      <pubDate>Tue, 06 Oct 2026 20:44:59 +0000</pubDate>
      <link>https://dev.to/alamzebkhan/building-a-wordle-solver-letter-frequency-analysis-in-javascript-22np</link>
      <guid>https://dev.to/alamzebkhan/building-a-wordle-solver-letter-frequency-analysis-in-javascript-22np</guid>
      <description>&lt;p&gt;Wordle looks simple. Five letters, six guesses. But building a &lt;em&gt;good&lt;/em&gt; solver — one that consistently cracks the puzzle in 3-4 guesses — is a genuinely interesting algorithms problem. Here's how I approached it when building the &lt;a href="https://wordylab.com/wordle/answer-today/" rel="noopener noreferrer"&gt;Wordle Solver&lt;/a&gt; on &lt;a href="https://wordylab.com" rel="noopener noreferrer"&gt;WordyLab&lt;/a&gt;, my free word game toolkit.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Problem
&lt;/h2&gt;

&lt;p&gt;Given a set of constraints (green = correct letter + position, yellow = correct letter wrong position, gray = letter not in word), narrow 12,000+ possible words down to the answer in as few guesses as possible.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 1: Letter Frequency Scoring
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;scoreWord&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;word&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;frequencyMap&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;uniqueLetters&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Set&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;word&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;score&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;letter&lt;/span&gt; &lt;span class="k"&gt;of&lt;/span&gt; &lt;span class="nx"&gt;uniqueLetters&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;score&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="nx"&gt;frequencyMap&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;letter&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;score&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Note the &lt;code&gt;Set&lt;/code&gt; — we count each letter once per word. A word like "speed" shouldn't get double credit for the double-e; what matters is &lt;em&gt;coverage&lt;/em&gt; of the alphabet.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 2: The Optimal Opening Guess
&lt;/h2&gt;

&lt;p&gt;Run frequency analysis on the full word list and the best openers surface: words with five unique, high-frequency letters. Classics like "adieu," "audio," and "raise" all score well because they test the most common vowels and consonants simultaneously.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 3: Constraint Propagation
&lt;/h2&gt;

&lt;p&gt;After each guess, filter ruthlessly — greens must match exact position, yellows must be present but elsewhere, grays must be absent. The gray-letter edge case is where most DIY solvers break: if you guess "speed" and the first E is green but the second E is gray, the E is &lt;em&gt;in&lt;/em&gt; the word, just once. Your filter has to handle letter counts, not just presence.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 4: Entropy (The Real Secret)
&lt;/h2&gt;

&lt;p&gt;Frequency scoring gets you to 4-5 guesses. To hit 3 consistently, you need &lt;strong&gt;information theory&lt;/strong&gt;. The optimal guess isn't the most likely answer — it's the guess that &lt;em&gt;eliminates the most candidates&lt;/em&gt; regardless of the outcome. For each candidate guess, simulate every possible feedback pattern and pick the one with minimum expected remaining pool.&lt;/p&gt;

&lt;h2&gt;
  
  
  Beyond Wordle
&lt;/h2&gt;

&lt;p&gt;The same toolkit powers anagram solvers, Scrabble word finders, and crossword solvers. All client-side, all instant. Try the working solver at &lt;a href="https://wordylab.com" rel="noopener noreferrer"&gt;wordylab.com&lt;/a&gt; — free, no signup.&lt;/p&gt;

</description>
      <category>javascript</category>
      <category>algorithms</category>
      <category>tutorial</category>
      <category>gamedev</category>
    </item>
    <item>
      <title>I Built a Free Word Counter That Runs 100% in Your Browser — Here's What I Learned About Client-Side Text Analysis</title>
      <dc:creator>Alamzeb Khan</dc:creator>
      <pubDate>Tue, 06 Oct 2026 20:43:54 +0000</pubDate>
      <link>https://dev.to/alamzebkhan/i-built-a-free-word-counter-that-runs-100-in-your-browser-heres-what-i-learned-about-399n</link>
      <guid>https://dev.to/alamzebkhan/i-built-a-free-word-counter-that-runs-100-in-your-browser-heres-what-i-learned-about-399n</guid>
      <description>&lt;p&gt;Counting words sounds trivial. It's not.&lt;/p&gt;

&lt;p&gt;When I set out to build &lt;a href="https://wordcountersuite.com" rel="noopener noreferrer"&gt;Word Counter Suite&lt;/a&gt; — a free toolkit with 15+ text analysis tools — I assumed the word counting part would take an afternoon. It took weeks. Here's why, and what I learned about doing real text analysis entirely client-side.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Naive Approach (and Why It Breaks)
&lt;/h2&gt;

&lt;p&gt;Most tutorials tell you this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;wordCount&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;trim&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;split&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sr"&gt;/&lt;/span&gt;&lt;span class="se"&gt;\s&lt;/span&gt;&lt;span class="sr"&gt;+/&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This breaks in at least five ways:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Hyphenated words&lt;/strong&gt; — is "state-of-the-art" one word or four? Microsoft Word says one. Google Docs says one. Your regex says four.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Em dashes and en dashes&lt;/strong&gt; — "word—another" with no spaces. One word or two?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Numbers&lt;/strong&gt; — "3.14" contains a period. Is it a sentence boundary?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;URLs and emails&lt;/strong&gt; — "&lt;a href="mailto:user@example.com"&gt;user@example.com&lt;/a&gt;" has no spaces but isn't a normal word.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Unicode&lt;/strong&gt; — &lt;code&gt;\s&lt;/code&gt; doesn't catch all Unicode whitespace. Neither does &lt;code&gt;\w&lt;/code&gt; handle all word characters.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  What Production Tools Actually Do
&lt;/h2&gt;

&lt;p&gt;Real word counters (Word, Google Docs, professional tools) follow published segmentation rules. The closest public standard is &lt;a href="https://unicode.org/reports/tr29/" rel="noopener noreferrer"&gt;Unicode Text Segmentation (UAX #29)&lt;/a&gt;, which defines word boundaries across languages.&lt;/p&gt;

&lt;p&gt;I ended up implementing a rule-based tokenizer that handles apostrophes within words, hyphens within compounds, decimal numbers, URLs and emails as single tokens, CJK characters, and emoji sequences.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;tokenize&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;normalized&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;normalize&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;NFC&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;wordPattern&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sr"&gt;/&lt;/span&gt;&lt;span class="se"&gt;[\p&lt;/span&gt;&lt;span class="sr"&gt;{L}&lt;/span&gt;&lt;span class="se"&gt;\p&lt;/span&gt;&lt;span class="sr"&gt;{N}&lt;/span&gt;&lt;span class="se"&gt;]&lt;/span&gt;&lt;span class="sr"&gt;+&lt;/span&gt;&lt;span class="se"&gt;(?:[&lt;/span&gt;&lt;span class="sr"&gt;''&lt;/span&gt;&lt;span class="se"&gt;\-&lt;/span&gt;&lt;span class="sr"&gt;‑–&lt;/span&gt;&lt;span class="se"&gt;][\p&lt;/span&gt;&lt;span class="sr"&gt;{L}&lt;/span&gt;&lt;span class="se"&gt;\p&lt;/span&gt;&lt;span class="sr"&gt;{N}&lt;/span&gt;&lt;span class="se"&gt;]&lt;/span&gt;&lt;span class="sr"&gt;+&lt;/span&gt;&lt;span class="se"&gt;)&lt;/span&gt;&lt;span class="sr"&gt;*/gu&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;normalized&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;match&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;wordPattern&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="p"&gt;[];&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;\p{L}&lt;/code&gt; and &lt;code&gt;\p{N}&lt;/code&gt; Unicode property escapes do the heavy lifting — they match letters and numbers in &lt;em&gt;any&lt;/em&gt; script, not just ASCII.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Client-Side Matters
&lt;/h2&gt;

&lt;p&gt;Every tool on Word Counter Suite runs entirely in the browser. No server round-trips, no data leaving the machine. Web Workers handle heavy analysis off the main thread, so pasting a 100,000-word manuscript doesn't freeze the UI. No backend means zero infrastructure cost and infinite scale.&lt;/p&gt;

&lt;h2&gt;
  
  
  Beyond Word Count
&lt;/h2&gt;

&lt;p&gt;Once you have a tokenizer, the rest follows: reading time (words ÷ WPM), keyword density (token frequency with stop-word filtering), speaking time, and readability scores like Flesch-Kincaid.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try It and Break It
&lt;/h2&gt;

&lt;p&gt;The toolkit is free at &lt;a href="https://wordcountersuite.com" rel="noopener noreferrer"&gt;wordcountersuite.com&lt;/a&gt; — no signup. If you find text that breaks the counter, I'd genuinely love to know. Edge cases in Unicode text segmentation are endless, and every weird input makes the tokenizer better.&lt;/p&gt;

</description>
      <category>javascript</category>
      <category>webdev</category>
      <category>tutorial</category>
      <category>opensource</category>
    </item>
  </channel>
</rss>
