DEV Community

Cover image for Flesch-Kincaid & Lexical Diversity: Automating Content Readability Analysis
Sameer Hassan
Sameer Hassan

Posted on

Flesch-Kincaid & Lexical Diversity: Automating Content Readability Analysis

Engineering teams writing technical documentation and landing pages often fall into one of two extremes:

  1. Academic Over-Complexity: Writing dense, multi-clause paragraphs packed with passive voice that exhaust the reader.
  2. Oversimplified Marketing Fluff: Using buzzwords without concrete technical substance, triggering zero interest from real developers.

In ⚡ PLYXO (CRO • SEO • AIO • AEO • GEO), our content audit pipeline programmatically measures Flesch-Kincaid Reading Ease and Type-Token Lexical Diversity (TTR) to ensure technical copy converts at maximum efficiency.


1. The Mathematics of Content Readability

The Flesch-Kincaid Reading Ease Formula:

$$\text{Score} = 206.835 - 1.015 \left( \frac{\text{Total Words}}{\text{Total Sentences}} \right) - 84.6 \left( \frac{\text{Total Syllables}}{\text{Total Words}} \right)$$

  • Score 90–100: 5th grade level (Comics, extremely simple).
  • Score 60–70: Plain English (Recommended for high-converting SaaS landing pages).
  • Score 30–50: College level (Academic papers).
  • Score < 30: Extremely difficult (Legal contracts).

If your landing page scores < 45, visitors experience cognitive fatigue, bounce rates spike, and conversion rates plummet.


2. Automated Readability Analyzer in TypeScript

Here is our production syllable counting and reading ease implementation:

export interface ReadabilityMetrics {
  readingEase: number;
  gradeLevel: number;
  wordCount: number;
  sentenceCount: number;
  averageWordsPerSentence: number;
  status: 'OPTIMAL' | 'TOO_COMPLEX' | 'TOO_SIMPLE';
}

function countSyllables(word: string): number {
  word = word.toLowerCase().trim();
  if (word.length <= 3) return 1;
  word = word.replace(/(?:[^laeiouy]|ed|es|e)$/, '');
  word = word.replace(/^y/, '');
  const matches = word.match(/[aeiouy]{1,2}/g);
  return matches ? matches.length : 1;
}

export function analyzeReadability(text: string): ReadabilityMetrics {
  const sentences = text.split(/[.!?]+/).filter(s => s.trim().length > 0);
  const words = text.split(/\s+/).filter(w => w.trim().length > 0);

  const wordCount = words.length;
  const sentenceCount = Math.max(1, sentences.length);
  const syllableCount = words.reduce((acc, word) => acc + countSyllables(word), 0);

  const wordsPerSentence = wordCount / sentenceCount;
  const syllablesPerWord = syllableCount / wordCount;

  // Flesch Reading Ease Formula
  const readingEase = Math.round(
    206.835 - 1.015 * wordsPerSentence - 84.6 * syllablesPerWord
  );

  // Flesch-Kincaid Grade Level
  const gradeLevel = Math.round(
    0.39 * wordsPerSentence + 11.8 * syllablesPerWord - 15.59
  );

  let status: ReadabilityMetrics['status'] = 'OPTIMAL';
  if (readingEase < 50) status = 'TOO_COMPLEX';
  if (readingEase > 80) status = 'TOO_SIMPLE';

  return {
    readingEase: Math.max(0, Math.min(100, readingEase)),
    gradeLevel: Math.max(1, gradeLevel),
    wordCount,
    sentenceCount,
    averageWordsPerSentence: Math.round(wordsPerSentence * 10) / 10,
    status,
  };
}
Enter fullscreen mode Exit fullscreen mode

3. Why LLMs Prioritize High-Readability Technical Content

When Answer Engines (Perplexity AI, ChatGPT Search) summarize web documentation, they favor content that explains complex concepts in clear, active sentences. Articles with optimal readability scores are selected for LLM citation summaries 3.4x more frequently than dense, convoluted text.

👉 Test your content readability with Plyxo on GitHub

Top comments (0)