Nobody types a famous quote exactly. They type it the way they remember it:
insanity is doing the same thing over and over and expecting different results
The usual wording has an "again" in it. Some people add "the definition of". Some put the attribution on the end: "… - Albert Einstein". A search box that needs every word, or the exact phrase, misses most of these.
I run GraciousQuotes, where each quote carries a verdict (verified, disputed, misattributed and so on) with its source. The Who said this quote? checker takes whatever someone pastes and answers with the closest quote on the site and its verdict. For the line above, it answers:
| Match | Credited to | Verdict | Quote |
|---|---|---|---|
| close (0.846) | Albert Einstein | disputed | Insanity is doing the same thing over and over again and expecting different results. |
Here's how it gets there, in two stages: a cheap database search to find candidates, then a careful PHP scorer to rank them.
Why plain search wasn't enough
The site's existing search was word-AND: return quotes containing every word typed. For checking quotes that fails in both directions:
- It misses. One extra word ("again", "the definition of") and the real quote is gone.
- It matches the wrong thing. Common words appear everywhere, so a quote that happens to contain all the same words, in a different order and meaning something else, counts as a hit.
What we actually want to know is: is this the same sentence, give or take a few words? That's about word order, not just shared words.
Step 1: normalise both sides
Before comparing anything, the query and every quote go through the same normalisation: decode HTML entities, strip accents, lowercase, delete apostrophes ("don't" and "dont" are the same word), and turn everything else that isn't a letter or digit into a space.
function gq_qc_norm( $s ) {
$s = html_entity_decode( wp_strip_all_tags( (string) $s ), ENT_QUOTES, 'UTF-8' );
$s = remove_accents( $s );
$s = strtolower( str_replace( array( "'", "\u{2019}", "\u{2018}" ), '', $s ) );
$s = preg_replace( '/[^a-z0-9]+/', ' ', $s );
return trim( preg_replace( '/\s+/', ' ', $s ) );
}
The query also loses a trailing credit ("… - Albert Einstein", "… —Mark Twain") and its wrapping quote marks, because people paste them along with the quote.
Step 2: candidates from MySQL FULLTEXT
Scoring every quote in PHP would be slow, so MySQL narrows it down first. Quotes live in a separate FULLTEXT table rebuilt nightly, and the checker asks it in natural-language mode, which ranks by relevance instead of requiring every word:
$rows = $wpdb->get_results( $wpdb->prepare(
"SELECT quote_id, text, MATCH(text) AGAINST (%s IN NATURAL LANGUAGE MODE) sc
FROM $t WHERE MATCH(text) AGAINST (%s IN NATURAL LANGUAGE MODE)
ORDER BY sc DESC LIMIT 120", $qn, $qn ) );
120 candidates is generous. FULLTEXT is good at "these share unusual words" and bad at "these are the same sentence", so the real decision happens in the next step.
Step 3: score word order and coverage
For each candidate, three measures between 0 and 1:
- Run: the longest stretch of consecutive words the query and the quote share, divided by the query's length. This is what captures "same sentence": word order matters.
- Query coverage: how much of what they typed is in the quote.
- Quote coverage: how much of the quote they typed.
The longest common run is the classic dynamic-programming longest-common-substring, over words instead of characters, keeping only the previous row:
function gq_qc_lcs_run( $a, $b ) {
$best = 0;
$prev = array_fill( 0, count( $b ) + 1, 0 );
foreach ( $a as $i => $wa ) {
$cur = array_fill( 0, count( $b ) + 1, 0 );
foreach ( $b as $j => $wb ) {
if ( $wa === $wb ) {
$cur[ $j + 1 ] = $prev[ $j ] + 1;
if ( $cur[ $j + 1 ] > $best ) $best = $cur[ $j + 1 ];
}
}
$prev = $cur;
}
return $best;
}
Then a weighted score, with the run counting most:
$score = 0.5 * $run + 0.3 * $cov_q + 0.2 * min( 1, $cov_c * 1.5 );
$kind = $score >= 0.7 && $run >= 0.6 ? 'close' : ( $score >= 0.55 ? 'possible' : 'none' );
Worked through for the insanity example (13 words typed; the quote has 14):
- the longest shared run is "insanity is doing the same thing over and over": 9 of 13 words → run 0.692
- all 13 typed words are in the quote → query coverage 1.0
- 13 of the quote's 14 words were typed → quote coverage 0.929, × 1.5 capped at 1 → 1.0
- score = 0.5 × 0.692 + 0.3 × 1.0 + 0.2 × 1.0 = 0.846, and the run is above 0.6, so it's close
The × 1.5 on quote coverage means typing most of a long quote counts fully. People usually remember the famous half.
Before any of that, one shortcut: if the typed words appear in order inside the quote (or a quote of at least 4 words appears inside the query), it's an exact match with score 1.0. "Float like a butterfly sting like a bee" is exact against Muhammad Ali's full line, which continues "Rumble, young man, rumble."
Step 4: present answers, not search results
A checker should give one answer, not ten results, so the ranking ends with a few rules:
- Ties go to the checked verdict. If two quotes score the same, a verified or misattributed record outranks an unreviewed one.
- One card per wording. Quotes whose first 80 normalised characters match are the same quote, shown once.
- Look-alikes are noise next to a real match. After an exact match, anything under 0.85 is dropped; after a close match, anything under 0.7.
- Variants fold in. Other wordings with the same author and verdict are listed under the top answer instead of repeating it.
Learning from misses
When the checker finds nothing, the query is counted by week (normalised text, first 300 characters, outcome) with no IP address, cookie or account, and rows older than 12 months are deleted. Those misses are the to-do list: they show which quotes people ask about that the site doesn't have yet.
Try it
The checker is free at graciousquotes.com/who-said-this-quote. Paste a quote the way you remember it. The misquotes are the interesting ones.
Top comments (0)