DEV Community

Anakin
Anakin

Posted on

Product Matching Is Mostly About Avoiding Confident Wrong Answers

You have two product feeds that look close enough to compare prices. Then someone notices you matched a 256 GB phone against the 128 GB version, or a monitor bundle against the bare monitor. The pricing logic was fine. The bad decision happened earlier, when the catalog said two listings were the same product.

Product matching sounds simple until you deal with real ecommerce data. Titles are inconsistent, sellers stuff keywords into descriptions, variants hide in different fields, and regional listings use different units. If you treat matching as string similarity, you will get convincing false positives.

Matching is not search

Search can return similar things. Product matching needs to decide whether two listings represent the same sellable item.

These two are probably the same product:

Apple AirPods Pro 2nd Generation with MagSafe Charging Case
Apple AirPods Pro (2nd Gen), MagSafe Case, USB-C
Enter fullscreen mode Exit fullscreen mode

These two are not the same product, even though the titles overlap heavily:

HP Pavilion 15 Laptop, Ryzen 5, 8GB RAM, 512GB SSD
HP Pavilion 15 Laptop, Intel i5, 8GB RAM, 512GB SSD
Enter fullscreen mode Exit fullscreen mode

The difference is small in text and large in business impact. If you price the Ryzen SKU against the Intel SKU, your downstream reports will look precise while being wrong.

A practical matcher usually has four stages:

  1. Normalize raw fields
  2. Extract product attributes
  3. Generate likely candidates
  4. Score candidates with explicit mismatch rules

The important part is that candidate generation and final matching are different jobs. Candidate generation can be fuzzy. Final matching should be strict.

Normalize before comparing anything

Raw titles carry useful signals, but they also carry noise. Normalize units, casing, punctuation, common abbreviations, and known brand aliases before you compute similarity.

Here is a small example in Python. It is not a complete matcher, but it shows the shape of the problem:

import re
from dataclasses import dataclass

@dataclass
class Product:
    title: str
    brand: str | None = None
    model: str | None = None
    storage_gb: int | None = None
    color: str | None = None
    cpu: str | None = None

BRAND_ALIASES = {
    "hewlett packard": "hp",
    "apple inc": "apple",
}

COLORS = {"black", "white", "silver", "blue", "red", "space gray"}

CPU_PATTERNS = [
    (r"\bryzen\s*5\b", "ryzen 5"),
    (r"\bryzen\s*7\b", "ryzen 7"),
    (r"\bi5\b|\bintel\s+core\s+i5\b", "intel i5"),
    (r"\bi7\b|\bintel\s+core\s+i7\b", "intel i7"),
]

def normalize_text(value: str) -> str:
    value = value.lower()
    value = value.replace("-", " ")
    value = re.sub(r"[^a-z0-9\s]", " ", value)
    value = re.sub(r"\s+", " ", value).strip()
    return value

def extract_product(title: str) -> Product:
    text = normalize_text(title)

    brand = None
    for candidate in ["hewlett packard", "hp", "apple inc", "apple", "samsung"]:
        if candidate in text:
            brand = BRAND_ALIASES.get(candidate, candidate)
            break

    storage_gb = None
    storage_match = re.search(r"(\d+)\s*(gb|tb)\b", text)
    if storage_match:
        amount = int(storage_match.group(1))
        storage_gb = amount * 1024 if storage_match.group(2) == "tb" else amount

    cpu = None
    for pattern, normalized in CPU_PATTERNS:
        if re.search(pattern, text):
            cpu = normalized
            break

    color = next((c for c in COLORS if c in text), None)

    model_match = re.search(r"\b([a-z]+\s?\d{2,4}[a-z]*)\b", text)
    model = model_match.group(1).replace(" ", "") if model_match else None

    return Product(title=title, brand=brand, model=model, storage_gb=storage_gb, color=color, cpu=cpu)

p1 = extract_product("HP Pavilion 15 Laptop, Ryzen 5, 8GB RAM, 512GB SSD")
p2 = extract_product("HP Pavilion 15 Laptop, Intel Core i5, 8GB RAM, 512GB SSD")

print(p1)
print(p2)
Enter fullscreen mode Exit fullscreen mode

This code intentionally extracts attributes instead of producing one big similarity score. That lets you treat some differences as fatal. CPU, storage size, pack size, region, and included accessories often matter more than title similarity.

Use blocking so you do not compare everything

If catalog A has 100,000 products and catalog B has 100,000 products, a full pairwise comparison means 10 billion comparisons. You need blocking, which means generating a smaller candidate set using cheap keys.

Common blocking keys include:

brand + model
brand + normalized title tokens
GTIN / UPC / EAN if available
brand + category + storage size
image hash bucket for visually similar listings
Enter fullscreen mode Exit fullscreen mode

Blocking can miss matches if you make it too strict. For example, one seller may omit the model number from the title but include it in a specs table. A good pipeline usually creates multiple candidate sources, then merges them.

This is also where data extraction quality matters. If your scraper misses the specs table or drops bundled accessories from the description, the matcher never gets a chance to make the right decision. Wire is relevant in this part of the workflow when competitor catalog extraction needs to preserve attributes, variants, and match confidence instead of returning only raw listing rows.

Score matches with hard vetoes

After candidate generation, compute a score. But do not rely only on the score.

A pair can have a 0.94 title similarity and still be wrong. Add hard vetoes for attributes that define the SKU:

def is_match(a: Product, b: Product) -> tuple[bool, str]:
    if a.brand and b.brand and a.brand != b.brand:
        return False, "brand mismatch"

    if a.model and b.model and a.model != b.model:
        return False, "model mismatch"

    if a.storage_gb and b.storage_gb and a.storage_gb != b.storage_gb:
        return False, "storage mismatch"

    if a.cpu and b.cpu and a.cpu != b.cpu:
        return False, "cpu mismatch"

    return True, "candidate match"
Enter fullscreen mode Exit fullscreen mode

This looks basic, but it prevents a common failure mode: a matcher that says yes because most tokens overlap. In ecommerce, the rare token often matters most.

You can still use embeddings or a trained classifier for the final score. Just make the model explain itself in terms of extracted attributes. If it cannot tell you why two products match, you will struggle to debug pricing errors later.

Treat bundles and regional differences as product data

Bundles are easy to miss because the title often stays almost identical:

Dell 27 Monitor S2721QS 4K UHD
Dell 27 Monitor S2721QS 4K UHD with HDMI Cable and Cleaning Kit
Enter fullscreen mode Exit fullscreen mode

If the competitor includes accessories, matching their price directly may make your bare listing look overpriced. Extract included items into a separate field and compare them explicitly.

Regional differences cause a different kind of bug. A product may be the same even when units differ:

10 inch tablet
25.4 cm tablet
Enter fullscreen mode Exit fullscreen mode

Normalize measurements before matching. But be careful with compliance labels, chargers, warranties, and radio bands. Those may represent real product differences, not formatting differences.

Counterfeits and unauthorized imports add another layer. The title and image can match perfectly while the seller, warranty, or fulfillment channel changes the risk. In that case, the output should not be just match=true. Use something like:

{
  "match": true,
  "confidence": 0.982,
  "risk_flags": ["unauthorized_seller", "warranty_missing"]
}
Enter fullscreen mode Exit fullscreen mode

Build for review, not blind automation

A useful matcher should produce three buckets:

match: safe to use for pricing or catalog merge
no_match: safe to ignore
review: too close to discard, too risky to trust
Enter fullscreen mode Exit fullscreen mode

The review bucket is not a failure. It is how you avoid pretending uncertain data is certain. Track review reasons like variant_conflict, missing_model, bundle_unclear, or regional_difference. Those labels help you improve extraction and rules over time.

A practical next step: take 200 known product pairs from your own catalog, label them as match or no match, then run a simple attribute-based matcher before trying embeddings or ML. The mistakes in that first pass will tell you which attributes actually define identity in your category.

Top comments (0)