You have two product feeds that look close enough to compare prices. Then someone notices you matched a 256 GB phone against the 128 GB version, or a monitor bundle against the bare monitor. The pricing logic was fine. The bad decision happened earlier, when the catalog said two listings were the same product.
Product matching sounds simple until you deal with real ecommerce data. Titles are inconsistent, sellers stuff keywords into descriptions, variants hide in different fields, and regional listings use different units. If you treat matching as string similarity, you will get convincing false positives.
Matching is not search
Search can return similar things. Product matching needs to decide whether two listings represent the same sellable item.
These two are probably the same product:
Apple AirPods Pro 2nd Generation with MagSafe Charging Case
Apple AirPods Pro (2nd Gen), MagSafe Case, USB-C
These two are not the same product, even though the titles overlap heavily:
HP Pavilion 15 Laptop, Ryzen 5, 8GB RAM, 512GB SSD
HP Pavilion 15 Laptop, Intel i5, 8GB RAM, 512GB SSD
The difference is small in text and large in business impact. If you price the Ryzen SKU against the Intel SKU, your downstream reports will look precise while being wrong.
A practical matcher usually has four stages:
- Normalize raw fields
- Extract product attributes
- Generate likely candidates
- Score candidates with explicit mismatch rules
The important part is that candidate generation and final matching are different jobs. Candidate generation can be fuzzy. Final matching should be strict.
Normalize before comparing anything
Raw titles carry useful signals, but they also carry noise. Normalize units, casing, punctuation, common abbreviations, and known brand aliases before you compute similarity.
Here is a small example in Python. It is not a complete matcher, but it shows the shape of the problem:
import re
from dataclasses import dataclass
@dataclass
class Product:
title: str
brand: str | None = None
model: str | None = None
storage_gb: int | None = None
color: str | None = None
cpu: str | None = None
BRAND_ALIASES = {
"hewlett packard": "hp",
"apple inc": "apple",
}
COLORS = {"black", "white", "silver", "blue", "red", "space gray"}
CPU_PATTERNS = [
(r"\bryzen\s*5\b", "ryzen 5"),
(r"\bryzen\s*7\b", "ryzen 7"),
(r"\bi5\b|\bintel\s+core\s+i5\b", "intel i5"),
(r"\bi7\b|\bintel\s+core\s+i7\b", "intel i7"),
]
def normalize_text(value: str) -> str:
value = value.lower()
value = value.replace("-", " ")
value = re.sub(r"[^a-z0-9\s]", " ", value)
value = re.sub(r"\s+", " ", value).strip()
return value
def extract_product(title: str) -> Product:
text = normalize_text(title)
brand = None
for candidate in ["hewlett packard", "hp", "apple inc", "apple", "samsung"]:
if candidate in text:
brand = BRAND_ALIASES.get(candidate, candidate)
break
storage_gb = None
storage_match = re.search(r"(\d+)\s*(gb|tb)\b", text)
if storage_match:
amount = int(storage_match.group(1))
storage_gb = amount * 1024 if storage_match.group(2) == "tb" else amount
cpu = None
for pattern, normalized in CPU_PATTERNS:
if re.search(pattern, text):
cpu = normalized
break
color = next((c for c in COLORS if c in text), None)
model_match = re.search(r"\b([a-z]+\s?\d{2,4}[a-z]*)\b", text)
model = model_match.group(1).replace(" ", "") if model_match else None
return Product(title=title, brand=brand, model=model, storage_gb=storage_gb, color=color, cpu=cpu)
p1 = extract_product("HP Pavilion 15 Laptop, Ryzen 5, 8GB RAM, 512GB SSD")
p2 = extract_product("HP Pavilion 15 Laptop, Intel Core i5, 8GB RAM, 512GB SSD")
print(p1)
print(p2)
This code intentionally extracts attributes instead of producing one big similarity score. That lets you treat some differences as fatal. CPU, storage size, pack size, region, and included accessories often matter more than title similarity.
Use blocking so you do not compare everything
If catalog A has 100,000 products and catalog B has 100,000 products, a full pairwise comparison means 10 billion comparisons. You need blocking, which means generating a smaller candidate set using cheap keys.
Common blocking keys include:
brand + model
brand + normalized title tokens
GTIN / UPC / EAN if available
brand + category + storage size
image hash bucket for visually similar listings
Blocking can miss matches if you make it too strict. For example, one seller may omit the model number from the title but include it in a specs table. A good pipeline usually creates multiple candidate sources, then merges them.
This is also where data extraction quality matters. If your scraper misses the specs table or drops bundled accessories from the description, the matcher never gets a chance to make the right decision. Wire is relevant in this part of the workflow when competitor catalog extraction needs to preserve attributes, variants, and match confidence instead of returning only raw listing rows.
Score matches with hard vetoes
After candidate generation, compute a score. But do not rely only on the score.
A pair can have a 0.94 title similarity and still be wrong. Add hard vetoes for attributes that define the SKU:
def is_match(a: Product, b: Product) -> tuple[bool, str]:
if a.brand and b.brand and a.brand != b.brand:
return False, "brand mismatch"
if a.model and b.model and a.model != b.model:
return False, "model mismatch"
if a.storage_gb and b.storage_gb and a.storage_gb != b.storage_gb:
return False, "storage mismatch"
if a.cpu and b.cpu and a.cpu != b.cpu:
return False, "cpu mismatch"
return True, "candidate match"
This looks basic, but it prevents a common failure mode: a matcher that says yes because most tokens overlap. In ecommerce, the rare token often matters most.
You can still use embeddings or a trained classifier for the final score. Just make the model explain itself in terms of extracted attributes. If it cannot tell you why two products match, you will struggle to debug pricing errors later.
Treat bundles and regional differences as product data
Bundles are easy to miss because the title often stays almost identical:
Dell 27 Monitor S2721QS 4K UHD
Dell 27 Monitor S2721QS 4K UHD with HDMI Cable and Cleaning Kit
If the competitor includes accessories, matching their price directly may make your bare listing look overpriced. Extract included items into a separate field and compare them explicitly.
Regional differences cause a different kind of bug. A product may be the same even when units differ:
10 inch tablet
25.4 cm tablet
Normalize measurements before matching. But be careful with compliance labels, chargers, warranties, and radio bands. Those may represent real product differences, not formatting differences.
Counterfeits and unauthorized imports add another layer. The title and image can match perfectly while the seller, warranty, or fulfillment channel changes the risk. In that case, the output should not be just match=true. Use something like:
{
"match": true,
"confidence": 0.982,
"risk_flags": ["unauthorized_seller", "warranty_missing"]
}
Build for review, not blind automation
A useful matcher should produce three buckets:
match: safe to use for pricing or catalog merge
no_match: safe to ignore
review: too close to discard, too risky to trust
The review bucket is not a failure. It is how you avoid pretending uncertain data is certain. Track review reasons like variant_conflict, missing_model, bundle_unclear, or regional_difference. Those labels help you improve extraction and rules over time.
A practical next step: take 200 known product pairs from your own catalog, label them as match or no match, then run a simple attribute-based matcher before trying embeddings or ML. The mistakes in that first pass will tell you which attributes actually define identity in your category.
Top comments (0)