DEV Community

Cover image for Product matching: let code read the offers, let a head decide, and enforce one counterpart per product
Maxim Kuznetsov
Maxim Kuznetsov

Posted on AI-assisted

Product matching: let code read the offers, let a head decide, and enforce one counterpart per product

Two shops sell the same TV. One calls it "Sony KDL-40V512 LCD TV - black", the other "SONY Lcd Tv kdl40v512 (black)" at a slightly different price, next to its neighbour model kdl40v513. Is it the same product?

A language model asked pair by pair does surprisingly well at this. It also happily says "yes" to two different offers for the same product, because each question is asked alone. This post builds a matcher where code reads the offers, a small head decides, a promise says which answers can go out without a person, and a rule across all the answers makes "one counterpart per offer" hold. I wrote solvi, the library used; everything runs on pip install solvi==1.0.0, no model, no network. Part 2 of 5 (overview, part 1).

Why this way

Three observations from solvi's public task stand (Abt-Buy, 1,916 product pairs; measured with solvi 0.8.0, LLM gpt-oss-120b) shaped this:

  1. Comparison is code's job. Normalising model codes, comparing sizes and prices: deterministic, testable, cheap.
  2. A rule across items belongs after the items' decisions. "Each offer has at most one counterpart" enforced on the answers raised F1 for every solver tried, and helped the weak ones most: 0.872 → 0.909 for an LLM asked each pair, 0.860 → 0.870 for the same LLM through solvi, 0.931 → 0.933 for a fitted head.
  3. A promise needs a signal that separates right from wrong. A head's probability over good facts does.

On the stand, the full recipe scored F1 0.933 against 0.872 for the LLM asked per pair, with 0 offers matched to two counterparts instead of 19, and under a 1% risk promise it answered 96.9% of the pairs alone with 0.59% of all pairs answered alone and wrong. Honest caveat up front: that head is supervised (fitted on 5,743 labelled pairs) and the LLM baseline is zero-shot. If you have no labels, this is not your recipe.

Charts for Abt-Buy, 1,916 product pairs. F1 as given and with one counterpart per offer: LLM asked each pair 0.872 to 0.909, the same LLM through solvi 0.860 to 0.870, a fitted head over comparison facts 0.931 to 0.933. Offers matched to two counterparts: 19 for the plain LLM, 0 for the solvi solution. Under a 1% risk promise, 96.9% of pairs answered alone and 0.59% answered alone and wrong.

Facts by code

Here the offers are synthetic (six brands, made-up model codes; shop B rewrites codes its own way, drops them 15% of the time, and prices drift ±15%). Each candidate pair is a state with four given values: a_name, a_price, b_name, b_price. The catalog turns them into facts:

cat = Catalog()


def codes(text):
    return {re.sub(r"[^a-z0-9]", "", m.lower()) for m in re.findall(r"[A-Za-z]{2,3}[- ]?\d{2}[A-Za-z]\d{3}", text)}


@cat.fn
def code_equal(a_name, b_name) -> bool:          # the model codes, normalised, are the same
    return bool(codes(a_name) & codes(b_name))


@cat.fn
def code_missing(b_name) -> bool:                # shop B left the code out
    return not codes(b_name)


@cat.fn
def same_brand(a_name, b_name) -> bool:
    return a_name.split()[0].lower() == b_name.split()[0].lower()


@cat.fn
def same_colour(a_name, b_name) -> bool:
    return a_name.rsplit(" ", 1)[-1].lower() == b_name.rsplit("(", 1)[-1].strip(")").lower()


@cat.fn
def price_gap(a_price, b_price) -> float:        # relative price difference
    return abs(a_price - b_price) / max(a_price, b_price)
Enter fullscreen mode Exit fullscreen mode

Every candidate set holds the true counterpart (90% of the time) and two near misses: same brand and kind, a neighbouring model code, another colour or price. That is where pair-by-pair matching goes wrong.

A head, then a promise

No rule answers match, so System.fit learns one: a closed-form ridge head over the facts, fitted in well under a second. select=False keeps every fact. Then System.guarantee calibrates a threshold on pairs the head did not see:

system = System(cat, [Question("match", "Are these two offers the same product?", Answer.yes_no())])
head = system.fit("match", [(x, y) for x, y, _ in train], select=False)
rep = system.guarantee("match", [(x, y) for x, y, _ in valid], max_risk=0.01)
Enter fullscreen mode Exit fullscreen mode

max_risk=0.01 means: the share of all pairs that are answered alone and wrong is at most 1%, for pairs like the calibration ones. Everything below the threshold is still answered, but with status != "ok", meaning "a person should look".

One counterpart per offer

The head answers each pair on its own. solvi.core.sets.decide_set then takes every pair's probabilities and keeps the most probable combination in which each offer matches at most one candidate — exactly, not greedily:

items = [Item((r["a"], r["b"]), {"yes": r["p"], "no": 1 - r["p"]}, keys={"a": r["a"]}) for r in rows]
out = decide_set(items, [AtMostOne("a")])
set_pred = [out[(r["a"], r["b"])].answer == "yes" for r in rows]
Enter fullscreen mode Exit fullscreen mode

The output

head over 7 facts, fitted on 866 pairs
guarantee on 878 pairs: threshold 0.6053, answered alone 96.5%
  P(answered alone and wrong) ≤ 0.01 — a share of all inputs — for inputs like the calibration examples
test, 874 pairs: F1 head 0.964 (offers with two counterparts: 13) -> with one counterpart per offer 0.987 (0); 13 answers changed, exact: True
answered alone 96.3%, wrong among them 9 (1.03% of all pairs; promise 1%)
Enter fullscreen mode Exit fullscreen mode

The set rule changed 13 answers and removed all 13 offers that had two counterparts; F1 went from 0.964 to 0.987. That is the same direction as on Abt-Buy, larger because the toy is built around near misses.

Now look at the last line, because it is the kind of number a launch post would usually hide: on this test set 9 of 874 pairs, 1.03%, were answered alone and wrong, against a promise of 1%. That is not a bug in the promise. The method behind max_risk (conformal risk control) bounds the share on average over calibration sets; one calibration set and one test set of 874 pairs fluctuate around it. If you need "at most 1% with high probability", use max_error= with learn-then-test, which holds with probability ≥ 1 − delta and lets fewer pairs through. And if most of your pairs are easy non-matches, read the guide's warning: "1% of all pairs" can hold while the matches given alone are wrong far more often; groups="answer" puts the promise inside each answer.

What it does not do

  • It does not read the offers for you. On Abt-Buy the comparison code (pairfacts.py) is about 200 lines of ordinary Python; solvi did not write it.
  • It is supervised. The head needs labelled pairs; the stand used 5,743. A zero-shot LLM baseline is a different trade-off, not a weaker version of this one.
  • The promise covers pairs like the calibration ones. A new product category, a shop that writes codes a new way: calibrate again.
  • Refits move the threshold. solvi's best practices record a case where refitting the head on more labels picked a much weaker ridge penalty and broke the promise's threshold, though the labels were right; with a fixed penalty (lam=100, pairs=False) the guarantee kept answering about 90%. Recalibrate after every refit.
  • It does not get better on its own over time. On product matching, facts added from a person's answers to escalations gave no measurable gain in solvi's tests; a static head already answers most pairs. What can help, measured: letting the verified answers of a slow path recalibrate the guarantee (the share answered alone in the second half of a replayed stream rose from 90.7% to 96.7%, error among them 1.4% under a 2% promise).

Using an LLM instead of a head

If you have no labels yet, the same structure works with an LLM as the per-pair decider (solvi.models.llm(...) behind a decision part) and the set rule after it. Two measured tips from the stand: with a reasoning model, do not force a JSON schema on the reply (F1 0.837 forced vs 0.860 with the reply contract in the prompt, on Abt-Buy), and give it max_tokens room (at 400, 26 of 1,916 replies were cut off; solvi's default with reasoning is 2,048).

Where next

If you match entities today: which cross-item rules do you have (one counterpart, same price band, one address per customer), and where do they live now — in the model's prompt, in a cleanup script, or nowhere? And would you accept a promise "on average" for this, or does your use need "with high probability"?

Top comments (2)

Collapse
 
arhancanli profile image
Arhan Canli •

Printing the 9-of-874 against the 1% promise instead of hiding it is the useful part of this post. A reader can now see what an expected-risk bound looks like on one draw. With 9 wrong out of about 842 answered alone, an exact interval on that test rate runs from roughly 0.5% to 2%, so a 1.03% result sits comfortably inside the noise of a 1% target.

One question on the set rule: AtMostOne("a") enforces one counterpart per shop-A offer. Abt-Buy is one-to-one, so a shop-B offer should also match at most one A offer. Is the same constraint declared on the B side, and if so does decide_set handle both at once (that is an assignment problem, not a per-group argmax)? If only "a" is keyed, the 13 offers fixed here are the A-side duplicates, and the B-side ones could still be hiding in the F1.

I'd also add a line on the 10% of candidate sets with no true counterpart. At-most-one allows "no match", but a head that is confident about a near miss will still produce a false positive there, and the set rule can't fix it.

Collapse
 
makeev profile image
Mikhail Makeev •

On where the cross-item rules live: ours run after the per-item decisions, in the code that groups news into stories. For one pairing the rule lived nowhere until last week. An 8-K and the company's own press release about the same event showed up as two stories, because the grouping kept SEC filings out of press clusters by design. Now it's one root per event. If the release headline appears word for word in the 8-K's EX-99.1 exhibit from the same issuer within a few hours, the press release is the root and the filing joins it. On a week of production pairs that key matched 33 of 33. In Abt-Buy, how often does the normalised model code alone settle a pair before the head runs?