Search Relevance & ML Ranking: Getting Search Right
You built semantic search. But results still feel wrong.
Users search "waterproof hiking boots" and see "dress shoes". They search "budget gaming laptop" and see $3,000 machines.
This is the gap between retrieval (find candidates) and ranking (order by relevance).
The Search Pipeline
Query: "waterproof hiking boots"
↓
RETRIEVAL: Find candidates
Vector search → 10,000 results
Keyword search → 50,000 results
Category filter → Boots category → 8,000 results
↓
RANKING: Order by relevance
BM25 score
+ semantic similarity
+ category match
+ price alignment
+ rating/popularity
+ freshness
↓
DISPLAY: Top 10 most relevant
Relevance Signals
Beyond semantic similarity, what makes a result relevant?
1. Category Matching
- Query: hiking boots
- Result: boots
- Boost: +0.5
2. Price Alignment
- Query budget: $150
- Result $120 (in budget): +0.3
- Result $500 (over budget): -0.5
3. Freshness
- Newer products ranked slightly higher
- Stale inventory penalized
4. Popularity
- High ratings boosted
- User reviews signal quality
5. Business Objectives
- Margins matter (higher margin products slightly boosted)
- New launches highlighted
6. User Context
- Browse history signals intent
- Returning customer preferences
Traditional Ranking: BM25 + Manual Boosting
GET /products/_search
{
"query": {
"bool": {
"must": [
{ "match": { "title": "hiking boots" } }
],
"should": [
{ "term": { "category": "boots", "boost": 2.0 } },
{ "range": { "price": { "lte": 150, "boost": 1.5 } } },
{ "range": { "rating": { "gte": 4.0, "boost": 1.2 } } }
]
}
}
}
Problem: Manual tuning. How much should category boost vs. price? Endless A/B tests.
Learning-to-Rank (LTR): ML Approach
Train an ML model on relevance signals.
Step 1: Collect Training Data
query: "waterproof hiking boots"
product_id: 123
category: "boots"
price: 129
rating: 4.5
bestseller: true
human_relevance_score: 0.92 // 1.0 = perfect, 0.0 = irrelevant
Collect thousands with human labels.
Step 2: Feature Engineering
def extract_features(query, product):
return {
'bm25_score': bm25(query, product),
'semantic_score': cos_sim(query_emb, product_emb),
'category_match': 1.0 if matches(query, product) else 0.0,
'price_aligned': 1.0 if in_budget(query, product) else 0.0,
'popularity': product.rating / 5.0,
'is_bestseller': 1.0 if product.bestseller else 0.0
}
Step 3: Train Ranking Model
from sklearn.ensemble import GradientBoostingRanker
X = [extract_features(q, p) for q, p in training_data]
y = [relevance_score for _, _, relevance_score in training_data]
ranker = GradientBoostingRanker(n_estimators=100)
ranker.fit(X, y)
Step 4: Deploy Two-Stage Ranking
@GetMapping("/search")
public List<Product> search(@RequestParam String query) {
// Stage 1: Retrieve 100 candidates (fast, rough)
List<Product> candidates = elasticsearch.search(query, 100);
// Stage 2: Re-rank with ML (slow, precise)
return candidates.stream()
.map(p -> new Ranked(p, mlRanker.predict(features(query, p))))
.sorted(Comparator.comparingDouble(Ranked::score).reversed())
.limit(10)
.map(Ranked::product)
.collect(Collectors.toList());
}
Measuring Relevance Quality
Don't guess. Measure.
1. NDCG (Normalized Discounted Cumulative Gain)
Rewards ranking where most relevant results are at top.
relevance = [0.8, 0.2, 0.7, 0.3, 1.0] # Top 5 results
ndcg = ndcg_score([1], [relevance])
# NDCG 0.85 = excellent (target: 0.8+)
2. Click-Through Rate (CTR)
Implicit signal. If users click, it's relevant.
3. A/B Testing
Old ranker (50% traffic) vs new ranker (50% traffic).
Measure: CTR, add-to-cart, revenue.
Deploy if new ranker wins.
Production Patterns
Two-Stage Ranking
- Stage 1: 10,000 candidates (BM25, vector, filters) — fast
- Stage 2: 100 candidates (ML model) — accurate but slower
Online Learning
Retrain nightly incorporating user feedback (clicks, purchases).
Monitoring
Alert if NDCG drops >5% (model regression detected).
Conclusion
Great search is intentional. It requires:
- Measuring baseline (NDCG)
- Collecting human labels
- Training ML models
- Deploying two-stage ranking
- Monitoring quality metrics
- Iterating based on user signals
From "close enough" to "users love it" is ML ranking.
Top comments (0)