If you've ever built a social app, you've hit this problem: you have posts, you have users, and you need to decide what shows up first. Sorting by created_at DESC works for exactly one afternoon. After that, your most recent post buries something way more relevant from three hours ago, and your users start wondering why the app feels dead.
This post walks through building a real, working feed ranking algorithm from scratch — not a black-box ML model, just the actual logic (with code) behind how platforms like early Facebook, early Twitter, and most consumer feeds decide what goes where. We'll keep it simplified, but every piece here is a real technique used in production systems, just without the years of A/B-tested tuning layered on top.
The Core Idea: Ranking Is Just Scoring + Sorting
Strip away the mystique and a feed ranking algorithm is two steps:
- Score every candidate post for a given user
- Sort by that score, descending
That's it. The entire game is in how you calculate the score. So let's build a scoring function piece by piece.
Step 1: Start With a Naive Baseline
Before adding intelligence, let's establish what we're improving on:
def naive_score(post):
return post["created_at"].timestamp()
feed = sorted(posts, key=naive_score, reverse=True)
This is recency-only sorting. It's simple, but it has an obvious flaw: a mediocre post from 2 minutes ago beats a genuinely great post from 2 hours ago, every time. We need to factor in quality and relevance, not just freshness.
Step 2: Define Your Ranking Signals
Every real feed ranking system is built on three categories of signals:
- Engagement signals — likes, comments, shares, saves (how good is this post, generally?)
- Relevance signals — how related is this post to this specific user?
- Recency signals — how fresh is it?
Let's build a signal for each.
Engagement Score
def engagement_score(post):
likes = post.get("likes", 0)
comments = post.get("comments", 0)
shares = post.get("shares", 0)
# Comments and shares are stronger signals than likes —
# they take more effort, so weight them higher
return (likes * 1.0) + (comments * 3.0) + (shares * 5.0)
Weighting matters here. A like takes one tap. A comment takes intent. A share means someone found it worth broadcasting to their own network. Most real systems weight these very differently — this is a simplified version of that logic.
Relevance Score
This is where personalization comes in. A basic version compares the post author to the user's past interaction history:
def relevance_score(post, user_history):
author_id = post["author_id"]
# How many times has this user engaged with this author before?
past_engagements = user_history.get(author_id, 0)
# Diminishing returns — 10 past likes shouldn't count 10x more
# than 3 past likes, or the feed becomes an echo chamber
import math
return math.log1p(past_engagements) * 2.0
Using log1p (log(1 + x)) instead of raw counts prevents a single over-followed account from permanently dominating the user's feed — a real problem if you don't cap this signal.
Recency Score with Decay
Instead of pure timestamp sorting, we apply time decay — a post's relevance score shrinks the older it gets, but doesn't disappear immediately:
import time
def recency_score(post, half_life_hours=6):
age_hours = (time.time() - post["created_at"].timestamp()) / 3600
decay = 0.5 ** (age_hours / half_life_hours)
return decay
A half_life_hours=6 means a post's recency score is cut in half every 6 hours. This is the same math used to model radioactive decay — and it's a genuinely good fit here, because it lets old-but-great content still surface without letting stale content dominate.
Step 3: Combine the Signals
Now we combine everything into a single weighted score:
def rank_score(post, user_history, weights=None):
weights = weights or {
"engagement": 0.4,
"relevance": 0.4,
"recency": 0.2,
}
e_score = engagement_score(post)
r_score = relevance_score(post, user_history)
t_score = recency_score(post)
# Normalize engagement so it doesn't dwarf the other signals
# (a viral post with 50,000 likes shouldn't break the math)
import math
e_score_normalized = math.log1p(e_score)
final_score = (
weights["engagement"] * e_score_normalized +
weights["relevance"] * r_score +
weights["recency"] * t_score
)
return final_score
Notice the weights dict is exposed as a parameter, not hardcoded. This matters a lot in practice — these weights are exactly what platforms tune constantly through A/B testing. Your first guess at 0.4/0.4/0.2 will be wrong. That's fine. The architecture just needs to make it easy to change.
Step 4: Build the Actual Feed
def build_feed(posts, user_history):
scored = [
(post, rank_score(post, user_history))
for post in posts
]
scored.sort(key=lambda x: x[1], reverse=True)
return [post for post, score in scored]
feed = build_feed(candidate_posts, current_user_history)
At this point, you have a working, personalized, decay-aware feed ranking system. It's simplified, but structurally, it's not far off from what real early-stage social platforms ran in production.
Common Pitfalls (Learned the Hard Way)
1. Forgetting to normalize wildly different scales.
Engagement counts can range from 0 to 100,000+. Relevance and recency scores typically live between 0 and 1. If you don't normalize (notice the log1p calls above), your engagement signal will silently dominate everything else and your "personalized" feed becomes "just show whatever's viral."
2. No diversity control.
A pure score-and-sort approach will happily show a user 8 posts in a row from the same account if that account scores highest. Real systems add a diversity pass after scoring — for example, capping consecutive posts from the same author:
def diversify(ranked_posts, max_consecutive=2):
result = []
last_author = None
consecutive_count = 0
for post in ranked_posts:
if post["author_id"] == last_author:
consecutive_count += 1
else:
consecutive_count = 1
last_author = post["author_id"]
if consecutive_count <= max_consecutive:
result.append(post)
return result
3. Ignoring negative signals.
"Hide post," "unfollow," and fast scroll-past behavior are just as valuable as likes — arguably more valuable, since they're explicit. A more mature version of relevance_score should subtract for these, not just add for positive engagement.
4. Recomputing the entire feed on every request.
This naive implementation re-scores every candidate post per request, which doesn't scale. Production systems typically pre-compute candidate sets (via a "candidate generation" stage using approximate nearest-neighbor search or simple following-graph queries) and only run the expensive ranking pass on a few hundred candidates, not your entire posts table.
Where This Goes Next
This is deliberately the "classical" approach — weighted scoring functions with hand-tuned coefficients. It's exactly how most feed ranking worked before deep learning became cheap enough to run at inference time for every request.
The natural next step, once you have real usage data, is replacing hand-picked weights with a learned model — typically a gradient-boosted tree (like XGBoost) or a small neural network trained on historical engagement data to predict P(user will engage with this post), then ranking by that predicted probability instead of a manual formula.
But here's the thing worth saying plainly: don't start there. The hand-tuned version above will outperform a half-trained ML model for a long time, because ML models need volume and clean labels to beat a well-reasoned heuristic. Ship the simple version, watch how users actually behave, and let the data tell you when it's time to graduate to a learned ranker — not before.
Full Code Recap
import math
import time
def engagement_score(post):
likes = post.get("likes", 0)
comments = post.get("comments", 0)
shares = post.get("shares", 0)
return (likes * 1.0) + (comments * 3.0) + (shares * 5.0)
def relevance_score(post, user_history):
past_engagements = user_history.get(post["author_id"], 0)
return math.log1p(past_engagements) * 2.0
def recency_score(post, half_life_hours=6):
age_hours = (time.time() - post["created_at"].timestamp()) / 3600
return 0.5 ** (age_hours / half_life_hours)
def rank_score(post, user_history, weights=None):
weights = weights or {"engagement": 0.4, "relevance": 0.4, "recency": 0.2}
e = math.log1p(engagement_score(post))
r = relevance_score(post, user_history)
t = recency_score(post)
return weights["engagement"] * e + weights["relevance"] * r + weights["recency"] * t
def build_feed(posts, user_history):
scored = [(p, rank_score(p, user_history)) for p in posts]
scored.sort(key=lambda x: x[1], reverse=True)
return [p for p, s in scored]
Copy this, plug in your own post data, and you've got a real, working starting point — not a toy example, an actual foundation you can iterate on as you learn what your users respond to.
Top comments (0)