DEV Community

devansh jaiswal
devansh jaiswal

Posted on

Cosine Similarity Doesn't Know What Time It Is

Vector search answers "what is most similar to my query?" Many real systems also need "what is most true right now?" Those are different questions.

While building a co-pilot for pediatric therapists, my plain vector index kept ranking a four-month-old note ("tolerated musical games well") above a note from 90 minutes earlier about an acute auditory crisis. The old note was semantically closer to the query, so it won. In a clinical setting, that's the wrong answer.

Why it happens

Embeddings encode meaning, not validity. A note's timestamp isn't part of the vector, so a stale fact and a fresh one compete only on wording.

Fix 1: blend similarity with recency

Re-rank the retrieved hits with an exponential-decay recency term:

def rescore(hits, now, half_life_hours=48, alpha=0.6):
    ranked = []
    for h in hits:
        age_h = (now - h["ts"]) / 3600  # ts = epoch seconds
        recency = 0.5 ** (age_h / half_life_hours)
        score = alpha * h["sim"] + (1 - alpha) * recency
        ranked.append({**h, "score": score})
    return sorted(ranked, key=lambda x: x["score"], reverse=True)
Enter fullscreen mode Exit fullscreen mode

A note loses half its recency weight every half_life_hours. alpha controls how much similarity matters versus freshness.

Fix 2: not everything should decay

A crisis note goes stale in hours. A note like "weighted lap pad resolves agitation in about 4 minutes" is a durable protocol and shouldn't fade after a week. So give each note type its own half-life:

HALF_LIFE = {
    "acute_event": 6,        # hours
    "sleep_log": 72,
    "protocol": 24 * 90,
}

def rescore(hits, now, alpha=0.6, default_half_life=48):
    ranked = []
    for h in hits:
        half_life = HALF_LIFE.get(h["type"], default_half_life)
        age_h = (now - h["ts"]) / 3600
        recency = 0.5 ** (age_h / half_life)
        score = alpha * h["sim"] + (1 - alpha) * recency
        ranked.append({**h, "score": score})
    return sorted(ranked, key=lambda x: x["score"], reverse=True)
Enter fullscreen mode Exit fullscreen mode

Caveats

  • Tune, don't guess. Pick alpha and the half-lives by evaluating against real queries with known-good answers.
  • Keep scales comparable. sim and recency should both sit in roughly 0 to 1, or one term will silently dominate.
  • Re-ranking only reorders. If the fresh note isn't in your top-k candidates, decay can't rescue it. Retrieve a wider set first (say, top 50), then rescore.
  • Some facts need overriding, not decaying. If a new note contradicts an old one, consider explicit supersession rather than relying on the score gap.

(The examples here are illustrative, not clinical guidance, and contain no real patient data.)

Takeaways

  1. Similarity is not validity. Treat time as a first-class signal.
  2. Decay rates should follow the type of fact, not one global constant.
  3. Recency can be added as a re-ranking step over any existing vector index, or handled by an agent-memory layer if you'd rather not maintain it yourself.

Top comments (0)