DEV Community

Cover image for Your Vector Search Knows “Bank” Is Related to “Bank.” It Doesn’t Know Which Bank You Mean.
Joel Trout II
Joel Trout II

Posted on

Your Vector Search Knows “Bank” Is Related to “Bank.” It Doesn’t Know Which Bank You Mean.

Vector search is very good at similarity.

Similarity is not always the same thing as meaning.

That distinction starts becoming expensive when retrieval feeds an LLM.

Consider:

financial bank
river bank
Enter fullscreen mode Exit fullscreen mode

They share the same word.

A similarity system has every reason to place them near each other.

A useful reasoning system often needs to do the opposite.

It needs to separate the senses.

That problem is one of the reasons I built ARBITER.

ARBITER is a deterministic measurement engine. You give it:

context
+
a field of possibilities
Enter fullscreen mode Exit fullscreen mode

and it returns a coherence-ordered field.

It does not generate an answer.

It measures the possibilities you supplied.

Why this matters for RAG

A common RAG pipeline looks roughly like:

query
  ↓
vector retrieval
  ↓
top N chunks
  ↓
LLM
Enter fullscreen mode Exit fullscreen mode

The retrieval stage is intentionally broad.

That is useful, but it also means bad context can survive long enough to reach generation.

Once incorrect-but-related context enters the prompt, the generator has to reason around it.

A different pipeline is:

query
  ↓
vector retrieval
  ↓
candidate field
  ↓
ARBITER
  ↓
coherence-ordered field
  ↓
LLM
Enter fullscreen mode Exit fullscreen mode

ARBITER does not replace retrieval.

It gives you a deterministic measurement step between retrieval and generation.

A simple example

Here is a live ARBITER call:

curl -sS -X POST https://arbiter.grip.fyi/v1/compare \
  -H 'content-type: application/json' \
  --data '{
    "query":"python memory",
    "candidates":[
      "garbage collection",
      "malloc",
      "snake habitat"
    ],
    "top_k":3
  }'
Enter fullscreen mode Exit fullscreen mode

The resulting ordering:

0.483573  garbage collection
0.324794  malloc
0.167992  snake habitat
Enter fullscreen mode Exit fullscreen mode

Same interface:

state / intent / context
+
field of possibilities
→
ARBITER
→
ranked resonance field
Enter fullscreen mode Exit fullscreen mode

The field could contain:

documents
tools
routes
robot actions
suppliers
code paths
hypotheses
agents
products
Enter fullscreen mode Exit fullscreen mode

The primitive does not change. :chatgpt-content-reference{index="0"}

The harder test: word senses

An earlier ARBITER compression/disambiguation benchmark compared a 768-dimensional source representation compressed to 72 dimensions using PCA versus ARBITER.

The similarity-retention result was:

PCA       0.8693
ARBITER   0.9653
Enter fullscreen mode Exit fullscreen mode

But the more interesting result was sense separation.

For ambiguous words:

                PCA       ARBITER

bank            ~0.85      0.066
bat             ~0.85      0.073
Enter fullscreen mode Exit fullscreen mode

Lower here means better separation between competing senses.

So river bank and financial bank remained strongly entangled after PCA compression, while ARBITER separated them much more sharply.

The same benchmark reduced the representation from 768 dimensions to 72: a 10.7× dimensional reduction. :chatgpt-content-reference{index="1"}

That is the part I care about.

Not simply:

Can I preserve similarity?

But:

Can the representation preserve enough structure to distinguish what something means in context?

More examples

The same behavior shows up in ordinary ambiguous language.

For:

Best bass fishing spots in freshwater lakes
Enter fullscreen mode Exit fullscreen mode

ARBITER produced:

0.772  Largemouth bass in shallow weedy areas
0.542  Bass amplifiers and speaker impedance
0.293  Bass clef instruments in orchestra
0.272  Bass guitar string gauges
Enter fullscreen mode Exit fullscreen mode

For:

Crane safety regulations on construction sites
Enter fullscreen mode Exit fullscreen mode

it produced:

0.828  Tower cranes require certified operators
0.325  Sandhill cranes migrate through Nebraska
0.274  Origami cranes symbolize peace in Japan
0.173  Crane flies are harmless insects
Enter fullscreen mode Exit fullscreen mode

And:

Cell division rates in tumor growth analysis
Enter fullscreen mode Exit fullscreen mode

returned:

0.823  Mitotic cell division in tumor tissue
0.312  Prison cell division protocols
0.287  Cellular network division coverage
Enter fullscreen mode Exit fullscreen mode

These are not generated answers.

They are measurements over an explicit candidate field. :chatgpt-content-reference{index="2"}

Context can reorganize the same field

This is where things get more interesting.

Take Python.

Without extra context:

Programming   0.796
Snakes        0.284
Enter fullscreen mode Exit fullscreen mode

Now change the supplied perspective:

"As a herpetologist..."
Enter fullscreen mode Exit fullscreen mode

and the same meanings reorganize:

Snakes        0.700
Programming   0.422
Enter fullscreen mode Exit fullscreen mode

Apple behaves similarly:

baseline:
Tech company  0.861
Fruit         0.252
Enter fullscreen mode Exit fullscreen mode

With:

"As a chef..."
Enter fullscreen mode Exit fullscreen mode

the ordering flips:

Fruit         0.668
Tech company  0.397
Enter fullscreen mode Exit fullscreen mode

No retraining.

The supplied context changed, so the field changed. :chatgpt-content-reference{index="3"}

Why I think this matters

A lot of current AI infrastructure treats representation as a lookup problem:

Which stored object is closest?
Enter fullscreen mode Exit fullscreen mode

But many useful machine decisions are closer to:

Given this exact state,
which of these possibilities fits best?
Enter fullscreen mode Exit fullscreen mode

Those are not identical questions.

RAG is an obvious place to use that distinction because retrieval already gives you a bounded field.

But the same operation applies to agent routing, tool selection, robotics, screening, planning, and other systems where the candidates already exist.

The generator does not always need to make the decision.

Sometimes the candidates are already there.

What you need is a measurement.

Try it

The live endpoint is:

POST https://arbiter.grip.fyi/v1/compare
Enter fullscreen mode Exit fullscreen mode

Or install the lightweight CLI:

curl -fsSL https://arbiter.grip.fyi/install | sh
Enter fullscreen mode Exit fullscreen mode

Then:

arb "python memory" \
  "garbage collection" \
  "malloc" \
  "snake habitat"
Enter fullscreen mode Exit fullscreen mode

The CLI is just the interface to the hosted ARBITER service.

The current developer surface includes 10 successful calls per day free, after which the same endpoint moves to native x402 payment at $0.01 per call. :chatgpt-content-reference{index="4"}

Try a field where you already know what the answer should be.

Ambiguous words are a good place to start.

ARBITER: https://arbiter.grip.fyi

Description:
Similarity is not the same thing as meaning. A deterministic measurement step for RAG, reranking, and bounded decision fields.

Tags:
ai, rag, machinelearning, programming

Top comments (0)