My RAG pipeline had a question it kept getting wrong. "How do I fix the replication lag alert on the orders DB?" The answer was in my runbooks. Retrieval found the right chunk at rank 3. Then the reranker moved it to rank 14, my top-5 cutoff threw it away, and the LLM confidently answered from a chunk about a different database.
The reranker wasn't dumb. It was blind. My bge-reranker truncates at 512 tokens, and the fix commands lived at the bottom of a chunk it never finished reading. When I counted, 31% of my chunks were longer than what the reranker could see.
This post is about that one mechanism: why a cross-encoder reranker truncates at 512 tokens, which part of your text it throws away, and what to do about it.
TL;DR
- Cross-encoder rerankers like
BAAI/bge-reranker-baseread the query and passage as one sequence capped at 512 tokens, including 4 special tokens. - When the pair is too long, the Hugging Face tokenizer uses
longest_firsttruncation, which trims the end of the passage. No error, no warning. - Your chunker, your embedding model and your reranker usually measure length with three different rulers. A chunk that fits one can overflow another.
- Detect it by tokenizing
(query, passage)with the reranker's own tokenizer and logging how often the length exceedsmax_length. - Fix it by chunking to the reranker's budget, or by scoring overlapping windows of the passage and taking the max.
How did my reranker bury the right chunk?
Here's the setup, because the details are where this bug hides:
- Corpus: about 2,400 chunks from internal runbooks and incident notes.
-
Chunker: LangChain's
RecursiveCharacterTextSplitter,chunk_size=2000. Characters, not tokens. -
Retrieval:
text-embedding-3-small, top 20 by cosine similarity. -
Reranking:
CrossEncoder("BAAI/bge-reranker-base")from sentence-transformers, keep top 5.
I had 60 hand-labeled questions. Retrieval put the correct chunk somewhere in the top 20 for 54 of them. After reranking, the correct chunk made the top 5 for only 41.
So the reranker was losing answers that retrieval had already found. That's backwards. The whole point of a reranker is to be the smarter, slower second pass.
I went through the 13 losses by hand. In 9 of them, the sentence that actually answered the question sat in the last third of the chunk. Runbooks are written that way: title, symptoms, context, and then, at the very bottom, the ## Resolution section with the commands you need.
Why does a cross-encoder reranker truncate at 512 tokens?
A cross-encoder reranker truncates at 512 tokens because it is a BERT-style encoder with a fixed number of position embeddings, and it reads the query and passage together as a single input. bge-reranker-base is built on XLM-RoBERTa, which was trained with 512 positions. Anything past that has no position to sit in, so the tokenizer cuts it before the model runs.
This is different from an embedding model (a bi-encoder). A bi-encoder embeds the query and the passage separately. A cross-encoder concatenates them:
<s> query tokens </s></s> passage tokens </s>
That's 4 special tokens for XLM-RoBERTa pairs. So your real passage budget is:
512 - 4 - len(query_tokens)
A 30-token question leaves 478 tokens for the passage. That joint attention over query and passage is exactly why cross-encoders rank better than cosine similarity. It's also why they have this hard ceiling.
Which part of the passage gets cut?
The end of the passage gets cut. sentence-transformers' CrossEncoder calls the tokenizer with truncation=True, which in Hugging Face means the longest_first strategy. It removes tokens one at a time from whichever sequence is currently longer. Your query is 30 tokens and your passage is 700, so the passage loses its tail every single time.
You can watch it happen:
from transformers import AutoTokenizer
tok = AutoTokenizer.from_pretrained("BAAI/bge-reranker-base")
enc = tok(query, passage, truncation=True, max_length=512)
seen = tok.decode(enc["input_ids"], skip_special_tokens=True)
print(seen[-300:]) # the last thing the reranker actually read
When I ran this on my replication-lag chunk, the last thing the reranker read was the middle of a paragraph explaining what replication lag is. The pg_stat_replication query and the fix were gone. The model scored "a chunk that explains the concept" against "a short chunk about a different DB that mentions the alert name in full," and the short chunk won. That's a reasonable judgment on the text it was given.
The silent part is what makes this nasty. With truncation=True, the tokenizer does not warn you. Your scores look like normal floats. Nothing in your logs says "I only read 60% of this."
Why didn't retrieval catch the problem?
Retrieval didn't catch it because every stage measured the chunk with a different ruler:
| Stage | Unit | Limit |
|---|---|---|
| Chunker | characters | 2,000 |
text-embedding-3-small |
OpenAI tokens | 8,191 |
bge-reranker-base |
XLM-RoBERTa tokens, query included | 512 |
The embedding model saw the whole chunk, including the Resolution section, and ranked it well. The reranker saw a truncated version of the same chunk and ranked it badly. The stage with the smallest window had the final say.
The character limit made it worse. 2,000 characters of English prose fits under 512 tokens comfortably. 2,000 characters of runbook does not. Shell commands, hostnames, UUIDs, stack traces and YAML break into many more subword tokens than ordinary words do. My chunks that overflowed were almost all the operational ones, which are exactly the ones people ask about.
How do you check if your reranker is truncating?
Tokenize every (query, chunk) pair with the reranker's tokenizer, without truncation, and count how many exceed the limit. It takes a few minutes on a few thousand chunks.
from transformers import AutoTokenizer
tok = AutoTokenizer.from_pretrained("BAAI/bge-reranker-base")
MAX_LEN = 512
TYPICAL_QUERY = "how do I fix the replication lag alert on the orders db"
def overflow(passage: str) -> int:
ids = tok(TYPICAL_QUERY, passage, truncation=False)["input_ids"]
return len(ids) - MAX_LEN
over = [c for c in chunks if overflow(c) > 0]
print(f"{len(over)}/{len(chunks)} chunks truncated "
f"({len(over) / len(chunks):.0%})")
Mine printed 31%. Two more checks worth adding:
- Log it in production. Record the untruncated pair length for every reranked candidate. A truncation rate is a metric you can alert on.
-
Check the
max_lengthyou actually pass. Wrappers and frameworks set their own defaults. If you swap in a reranker trained on longer inputs but your wrapper still passes 512, you're paying for a model you're not using.
How do you fix reranker truncation?
There are three fixes, and I ended up using the first two together.
1. Chunk to the reranker's budget, measured with the reranker's tokenizer. Reserve room for the query and the special tokens, then cap passages at what's left.
QUERY_RESERVE = 64 # longest query you expect, in reranker tokens
PASSAGE_BUDGET = 512 - 4 - QUERY_RESERVE # 444
def token_len(text: str) -> int:
return len(tok(text, add_special_tokens=False)["input_ids"])
splitter = RecursiveCharacterTextSplitter(
chunk_size=PASSAGE_BUDGET,
chunk_overlap=60,
length_function=token_len,
)
The one-line change that matters is length_function=token_len. Now the chunker and the reranker use the same ruler.
2. Score overlapping windows and take the max. Some documents shouldn't be split, or you can't re-index right now. Slide a window over the passage, score each window, and keep the highest. In IR research this is known as MaxP scoring.
def windowed_score(model, query, passage, window=400, stride=200):
ids = tok(passage, add_special_tokens=False)["input_ids"]
if len(ids) <= window:
return float(model.predict([(query, passage)])[0])
starts = list(range(0, len(ids) - window + 1, stride))
if starts[-1] + window < len(ids):
starts.append(len(ids) - window) # always cover the tail
windows = [tok.decode(ids[s:s + window]) for s in starts]
return float(max(model.predict([(query, w) for w in windows])))
Watch the starts.append line. Without it, a passage whose length isn't a neat multiple of the stride loses its last few dozen tokens. That would be the same bug you're fixing, just smaller.
3. Use a reranker trained on longer inputs. This works, but read the cost. Self-attention scales with the square of sequence length, and you run the reranker once per candidate. Going from 512 to 2,048 tokens on 20 candidates per query is a real latency change, so measure it on your own hardware before shipping.
After re-chunking with the reranker's tokenizer and using windowed scoring for the few long documents I kept whole, the correct chunk made the top 5 for 51 of my 60 questions, up from 41. Same embedding model, same reranker, same LLM. The only difference was that the reranker could now read the whole chunk.
What should you take away from this?
Every model in a RAG pipeline has a context window, and the smallest one controls quality. People check the LLM's window and maybe the embedding model's. Almost nobody checks the reranker's, because it returns a clean float either way.
If your documents put the important part at the end, like runbooks, API docs with examples at the bottom, or support tickets with the resolution last, you're in the worst case for longest_first truncation.
So does bge-reranker really truncate at 512 tokens?
Yes. bge-reranker-base and similar BERT- and XLM-RoBERTa-based cross-encoders read the query and passage as one sequence capped at 512 tokens, including special tokens. sentence-transformers truncates with longest_first, which silently drops the end of the passage. If your chunks are sized in characters, or in another model's tokens, a meaningful share of them may be partly invisible to your reranker, and any answer in that tail can't help the score. Measure overflow with the reranker's own tokenizer, size chunks to its budget, and use overlapping-window max scoring for passages you can't split.
Written by the developer behind Preterview, an interview prep platform.
Top comments (2)
The subword fragmentation on operational text makes this even worse than the character count suggests. SentencePiece on XLM-RoBERTa splits snake_case identifiers, path slashes, and UUIDs much more aggressively than cl100k_base, so a 50-line bash snippet can easily hit 400 tokens on its own. Passing length_function=token_len with the reranker's tokenizer during ingestion is the only reliable cure. Relying on character counts when your corpus contains shell commands is essentially asking for silent tail truncation.
tr.ee/dev-to