Quick answer: Choosing an embedding model for legal document search is not just about picking the model with the highest benchmark score. Legal documents contain defined terms, clause numbers, exceptions, and near-duplicate language, so the model needs to capture both semantic meaning and important lexical details.
Qwen3-Embedding-8B and BAAI bge-m3 are strong starting points, while Snowflake Arctic Embed L v2.0 offers a more efficient option for larger document collections. These open-weight models can be run through Superlinked Inference Engine (SIE), either on SIE Cloud or in your own environment.
Hosted legal models from Voyage and Isaacus are alternatives if you prefer a managed service.
The best approach is to test multiple models on your own contracts and queries, compare retrieval quality, latency, and infrastructure cost, and then choose the model that performs best for your specific legal search workload.
Imagine a lawyer searching a supply agreement for "Confidential Information."
Instead of finding the definition clause—the clause that actually determines what "Confidential Information" means—the search returns dozens of paragraphs that simply discuss confidentiality. The definition clause ends up at position 31.
Why does this happen?
The embedding model may understand "Confidential Information" as a general topic related to secrecy. The lawyer, however, is looking for a specific defined term, where capitalization and the exact wording can have contractual consequences.
This is one of the biggest challenges with legal search.
Legal documents are filled with terms that look ordinary but have very specific meanings within a contract. Defined terms, clause numbers, party names, exceptions, and near-identical boilerplate can all affect whether the right passage appears at the top of the results.
That makes the embedding model an important part of the retrieval pipeline—not just another component to swap in and out.
In this article, we look at seven open-weight embedding models that can be used commercially and are relevant to legal document search. We also compare them with hosted legal embedding APIs, look at what is involved in moving away from OpenAI embeddings, and show how these models can be evaluated and deployed through Superlinked Inference Engine.
Licensing note: Several otherwise strong models were left out because their licenses restrict commercial use. These include jina-colbert-v2, Linq-Embed-Mistral, SFR-Embedding, and Reason-ModernColBERT, all of which carry a CC-BY-NC-4.0 license.
1. Qwen3-Embedding-8B
Model card: Qwen/Qwen3-Embedding-8B on Hugging Face
- License: Apache 2.0
- Context length: 32K tokens
- Deployment: Open weights; available in the SIE self-hosted catalog
If retrieval quality is your primary concern and you have the infrastructure to support a larger model, Qwen3-Embedding-8B is one of the first models worth testing.
It is an LLM-based dense embedding model with a 32K-token context window and support for more than 100 languages.
Qwen's model card reports a score of 70.58 on the MTEB multilingual leaderboard as of June 5, 2025, along with a retrieval score of 69.44 on MTEB English v2.
Those are the highest MTEB figures among the models discussed here. However, there is an important caveat: the models in this comparison report results across different benchmarks and evaluation dates, so these numbers should not be treated as a direct apples-to-apples ranking.
Why Qwen3-Embedding-8B is interesting for legal search
The 32K context window is a major advantage when working with long contractual clauses.
A limitation-of-liability clause, indemnification provision, or termination section can contain multiple sub-clauses, exceptions, and carve-outs. A long-context model can preserve more of that information when generating the embedding.
But there is a catch.
More context does not automatically mean better retrieval.
If you put an entire multi-page contract section into a single vector, the important sentence can get diluted by surrounding boilerplate. For legal search, clause-level or otherwise carefully designed chunks can still perform better than very large chunks.
Qwen3-Embedding is also instruction-aware. You can tell the model what kind of query it is processing—for example, that the user is looking for a particular contractual clause.
According to the model card, leaving out the instruction can reduce retrieval performance by roughly 1–5%, depending on the task.
The trade-off
The biggest drawback is the model's size.
With roughly 8 billion parameters, Qwen3-Embedding-8B requires considerably more GPU resources than the smaller models in this list. That matters not only for query-time latency but also when indexing or re-indexing a large contract archive.
Its reported leaderboard position is also based on a dated evaluation, so it should be treated as a strong candidate rather than a guaranteed winner.
If you want to test this model through SIE Cloud, the smaller 4B sibling is the Qwen model listed there, while the 8B version is available in the self-hosted catalog.
2. BAAI bge-m3
Model card: BAAI/bge-m3 on Hugging Face
- License: MIT
- Context length: 8,192 tokens
- Deployment: Open weights; available on SIE Cloud and self-hosted
If you want a model that combines several retrieval approaches in one package, bge-m3 is an excellent place to start.
Unlike a traditional dense embedding model, bge-m3 can produce three types of representations:
- Dense vectors
- Sparse lexical weights
- ColBERT-style multi-vectors
That combination is particularly useful for legal search.
Why sparse retrieval matters
Consider a query such as:
"Section 12.4 confidentiality obligations"
A dense model understands the overall meaning of the query, but legal search often depends on exact terms too.
A defined term, statute number, section reference, party name, or specific contractual phrase may need to match lexically rather than just semantically.
That's where the sparse representation becomes useful.
bge-m3 supports more than 100 languages and has an 8,192-token context window, making it suitable for longer clauses and multilingual agreements.
BAAI also recommends combining retrieval approaches with reranking, which is a common pattern in production search systems.
The advantage here is that bge-m3 can provide both dense and sparse signals from the same model.
The trade-off
bge-m3 dates back to 2024, and newer embedding models report stronger results on some general-purpose benchmarks.
But benchmark leadership isn't necessarily the deciding factor for legal search.
Its combination of dense + sparse + multi-vector retrieval makes bge-m3 a particularly strong baseline. If another model cannot outperform it on your own legal evaluation set, switching may not be worth the additional complexity.
3. Snowflake Arctic Embed L v2.0
Model card: Snowflake/snowflake-arctic-embed-l-v2.0 on Hugging Face
- License: Apache 2.0
- Context length: 8,192 tokens
- Parameters: 568M
- Deployment: Open weights; available on SIE Cloud and self-hosted
Not every legal search application needs the largest model available.
If you are working with hundreds of thousands—or millions—of documents, indexing speed and infrastructure cost become just as important as retrieval quality.
That's where Snowflake's Arctic Embed L v2.0 becomes interesting.
The model has 568 million parameters, produces 1,024-dimensional vectors, and supports an 8,192-token context.
Snowflake reports:
- BEIR: 55.6
- MIRACL: 55.8
- CLEF focused: 52.9
- CLEF full: 54.3
These figures come from Snowflake's model card and use different benchmarks from some of the other models discussed here, so they should be used as reference points rather than direct rankings.
Why efficiency matters in legal search
Consider a legal archive containing hundreds of thousands of contracts.
Now imagine changing your chunking strategy.
That seemingly small change could require you to re-embed the entire collection.
A smaller model can make that operation significantly faster and less expensive.
Arctic Embed L v2.0 also provides multilingual support, which is useful for organizations working with agreements across multiple jurisdictions.
It is a good candidate for teams that care about throughput, indexing cost, and query latency while still wanting a strong general-purpose embedding model.
4. intfloat Multilingual-E5-Large-Instruct
Model card: intfloat/multilingual-e5-large-instruct on Hugging Face
- License: MIT
- Context length: 512 tokens
- Parameters: 560M
- Deployment: Open weights; available in the SIE self-hosted catalog
E5 is one of the better-known embedding families and makes a useful baseline for comparison.
The multilingual-e5-large-instruct model has 560 million parameters, 24 layers, 1,024-dimensional embeddings, and support for roughly 100 languages.
It also uses an instruction prefix for queries.
The biggest limitation for legal documents, however, is the 512-token context window.
That can become a problem surprisingly quickly.
A detailed limitation-of-liability clause or indemnification provision can easily exceed 512 tokens. You then have to either split the clause into smaller pieces or accept some loss of context.
Neither option is ideal.
That doesn't make E5 a poor model. It simply means that chunking becomes especially important if you use it for contracts.
E5 is particularly useful as a reference model because it has been widely studied and has a long track record in retrieval evaluations.
5. nomic-embed-text-v2-moe
Model card: nomic-ai/nomic-embed-text-v2-moe on Hugging Face
- License: Apache 2.0
- Context length: 512 tokens
- Parameters: 475M total / 305M active
- Deployment: Open weights; SIE self-hosted catalog
nomic-embed-text-v2-moe uses a mixture-of-experts architecture with approximately 475 million total parameters and around 305 million active parameters per token.
It supports roughly 100 languages and is another interesting option for multilingual collections.
One thing that sets Nomic apart is its focus on transparency. The organization publishes information about its training data and code, which can be valuable for legal-tech companies that need to explain their AI stack to customers.
The model also supports Matryoshka embeddings, allowing the vector size to be reduced from 768 dimensions down to 256 dimensions.
That can be useful when storage becomes a concern.
For large legal archives, reducing vector dimensions can have a meaningful impact on index size.
The limitation
As with multilingual-e5-large-instruct, the context window is only 512 tokens.
Long contractual provisions therefore need careful chunking, and some context may be lost when a clause is split.
It is a good candidate for teams that value multilingual support, transparent model development, and smaller vector indexes.
6. Alibaba-NLP GTE-ModernBERT-Base
Model card: Alibaba-NLP/gte-modernbert-base on Hugging Face
- License: Apache 2.0
- Context length: 8,192 tokens
- Parameters: 149M
- Deployment: Open weights; SIE self-hosted catalog
If model size and latency are major concerns, GTE-ModernBERT-base is worth a look.
At only 149 million parameters, it is one of the smallest models in this comparison while still supporting an 8,192-token context window.
Alibaba reports:
- MTEB English: 64.38
- BEIR: 55.33
- LoCo: 87.57
- CoIR: 79.31
The LoCo result is particularly interesting for long-document retrieval.
Why it could work well for legal search
Interactive search has a different requirement from offline document processing.
When a lawyer enters a query, they expect the results quickly.
A smaller model can help reduce inference latency and make re-indexing significantly cheaper.
The main limitation is language coverage. This model is designed for English, so it is better suited to English-only legal collections.
If your contracts are primarily in English and infrastructure efficiency matters, this is one of the models I would include in the benchmark.
7. LightOn GTE-ModernColBERT-v1
Model card: lightonai/GTE-ModernColBERT-v1 on Hugging Face
- License: Apache 2.0
- Deployment: Open weights; available on SIE Cloud and self-hosted
GTE-ModernColBERT-v1 takes a different approach from standard dense embedding models.
Instead of creating a single vector for an entire chunk, it uses late interaction, keeping token-level representations that can be matched against the query.
This can be particularly useful for legal documents where two clauses may be almost identical except for one important phrase.
For example, imagine two indemnification clauses that differ only in one negotiated exception.
A single-vector model may consider the two clauses almost identical.
A late-interaction model has more opportunity to identify the specific token-level difference.
LightOn reports a BEIR average of 54.67 and a LongEmbed average of 88.39 at a 32K document length.
There is an important detail behind that second number, though.
The model was trained with a 300-token document length by default. Long-document configurations require changing that setting, and the model card recommends running your own evaluation for lengths beyond 8K.
The trade-off
Late interaction can improve fine-grained matching, but it also comes with a larger multi-vector index.
The model is also English-only.
So this is a particularly interesting candidate when your main problem is near-duplicate clauses and contract templates, rather than multilingual search.
How Do These Models Compare With Hosted APIs?
Open-weight models aren't the only way to build legal search.
Hosted embedding APIs can be easier to operate because you don't need to provision GPUs, manage model versions, or maintain inference infrastructure.
Some options worth considering include:
Voyage Law-2
Voyage AI offers a legal-specific embedding model designed for legal search.
The listed price is $0.12 per million tokens. Voyage also provides deployment options through AWS Marketplace, although availability of the specific legal model should be confirmed before planning an architecture around it.
Isaacus Kanon 2 Embedder
Isaacus focuses specifically on legal AI.
Kanon 2 Embedder is a proprietary model with no open weights. Isaacus reports an 86% NDCG@10 score on its MLEB legal benchmark.
The company also offers private and air-gapped deployment options for organizations with stricter security requirements.
OpenAI text-embedding-3-large
OpenAI's text-embedding-3-large is a common choice for teams that want a managed API without managing their own inference infrastructure.
Its listed price is $0.13 per million tokens.
Cohere Embed 5
Cohere's Embed 5 family supports a 128K-token context and offers private VPC and on-premises deployment options.
The choice between these hosted services and open-weight models ultimately comes down to a familiar trade-off:
Do you want the convenience of a managed API, or the control of running the model yourself?
What Does It Take to Move Away From OpenAI Embeddings?
Switching embedding models is more involved than changing a configuration value.
If your current application uses text-embedding-3-large, every vector in your existing index was generated using that model.
A new embedding model means generating new vectors for the entire archive.
The vector dimensions may also change.
For example:
-
text-embedding-3-large: 3,072 dimensions -
bge-m3: 1,024 dimensions -
Arctic Embed L v2.0: 1,024 dimensions
Depending on your vector database, you may need to create a completely new collection rather than updating the existing one.
A sensible migration looks something like this:
- Re-embed the existing document archive.
- Create a new vector index.
- Keep the old and new indexes running in parallel.
- Test both against the same evaluation dataset.
- Compare retrieval quality, latency, and cost.
- Gradually move users or tenants to the new index.
In other words, changing the embedding model is a migration project, not a model-setting change.
How to Run These Models
The models discussed above are available through the Superlinked Inference Engine model catalog, making it possible to evaluate different embedding models without rebuilding the surrounding search pipeline each time.
You can run SIE as:
- SIE Cloud, as a managed service
- Self-hosted SIE on your own Kubernetes environment
- A local setup for development and experimentation
The same inference layer can then be used to compare different models against the same documents and queries.
For example, you can run an encode call for each candidate model and use is_query=True when generating the query-side embedding.
On SIE Cloud, Superlinked lists models including Qwen3-Embedding-4B, Arctic Embed L v2.0, bge-m3, and GTE-ModernColBERT with different per-million-token pricing.
Self-hosting gives organizations more control over where inference happens, which can be particularly important when contracts cannot be sent to an external processor.
That makes the deployment decision just as important as the model decision for many legal-tech applications.
How We Ranked These Models
There isn't one neutral benchmark that can tell us which embedding model is the best for legal search.
MLEB is specifically focused on legal retrieval, but it is published by a vendor whose own model performs strongly on the benchmark.
MTEB, BEIR, MIRACL, and similar benchmarks are broader and aren't specifically designed around contracts.
So instead of treating one leaderboard as the final answer, the models were considered based on several factors:
- Published retrieval performance
- Context length
- Dense, sparse, or multi-vector capabilities
- Multilingual support
- Model size
- Deployment options
- Suitability for legal document retrieval
The benchmark numbers should therefore be treated as starting points rather than final rankings.
Your own contracts and queries are more important.
If you don't have a legal evaluation dataset yet, CUAD provides 510 lawyer-labelled contracts under a CC-BY-4.0 license and can be a useful starting point for building one.
FAQs
Is an open-weight model on SIE a good choice for legal search?
It can be.
A strong legal retrieval pipeline usually combines the embedding model with clause-level chunking, sparse retrieval, and reranking.
For a first comparison, bge-m3 and Qwen3-Embedding are good candidates. Run both against a few hundred real queries with clause-level relevance judgments rather than relying only on public benchmark scores.
The goal is to find out which model works best for your contracts.
Why run an open embedding model on SIE?
One advantage is having embeddings, reranking, OCR, and extraction available through the same inference layer.
For organizations with strict data requirements, self-hosted SIE can keep the inference workload within their own cloud or infrastructure.
It also makes experimentation easier. You can change the embedding model without having to redesign the entire retrieval architecture.
How much context does a legal embedding model need?
Ideally, enough to capture the complete clause and its relevant sub-clauses.
That makes 512-token models such as multilingual-e5-large-instruct and nomic-embed-text-v2-moe more dependent on careful chunking.
Models supporting 8K tokens, including bge-m3, Arctic Embed L, and GTE-ModernBERT, provide more room for longer clauses.
Qwen3-Embedding-8B goes further with a 32K context window.
But there is an important distinction:
A larger context window doesn't mean you should embed an entire contract as one chunk.
For legal search, meaningful clause-level chunks often provide a better balance between context and retrieval precision.
Where Should You Start?
For most legal-tech teams, I would start with three models:
bge-m3 is the strongest baseline when you want dense and sparse retrieval together, particularly for defined terms and exact legal language.
Qwen3-Embedding-8B is the high-end candidate to test when infrastructure isn't a major constraint and retrieval quality is the priority.
Arctic Embed L v2.0 is the efficiency-focused option, particularly for large archives where re-indexing speed and infrastructure cost matter.
I'd also add GTE-ModernColBERT-v1 if your biggest problem is retrieving the exact negotiated clause among several almost-identical templates.
Ultimately, though, there is no universal winner.
The best embedding model for legal search is the one that performs best on your contracts, your queries, and your definition of relevance.
The most practical approach is to evaluate several candidates through the same inference layer and measure:
- Recall
- NDCG@10
- Precision
- Query latency
- Indexing cost
- Storage requirements
- Performance on defined terms
- Performance on near-duplicate clauses
That evaluation will tell you much more than a leaderboard ever can.
And once you've identified the strongest candidate, the same inference layer can be used to serve it in production—whether you choose a managed environment such as SIE Cloud or run the stack inside your own infrastructure.

Top comments (0)