E-Governance for Vernacular India: Moving from Lexical to Sparse-Vector Hybrid Search
A citizen searching for a government scheme doesn’t necessarily search the way the government wrote the scheme.
They might search:
farmer income support scheme
Or:
kisan ko har saal paisa milne wali scheme
Or:
किसानों को हर साल पैसे देने वाली योजना
Or even:
farmer ko 6000 rupees wali yojna
The underlying intent can be almost identical.
The words aren’t.
And this becomes an interesting retrieval problem when the corpus contains thousands of government schemes.
For this project, I wanted to understand how far a traditional lexical search system can take us, where it starts breaking, and whether combining learned sparse and dense representations can make retrieval more robust for vernacular queries.
The project brief for this experiment specifically focuses on moving from a BM25-based legacy search setup toward sparse-vector and hybrid retrieval, while testing multilingual, transliterated and vocabulary-mismatched queries.
So instead of starting with embeddings immediately, I started with the simplest possible question:
How well can keyword search actually do?
1. The Dataset
I didn’t want to create synthetic government schemes.
For the retrieval corpus, I used the Indian Government MyScheme Dataset.
The dataset contains 2,066 government schemes: 527 Central Government schemes and 1,539 State/UT schemes. It provides 16 fields including scheme name, ministry/state, categories, description, eligibility, benefits, application process, documents required, FAQs and official links. The dataset is derived from the MyScheme portal and is provided in CSV and JSON formats.
That gives us something much more useful than a toy corpus.
The first thing I did was load the dataset.
import pandas as pd
DATA_PATH = "gov_myscheme_data.csv"
df = pd.read_csv(DATA_PATH)
print("Rows:", len(df))
print("Columns:", len(df.columns))
df.head()
Before building any retrieval system, I wanted to understand what was actually available.
df.columns.tolist()
The important fields looked roughly like:
Scheme Name
Scheme Slug
Level
State / UT / Ministry
Application Mode
Tags / Categories
Description
Eligibility Criteria
Eligibility (General)
Exclusions / Ineligibility
Benefits
Application Process
Documents Required
Frequently Asked Questions (FAQs)
Official Link
MyScheme URL
This matters because a government scheme isn’t represented by one sentence.
There is information distributed across multiple fields.
For retrieval, I therefore created a single searchable text representation.
TEXT_COLUMNS = [
"Scheme Name",
"Tags / Categories",
"Description",
"Eligibility Criteria",
"Eligibility (General)",
"Benefits",
"Application Process",
"Documents Required",
"Frequently Asked Questions (FAQs)"
]
for column in TEXT_COLUMNS:
df[column] = df[column].fillna("")
df["search_text"] = (
df[TEXT_COLUMNS]
.astype(str)
.agg(" ".join, axis=1)
)
Now each scheme has a searchable document:
Scheme Name
+
Categories
+
Description
+
Eligibility
+
Benefits
+
Application
+
FAQs
This becomes the common input to every retrieval approach.
2. The Vernacular Vocabulary Problem
Let’s take a hypothetical scheme document containing:
Pradhan Mantri Kisan Samman Nidhi
The scheme provides income support to eligible
farmer families through financial assistance
of ₹6,000 per year in three equal installments.
A user could search:
PM Kisan farmer income support
There is obvious lexical overlap.
But what about:
yearly cash transfer for cultivators
Now the words are different.
Or:
kisan ko har saal 6000 rupaye kaise milega
Now we’re dealing with transliterated Hindi.
Or:
किसानों को हर साल 6000 रुपये की सहायता
Now the script itself has changed.
This gives us several different retrieval problems.
And this is why I didn’t want to jump directly to vector search.
I wanted to establish a baseline first.
3. Building the BM25 Baseline
BM25 is still a very useful baseline for text retrieval.
The basic idea is straightforward:
Query
↓
Tokenize
↓
Find matching terms
↓
Score documents
↓
Rank results
I used rank_bm25 for the baseline.
from rank_bm25 import BM25Okapi
First, I needed a tokenizer that wouldn’t completely ignore Indic Unicode characters.
import re
TOKEN_RE = re.compile(
r"[A-Za-z\u0900-\u097F\u0B80-\u0BFF]+|\d+",
re.UNICODE
)
def tokenize(text):
return [
token.lower()
for token in TOKEN_RE.findall(str(text))
]
Then I built the corpus.
tokenized_corpus = [
tokenize(text)
for text in df["search_text"]
]
bm25 = BM25Okapi(tokenized_corpus)
And searching is simple:
def bm25_search(query, top_k=5):
query_tokens = tokenize(query)
scores = bm25.get_scores(
query_tokens
)
ranked_indices = sorted(
range(len(scores)),
key=lambda i: scores[i],
reverse=True
)[:top_k]
return df.iloc[ranked_indices]
Now we can test:
query = "farmer income support scheme"
results = bm25_search(
query,
top_k=5
)
results[
[
"Scheme Name",
"Benefits"
]
]
For a query containing the terminology used by the document, BM25 can work surprisingly well.
That’s not the problem.
The problem starts when the citizen doesn’t know the terminology.
4. Why Lexical Search Starts Breaking
Consider:
Document: financial assistance to eligible farmer families
and:
Query: yearly cash transfer for cultivators
BM25 is looking for:
yearly
cash
transfer
cultivators
The document might instead contain:
financial
assistance
farmer
families
The concepts are related.
The vocabulary isn’t.
A traditional inverted index is fundamentally built around matching terms.
So the retrieval path looks like:
Query Term
↓
Inverted Index
↓
Matching Documents
It doesn’t inherently understand:
cash transfer
≈
financial assistance
unless some additional mechanism provides that connection.
And this is where learned sparse retrieval becomes interesting.
5. Moving Beyond Exact Terms
Instead of representing text purely as a list of words, we can use a learned sparse model.
For the sparse retrieval experiment, the model used is:
SPARSE_MODEL = (
"prithivida/Splade_PP_en_v1"
)
Using FastEmbed:
from fastembed import SparseTextEmbedding
sparse_model = SparseTextEmbedding(
model_name=SPARSE_MODEL
)
Now we can encode our documents:
sparse_embeddings = list(
sparse_model.embed(
df["search_text"].tolist(),
batch_size=32
)
)
A dense embedding might look conceptually like:
[0.12, -0.08, 0.31, ...]
A sparse representation instead contains a relatively small number of active dimensions:
indices = [12, 183, 902, ...]
values = [0.21, 0.74, 0.32, ...]
The important idea is that the model learns which terms/features are important rather than treating every token equally.
So the retrieval problem becomes:
Government Document
↓
Learned Sparse Encoder
↓
Sparse representation
↓
Search
And for the query:
Citizen Query
↓
Learned Sparse Encoder
↓
Sparse representation
↓
Search
The representation can therefore capture useful lexical and learned expansion signals without requiring a huge dense vector for every document.
6. One Important Limitation
There is a subtle but important detail here.
The sparse model I used is:
prithivida/Splade_PP_en_v1
So I wouldn’t describe this as a universal multilingual solution.
That’s important because our problem is vernacular India.
A sparse model trained primarily around English retrieval and a multilingual dense model are solving different parts of the problem.
That gives us another question:
What happens when the user changes the language?
7. Adding Multilingual Dense Embeddings
For the dense representation, I used:
from sentence_transformers import SentenceTransformer
dense_model = SentenceTransformer(
"BAAI/bge-m3"
)
The corpus can now be embedded:
dense_embeddings = dense_model.encode(
df["search_text"].tolist(),
normalize_embeddings=True,
show_progress_bar=True
)
And a query:
query_embedding = dense_model.encode(
[query],
normalize_embeddings=True
)
Now the representation is no longer based purely on matching words.
The model is attempting to represent the semantic meaning of the text.
So these queries:
yearly financial support for farmers
and:
kisan ko har saal paisa milne wali scheme
can potentially occupy nearby regions in embedding space.
And that’s the important difference.
BM25
↓
Do the terms match?
Sparse
↓
What lexical features are important?
Dense
↓
What does this query mean?
None of these questions is identical.
8. Putting Multiple Representations Together
At this point I had:
I wanted these representations to live together.
That is where the database architecture becomes useful.
Instead of maintaining one system for lexical-style retrieval and another system for semantic retrieval, I can store multiple named vectors for the same point.
Conceptually:
Qdrant Point
│
├── Dense Vector
│
├── Sparse Vector
│
└── Payload
├── Scheme Name
├── Ministry
├── Category
├── Benefits
├── Eligibility
└── Official URL
Qdrant’s hybrid-search API supports multiple named vectors and prefetch queries, allowing different representations to be searched and their results fused.
9. Creating the Collection
The collection can be configured with both dense and sparse vectors.
from qdrant_client import QdrantClient, models
client = QdrantClient(
location=":memory:"
)
Then:
client.create_collection(
collection_name="gov_schemes",
vectors_config={
"dense": models.VectorParams(
size=dense_embeddings.shape[1],
distance=models.Distance.COSINE
)
},
sparse_vectors_config={
"sparse": models.SparseVectorParams(
index=models.SparseIndexParams(
on_disk=False
)
)
}
)
The important part is not the configuration itself.
It’s the representation.
One scheme can now have:
10. Preparing the Payload
The vectors are useful for retrieval.
But the actual scheme information still needs to come back.
So I stored the relevant metadata as payload.
points = []
for i, row in df.iterrows():
sparse_vector = sparse_embeddings[i]
points.append(
models.PointStruct(
id=i,
vector={
"dense": dense_embeddings[i].tolist(),
"sparse": models.SparseVector(
indices=sparse_vector.indices,
values=sparse_vector.values
)
},
payload={
"scheme_name": row["Scheme Name"],
"level": row["Level"],
"ministry": row["State / UT / Ministry"],
"categories": row["Tags / Categories"],
"description": row["Description"],
"benefits": row["Benefits"],
"eligibility": row["Eligibility Criteria"],
"application": row["Application Process"],
"official_url": row["Official Link"],
"myscheme_url": row["MyScheme URL"]
}
)
)
And finally:
client.upsert(
collection_name="gov_schemes",
points=points
)
Now retrieval and metadata live together.
11. Searching the Dense Representation
Let’s start with dense search.
query = (
"government scheme giving financial "
"support to farmers every year"
)
query_dense = encode_query(
query
)
Then:
for result in dense_results:
print(
result.score,
result.payload["scheme_name"]
)
This is where semantic retrieval starts becoming useful.
The user doesn’t have to know the exact phrase appearing inside the document.
12. Searching the Sparse Representation
The same query goes through the sparse encoder.
query_sparse = list(
sparse_model.embed(
[query]
)
)[0]
Then:
sparse_results = client.query_points(
collection_name="gov_schemes",
query=models.SparseVector(
indices=query_sparse.indices,
values=query_sparse.values
),
using="sparse",
limit=5,
with_payload=True
).points
Now we have two ranked lists.
Dense
-----
1. Scheme A
2. Scheme B
3. Scheme C
4. Scheme D
5. Scheme E
Sparse
------
1. Scheme B
2. Scheme A
3. Scheme F
4. Scheme C
5. Scheme G
The question becomes:
If both systems have useful information, why choose only one?
13. Why I Didn’t Just Add the Scores
A tempting implementation would be:
final_score = (
0.5 * dense_score
+
0.5 * sparse_score
)
But dense and sparse scores don’t necessarily live on the same scale.
A direct weighted sum could therefore be dominated by score magnitude rather than retrieval quality.
This is one of the reasons rank-based fusion is useful.
14. Reciprocal Rank Fusion
Reciprocal Rank Fusion, or RRF, doesn’t need the raw scores to be directly comparable.
It looks at where a document appears in each ranked list.
Conceptually:
RRF(d) = Σ 1 / (k + rank(d))
A document appearing near the top in multiple retrieval systems receives a stronger combined signal.
Qdrant’s hybrid-query documentation describes RRF as a way to fuse results from multiple representations and specifically discusses combining sparse and dense retrieval through prefetch.
The query therefore becomes:
hybrid_results = client.query_points(
collection_name="gov_schemes",
prefetch=[
models.Prefetch(
query=query_dense,
using="dense",
limit=20
),
models.Prefetch(
query=models.SparseVector(
indices=query_sparse.indices,
values=query_sparse.values
),
using="sparse",
limit=20
)
],
query=models.RrfQuery(
rrf=models.Rrf()
),
limit=5,
with_payload=True
).points
Now the architecture is:
And that is the core of the experiment.
15. Testing Vernacular Queries
Now comes the part I actually care about.
Instead of testing only:
farmer scheme
I created queries designed around different retrieval problems.
The important distinction is:
Real scheme records
+
Custom evaluation queries
The scheme data itself comes from the public MyScheme-derived dataset.
The additional queries are an evaluation set created to deliberately stress different retrieval behaviours.
16. Evaluating Retrieval
A search system shouldn’t be evaluated by looking at three results and saying:
“This looks good.”
We need measurable retrieval metrics.
For example, Recall@K:
def recall_at_k(
retrieved_ids,
relevant_id,
k=5
):
return int(
relevant_id
in retrieved_ids[:k]
)
MRR:
def reciprocal_rank(
retrieved_ids,
relevant_id
):
for rank, doc_id in enumerate(
retrieved_ids,
start=1
):
if doc_id == relevant_id:
return 1.0 / rank
return 0.0
And nDCG:
import math
def ndcg_at_k(
retrieved_ids,
relevant_id,
k=5
):
for rank, doc_id in enumerate(
retrieved_ids[:k],
start=1
):
if doc_id == relevant_id:
dcg = 1 / math.log2(
rank + 1
)
return dcg
return 0.0
The evaluation table can then look like:
Query Type BM25 Sparse Dense Hybrid
-------------------------------------------------------
English ... ... ... ...
Hindi ... ... ... ...
Hinglish ... ... ... ...
Misspelling ... ... ... ...
Vocabulary mismatch ... ... ... ...
I would keep the actual values generated by the notebook rather than hard-coding illustrative numbers into the article.
That way the Medium article and GitHub repository remain reproducible.
17. What Each Retrieval Method Is Actually Doing
After running these experiments, the useful way to think about the systems isn’t:
BM25 vs Sparse vs Dense
as if one has to replace everything else.
It is:
BM25 is useful when terminology overlaps.
Sparse retrieval can provide a different learned lexical signal.
Dense retrieval gives us semantic and multilingual representation.
Hybrid retrieval allows these signals to participate in the same ranking process.
That’s a much more useful mental model than trying to find one magical retrieval technique.
18. The Architecture I Would Actually Deploy
The project brief describes the production problem as a migration from legacy Elasticsearch/Solr-style BM25 retrieval toward a sparse-vector hybrid architecture.
I wouldn’t approach that as:
Friday:
Elasticsearch
Monday:
Everything replaced
That’s asking for trouble.
Instead, I’d introduce the new retrieval path alongside the existing one.
Initially, the new system doesn’t have to control what the user sees.
It can run in shadow mode.
19. Making the Migration Reversible
This was the part I wanted to make clearer.
Instead of saying:
1% → 5% → 10% → 25% → 50% → 100%
without explaining what those numbers mean, think about it as traffic ownership.
Suppose the existing search system is serving 100% of users.
We introduce the new system.
Stage 1 — Shadow
100% users
|
+------> Legacy Search ------> User
|
+------> New Retrieval
|
Logged only
The new system receives the same queries, but its result is not shown to users.
We compare:
latency
retrieval quality
errors
empty results
ranking differences
Stage 2 — Small Live Slice
Once the shadow results look reasonable:
99% ──> Legacy
1% ──> New
Now 1% of real traffic actually sees the new ranking.
Stage 3 — Increase Ownership
If monitoring remains healthy:
95% ──> Legacy
5% ──> New
Then:
90% ──> Legacy
10% ──> New
Then:
75% ──> Legacy
25% ──> New
Then:
50% ──> Legacy
50% ──> New
Eventually:
100% ──> New
At every stage, the legacy system remains available as a fallback.
So the actual architecture is:
That’s what makes the migration reversible.
The percentages aren’t the architecture.
The ability to move traffic in either direction is.
20. What Happens Inside the New Search Path
The final retrieval architecture now looks like:
The advantage isn’t that every query suddenly becomes perfect.
The advantage is that the system has more than one way of understanding the query.
21. Why the Dataset Structure Matters
One thing I liked about using an actual government-scheme dataset is that retrieval isn’t happening over random paragraphs.
A scheme contains structured information.
For example:
Scheme Name
Benefits
Eligibility
Exclusions
Application Process
Documents
FAQs
That means the same retrieval system can eventually support more specific questions.
For example:
"Which farmer schemes provide financial assistance?"
might rely heavily on:
Benefits
Tags
Description
While:
"Who is eligible for PM-KISAN?"
can benefit from:
Eligibility Criteria
Eligibility (General)
Exclusions
FAQs
And:
"How do I apply?"
can retrieve:
Application Process
Documents Required
Official Link
So the searchable document isn’t just:
Scheme Name + Description
It can represent the entire information structure of the scheme.
22. The Bigger Picture
The interesting thing about vernacular e-governance isn’t simply that people speak different languages.
It’s that people also describe the same requirement differently.
Consider:
"scheme for women entrepreneurs"
"mahila business ke liye government help"
"महिलाओं को व्यवसाय शुरू करने के लिए सरकारी सहायता"
Three surface forms.
Potentially one intent.
A search system that depends entirely on exact terminology puts part of the burden back on the citizen.
A retrieval system should instead try to bridge:
Citizen language
↓
Retrieval representation
↓
Government terminology
And that’s where multiple retrieval representations become interesting.
23. What I Would Measure Before Calling This Production-Ready
A retrieval demo can return plausible results.
Production search needs more evidence.
I’d track at least:
metrics = {
"recall_at_3": [],
"recall_at_5": [],
"precision_at_5": [],
"mrr": [],
"ndcg_at_5": [],
"latency_ms": []
}
And evaluate by query type:
English
Hindi
Hinglish
Misspellings
Vocabulary mismatch
Long natural-language queries
Not just aggregate performance.
Because an aggregate score can hide a very important failure.
For example:
Overall Recall@5
0.90
sounds good.
But if:
English 0.96
Hindi 0.95
Hinglish 0.91
Misspellings 0.88
and:
Hard mismatch
0.62
then the system still has a specific retrieval weakness.
That’s much more useful information than one headline number.
24. The Final Architecture
After going through the entire experiment, the architecture is no longer:
Query
↓
Embedding
↓
Nearest Neighbours
It becomes:
And underneath that:
Qdrant’s documentation supports this pattern of multiple named vectors, prefetch, and rank fusion such as RRF for combining retrieval signals.
25. What I Actually Learned
The biggest takeaway from this project wasn’t:
“Qdrant is better than Elasticsearch.”
That isn’t what this experiment proves.
And it wasn’t:
“Dense vectors are always better.”
The results don’t support that either.
What the experiment actually showed me was something more interesting.
A search query can fail for completely different reasons.
Sometimes the user uses the same vocabulary as the document.
PM Kisan farmer income support
BM25 can handle that.
Sometimes the vocabulary changes.
yearly cash transfer for cultivators
That’s where learned sparse retrieval can provide a different signal.
Sometimes the language or script changes.
किसानों को हर साल छह हजार रुपये की सहायता
That’s where multilingual dense retrieval becomes much more important.
And sometimes several of these things happen simultaneously.
That’s where hybrid retrieval becomes useful.
So instead of thinking:
BM25 vs Vector Search
I started thinking:
Search Problem
|
+-------------+-------------+
| | |
Vocabulary Semantics Language
| | |
v v v
Sparse Dense Dense
| | |
+-------------+-------------+
|
Hybrid
That’s a much more useful mental model.
Conclusion
I started this project with a fairly simple question:
Can keyword search handle how people actually search for government schemes?
The answer isn’t simply yes or no.
BM25 remains useful when the query and document share terminology.
But vernacular search introduces other problems:
Different vocabulary
Different language
Different script
Transliteration
Misspellings
Paraphrasing
That changes the retrieval problem.
A learned sparse representation gives us another way to capture important lexical signals.
A multilingual dense representation gives us a way to model semantic relationships across languages.
And hybrid retrieval gives us a mechanism to combine those signals instead of forcing every query through a single representation.
The progression therefore becomes:
BM25
↓
Learned Sparse Retrieval
↓
Multilingual Dense Retrieval
↓
Hybrid Ranking
But the bigger lesson for me was that search quality is less about choosing one representation and more about understanding why a query failed in the first place.
If the user uses the exact government terminology, lexical retrieval can work.
If they use different terminology, learned sparse retrieval can provide another signal.
If they switch languages or scripts, multilingual dense retrieval becomes more relevant.
And when we don’t know which of these situations we’re going to get, combining the signals gives us a retrieval system that is less dependent on one particular way of asking a question.
For e-governance, that’s important.
The citizen shouldn’t have to know whether the official document calls something:
income support
or:
financial assistance
or:
cash transfer
They should be able to describe what they need in the language they naturally use.
The retrieval system should handle the translation between how people ask and how information is stored.
That’s the problem I wanted to explore with this project.
References
Dataset — Indian Government MyScheme Dataset
Aryan-Pardeshi/gov-myscheme-dataset — 2,066 structured government schemes derived from MyScheme.Government of India — MyScheme
myscheme.gov.inQdrant — Hybrid and Multi-Stage Queries
Qdrant Hybrid Queries Documentation — multiple named vectors, prefetch and RRF.Qdrant — Hybrid Search
Combining Semantic and Lexical SearchQdrant — FastEmbed
FastEmbed DocumentationBAAI — BGE-M3
BAAI/bge-m3 on Hugging FaceSPLADE PP en v1
prithivida/Splade_PP_en_v1AI4Bharat — IndicQA
IndicQA DatasetProject Repository — E-Governance
github.com/AKR4PC/E-Governance
















Top comments (0)