DEV Community

Cover image for E-Governance for Vernacular India: Moving from Lexical to Sparse-Vector Hybrid Search
AKSHAT KUMAR
AKSHAT KUMAR

Posted on

E-Governance for Vernacular India: Moving from Lexical to Sparse-Vector Hybrid Search

E-Governance for Vernacular India: Moving from Lexical to Sparse-Vector Hybrid Search

A citizen searching for a government scheme doesn’t necessarily search the way the government wrote the scheme.

Vernacular government scheme search

They might search:

farmer income support scheme
Enter fullscreen mode Exit fullscreen mode

Or:

kisan ko har saal paisa milne wali scheme
Enter fullscreen mode Exit fullscreen mode

Or:

किसानों को हर साल पैसे देने वाली योजना
Enter fullscreen mode Exit fullscreen mode

Or even:

farmer ko 6000 rupees wali yojna
Enter fullscreen mode Exit fullscreen mode

The underlying intent can be almost identical.

The words aren’t.

And this becomes an interesting retrieval problem when the corpus contains thousands of government schemes.

For this project, I wanted to understand how far a traditional lexical search system can take us, where it starts breaking, and whether combining learned sparse and dense representations can make retrieval more robust for vernacular queries.

The project brief for this experiment specifically focuses on moving from a BM25-based legacy search setup toward sparse-vector and hybrid retrieval, while testing multilingual, transliterated and vocabulary-mismatched queries.

So instead of starting with embeddings immediately, I started with the simplest possible question:

How well can keyword search actually do?


1. The Dataset

I didn’t want to create synthetic government schemes.

For the retrieval corpus, I used the Indian Government MyScheme Dataset.

The dataset contains 2,066 government schemes: 527 Central Government schemes and 1,539 State/UT schemes. It provides 16 fields including scheme name, ministry/state, categories, description, eligibility, benefits, application process, documents required, FAQs and official links. The dataset is derived from the MyScheme portal and is provided in CSV and JSON formats.

That gives us something much more useful than a toy corpus.

Indian Government MyScheme Dataset

The first thing I did was load the dataset.

import pandas as pd

DATA_PATH = "gov_myscheme_data.csv"

df = pd.read_csv(DATA_PATH)

print("Rows:", len(df))
print("Columns:", len(df.columns))

df.head()
Enter fullscreen mode Exit fullscreen mode

Before building any retrieval system, I wanted to understand what was actually available.

df.columns.tolist()
Enter fullscreen mode Exit fullscreen mode

The important fields looked roughly like:

Scheme Name
Scheme Slug
Level
State / UT / Ministry
Application Mode
Tags / Categories
Description
Eligibility Criteria
Eligibility (General)
Exclusions / Ineligibility
Benefits
Application Process
Documents Required
Frequently Asked Questions (FAQs)
Official Link
MyScheme URL
Enter fullscreen mode Exit fullscreen mode

This matters because a government scheme isn’t represented by one sentence.

There is information distributed across multiple fields.

For retrieval, I therefore created a single searchable text representation.

TEXT_COLUMNS = [
    "Scheme Name",
    "Tags / Categories",
    "Description",
    "Eligibility Criteria",
    "Eligibility (General)",
    "Benefits",
    "Application Process",
    "Documents Required",
    "Frequently Asked Questions (FAQs)"
]

for column in TEXT_COLUMNS:
    df[column] = df[column].fillna("")

df["search_text"] = (
    df[TEXT_COLUMNS]
    .astype(str)
    .agg(" ".join, axis=1)
)
Enter fullscreen mode Exit fullscreen mode

Now each scheme has a searchable document:

Scheme Name
+
Categories
+
Description
+
Eligibility
+
Benefits
+
Application
+
FAQs
Enter fullscreen mode Exit fullscreen mode

This becomes the common input to every retrieval approach.


2. The Vernacular Vocabulary Problem

Let’s take a hypothetical scheme document containing:

Pradhan Mantri Kisan Samman Nidhi

The scheme provides income support to eligible
farmer families through financial assistance
of ₹6,000 per year in three equal installments.
Enter fullscreen mode Exit fullscreen mode

A user could search:

PM Kisan farmer income support
Enter fullscreen mode Exit fullscreen mode

There is obvious lexical overlap.

But what about:

yearly cash transfer for cultivators
Enter fullscreen mode Exit fullscreen mode

Now the words are different.

Or:

kisan ko har saal 6000 rupaye kaise milega
Enter fullscreen mode Exit fullscreen mode

Now we’re dealing with transliterated Hindi.

Or:

किसानों को हर साल 6000 रुपये की सहायता
Enter fullscreen mode Exit fullscreen mode

Now the script itself has changed.

This gives us several different retrieval problems.

Vernacular search problem

And this is why I didn’t want to jump directly to vector search.

I wanted to establish a baseline first.


3. Building the BM25 Baseline

BM25 is still a very useful baseline for text retrieval.

The basic idea is straightforward:

Query
  ↓
Tokenize
  ↓
Find matching terms
  ↓
Score documents
  ↓
Rank results
Enter fullscreen mode Exit fullscreen mode

I used rank_bm25 for the baseline.

from rank_bm25 import BM25Okapi
Enter fullscreen mode Exit fullscreen mode

First, I needed a tokenizer that wouldn’t completely ignore Indic Unicode characters.

import re

TOKEN_RE = re.compile(
    r"[A-Za-z\u0900-\u097F\u0B80-\u0BFF]+|\d+",
    re.UNICODE
)

def tokenize(text):
    return [
        token.lower()
        for token in TOKEN_RE.findall(str(text))
    ]
Enter fullscreen mode Exit fullscreen mode

Then I built the corpus.

tokenized_corpus = [
    tokenize(text)
    for text in df["search_text"]
]

bm25 = BM25Okapi(tokenized_corpus)
Enter fullscreen mode Exit fullscreen mode

And searching is simple:

def bm25_search(query, top_k=5):

    query_tokens = tokenize(query)

    scores = bm25.get_scores(
        query_tokens
    )

    ranked_indices = sorted(
        range(len(scores)),
        key=lambda i: scores[i],
        reverse=True
    )[:top_k]

    return df.iloc[ranked_indices]
Enter fullscreen mode Exit fullscreen mode

Now we can test:

query = "farmer income support scheme"

results = bm25_search(
    query,
    top_k=5
)

results[
    [
        "Scheme Name",
        "Benefits"
    ]
]
Enter fullscreen mode Exit fullscreen mode

For a query containing the terminology used by the document, BM25 can work surprisingly well.

That’s not the problem.

The problem starts when the citizen doesn’t know the terminology.


4. Why Lexical Search Starts Breaking

Consider:

Document: financial assistance to eligible farmer families
Enter fullscreen mode Exit fullscreen mode

and:

Query: yearly cash transfer for cultivators
Enter fullscreen mode Exit fullscreen mode

BM25 is looking for:

yearly
cash
transfer
cultivators
Enter fullscreen mode Exit fullscreen mode

The document might instead contain:

financial
assistance
farmer
families
Enter fullscreen mode Exit fullscreen mode

The concepts are related.

The vocabulary isn’t.

A traditional inverted index is fundamentally built around matching terms.

So the retrieval path looks like:

Query Term
    ↓
Inverted Index
    ↓
Matching Documents
Enter fullscreen mode Exit fullscreen mode

It doesn’t inherently understand:

cash transfer
      ≈
financial assistance
Enter fullscreen mode Exit fullscreen mode

unless some additional mechanism provides that connection.

And this is where learned sparse retrieval becomes interesting.


5. Moving Beyond Exact Terms

Instead of representing text purely as a list of words, we can use a learned sparse model.

For the sparse retrieval experiment, the model used is:

SPARSE_MODEL = (
    "prithivida/Splade_PP_en_v1"
)
Enter fullscreen mode Exit fullscreen mode

Using FastEmbed:

from fastembed import SparseTextEmbedding

sparse_model = SparseTextEmbedding(
    model_name=SPARSE_MODEL
)
Enter fullscreen mode Exit fullscreen mode

Now we can encode our documents:

sparse_embeddings = list(
    sparse_model.embed(
        df["search_text"].tolist(),
        batch_size=32
    )
)
Enter fullscreen mode Exit fullscreen mode

A dense embedding might look conceptually like:

[0.12, -0.08, 0.31, ...]
Enter fullscreen mode Exit fullscreen mode

A sparse representation instead contains a relatively small number of active dimensions:

indices = [12, 183, 902, ...]
values  = [0.21, 0.74, 0.32, ...]
Enter fullscreen mode Exit fullscreen mode

The important idea is that the model learns which terms/features are important rather than treating every token equally.

So the retrieval problem becomes:

Government Document
        ↓
 Learned Sparse Encoder
        ↓
Sparse representation
        ↓
       Search
Enter fullscreen mode Exit fullscreen mode

And for the query:

Citizen Query
      ↓
Learned Sparse Encoder
      ↓
Sparse representation
      ↓
      Search
Enter fullscreen mode Exit fullscreen mode

The representation can therefore capture useful lexical and learned expansion signals without requiring a huge dense vector for every document.


6. One Important Limitation

There is a subtle but important detail here.

The sparse model I used is:

prithivida/Splade_PP_en_v1
Enter fullscreen mode Exit fullscreen mode

So I wouldn’t describe this as a universal multilingual solution.

That’s important because our problem is vernacular India.

A sparse model trained primarily around English retrieval and a multilingual dense model are solving different parts of the problem.

That gives us another question:

What happens when the user changes the language?


7. Adding Multilingual Dense Embeddings

For the dense representation, I used:

from sentence_transformers import SentenceTransformer

dense_model = SentenceTransformer(
    "BAAI/bge-m3"
)
Enter fullscreen mode Exit fullscreen mode

The corpus can now be embedded:

dense_embeddings = dense_model.encode(
    df["search_text"].tolist(),
    normalize_embeddings=True,
    show_progress_bar=True
)
Enter fullscreen mode Exit fullscreen mode

And a query:

query_embedding = dense_model.encode(
    [query],
    normalize_embeddings=True
)
Enter fullscreen mode Exit fullscreen mode

Now the representation is no longer based purely on matching words.

The model is attempting to represent the semantic meaning of the text.

So these queries:

yearly financial support for farmers
Enter fullscreen mode Exit fullscreen mode

and:

kisan ko har saal paisa milne wali scheme
Enter fullscreen mode Exit fullscreen mode

can potentially occupy nearby regions in embedding space.

And that’s the important difference.

BM25
    ↓
Do the terms match?

Sparse
    ↓
What lexical features are important?

Dense
    ↓
What does this query mean?
Enter fullscreen mode Exit fullscreen mode

None of these questions is identical.

Dense semantic retrieval


8. Putting Multiple Representations Together

At this point I had:

Multiple retrieval representations

I wanted these representations to live together.

That is where the database architecture becomes useful.

Instead of maintaining one system for lexical-style retrieval and another system for semantic retrieval, I can store multiple named vectors for the same point.

Conceptually:

Qdrant Point
│
├── Dense Vector
│
├── Sparse Vector
│
└── Payload
    ├── Scheme Name
    ├── Ministry
    ├── Category
    ├── Benefits
    ├── Eligibility
    └── Official URL
Enter fullscreen mode Exit fullscreen mode

Qdrant’s hybrid-search API supports multiple named vectors and prefetch queries, allowing different representations to be searched and their results fused.


9. Creating the Collection

The collection can be configured with both dense and sparse vectors.

from qdrant_client import QdrantClient, models

client = QdrantClient(
    location=":memory:"
)
Enter fullscreen mode Exit fullscreen mode

Then:

client.create_collection(
    collection_name="gov_schemes",

    vectors_config={
        "dense": models.VectorParams(
            size=dense_embeddings.shape[1],
            distance=models.Distance.COSINE
        )
    },

    sparse_vectors_config={
        "sparse": models.SparseVectorParams(
            index=models.SparseIndexParams(
                on_disk=False
            )
        )
    }
)
Enter fullscreen mode Exit fullscreen mode

The important part is not the configuration itself.

It’s the representation.

One scheme can now have:

Qdrant multi-vector representation


10. Preparing the Payload

The vectors are useful for retrieval.

But the actual scheme information still needs to come back.

So I stored the relevant metadata as payload.

points = []

for i, row in df.iterrows():

    sparse_vector = sparse_embeddings[i]

    points.append(
        models.PointStruct(

            id=i,

            vector={
                "dense": dense_embeddings[i].tolist(),

                "sparse": models.SparseVector(
                    indices=sparse_vector.indices,
                    values=sparse_vector.values
                )
            },

            payload={
                "scheme_name": row["Scheme Name"],
                "level": row["Level"],
                "ministry": row["State / UT / Ministry"],
                "categories": row["Tags / Categories"],
                "description": row["Description"],
                "benefits": row["Benefits"],
                "eligibility": row["Eligibility Criteria"],
                "application": row["Application Process"],
                "official_url": row["Official Link"],
                "myscheme_url": row["MyScheme URL"]
            }
        )
    )
Enter fullscreen mode Exit fullscreen mode

And finally:

client.upsert(
    collection_name="gov_schemes",
    points=points
)
Enter fullscreen mode Exit fullscreen mode

Now retrieval and metadata live together.


11. Searching the Dense Representation

Let’s start with dense search.

query = (
    "government scheme giving financial "
    "support to farmers every year"
)

query_dense = encode_query(
    query
)
Enter fullscreen mode Exit fullscreen mode

Then:

for result in dense_results:

    print(
        result.score,
        result.payload["scheme_name"]
    )
Enter fullscreen mode Exit fullscreen mode

This is where semantic retrieval starts becoming useful.

The user doesn’t have to know the exact phrase appearing inside the document.


12. Searching the Sparse Representation

The same query goes through the sparse encoder.

query_sparse = list(
    sparse_model.embed(
        [query]
    )
)[0]
Enter fullscreen mode Exit fullscreen mode

Then:

sparse_results = client.query_points(

    collection_name="gov_schemes",

    query=models.SparseVector(
        indices=query_sparse.indices,
        values=query_sparse.values
    ),

    using="sparse",

    limit=5,

    with_payload=True

).points
Enter fullscreen mode Exit fullscreen mode

Now we have two ranked lists.

Dense
-----
1. Scheme A
2. Scheme B
3. Scheme C
4. Scheme D
5. Scheme E


Sparse
------
1. Scheme B
2. Scheme A
3. Scheme F
4. Scheme C
5. Scheme G
Enter fullscreen mode Exit fullscreen mode

The question becomes:

If both systems have useful information, why choose only one?


13. Why I Didn’t Just Add the Scores

A tempting implementation would be:

final_score = (
    0.5 * dense_score
    +
    0.5 * sparse_score
)
Enter fullscreen mode Exit fullscreen mode

But dense and sparse scores don’t necessarily live on the same scale.

A direct weighted sum could therefore be dominated by score magnitude rather than retrieval quality.

This is one of the reasons rank-based fusion is useful.


14. Reciprocal Rank Fusion

Reciprocal Rank Fusion, or RRF, doesn’t need the raw scores to be directly comparable.

It looks at where a document appears in each ranked list.

Conceptually:

RRF(d) = Σ 1 / (k + rank(d))
Enter fullscreen mode Exit fullscreen mode

A document appearing near the top in multiple retrieval systems receives a stronger combined signal.

Qdrant’s hybrid-query documentation describes RRF as a way to fuse results from multiple representations and specifically discusses combining sparse and dense retrieval through prefetch.

The query therefore becomes:

hybrid_results = client.query_points(

    collection_name="gov_schemes",

    prefetch=[

        models.Prefetch(
            query=query_dense,
            using="dense",
            limit=20
        ),

        models.Prefetch(
            query=models.SparseVector(
                indices=query_sparse.indices,
                values=query_sparse.values
            ),
            using="sparse",
            limit=20
        )

    ],

    query=models.RrfQuery(
        rrf=models.Rrf()
    ),

    limit=5,

    with_payload=True

).points
Enter fullscreen mode Exit fullscreen mode

Now the architecture is:

Hybrid retrieval architecture

And that is the core of the experiment.


15. Testing Vernacular Queries

Now comes the part I actually care about.

Instead of testing only:

farmer scheme
Enter fullscreen mode Exit fullscreen mode

I created queries designed around different retrieval problems.

The important distinction is:

Real scheme records
        +
Custom evaluation queries
Enter fullscreen mode Exit fullscreen mode

The scheme data itself comes from the public MyScheme-derived dataset.

The additional queries are an evaluation set created to deliberately stress different retrieval behaviours.


16. Evaluating Retrieval

A search system shouldn’t be evaluated by looking at three results and saying:

“This looks good.”

We need measurable retrieval metrics.

For example, Recall@K:

def recall_at_k(
    retrieved_ids,
    relevant_id,
    k=5
):

    return int(
        relevant_id
        in retrieved_ids[:k]
    )
Enter fullscreen mode Exit fullscreen mode

MRR:

def reciprocal_rank(
    retrieved_ids,
    relevant_id
):

    for rank, doc_id in enumerate(
        retrieved_ids,
        start=1
    ):

        if doc_id == relevant_id:
            return 1.0 / rank

    return 0.0
Enter fullscreen mode Exit fullscreen mode

And nDCG:

import math

def ndcg_at_k(
    retrieved_ids,
    relevant_id,
    k=5
):

    for rank, doc_id in enumerate(
        retrieved_ids[:k],
        start=1
    ):

        if doc_id == relevant_id:

            dcg = 1 / math.log2(
                rank + 1
            )

            return dcg

    return 0.0
Enter fullscreen mode Exit fullscreen mode

The evaluation table can then look like:

Query Type          BM25    Sparse    Dense    Hybrid
-------------------------------------------------------
English              ...      ...      ...       ...
Hindi                ...      ...      ...       ...
Hinglish             ...      ...      ...       ...
Misspelling          ...      ...      ...       ...
Vocabulary mismatch  ...      ...      ...       ...
Enter fullscreen mode Exit fullscreen mode

I would keep the actual values generated by the notebook rather than hard-coding illustrative numbers into the article.

That way the Medium article and GitHub repository remain reproducible.


17. What Each Retrieval Method Is Actually Doing

After running these experiments, the useful way to think about the systems isn’t:

BM25 vs Sparse vs Dense
Enter fullscreen mode Exit fullscreen mode

as if one has to replace everything else.

It is:

Retrieval methods comparison

BM25 is useful when terminology overlaps.

Sparse retrieval can provide a different learned lexical signal.

Dense retrieval gives us semantic and multilingual representation.

Hybrid retrieval allows these signals to participate in the same ranking process.

That’s a much more useful mental model than trying to find one magical retrieval technique.


18. The Architecture I Would Actually Deploy

The project brief describes the production problem as a migration from legacy Elasticsearch/Solr-style BM25 retrieval toward a sparse-vector hybrid architecture.

I wouldn’t approach that as:

Friday:
Elasticsearch

Monday:
Everything replaced
Enter fullscreen mode Exit fullscreen mode

That’s asking for trouble.

Instead, I’d introduce the new retrieval path alongside the existing one.

Production migration architecture

Initially, the new system doesn’t have to control what the user sees.

It can run in shadow mode.


19. Making the Migration Reversible

This was the part I wanted to make clearer.

Instead of saying:

1% → 5% → 10% → 25% → 50% → 100%
Enter fullscreen mode Exit fullscreen mode

without explaining what those numbers mean, think about it as traffic ownership.

Suppose the existing search system is serving 100% of users.

We introduce the new system.

Stage 1 — Shadow

100% users
     |
     +------> Legacy Search ------> User
     |
     +------> New Retrieval
                 |
             Logged only
Enter fullscreen mode Exit fullscreen mode

The new system receives the same queries, but its result is not shown to users.

We compare:

latency
retrieval quality
errors
empty results
ranking differences
Enter fullscreen mode Exit fullscreen mode

Stage 2 — Small Live Slice

Once the shadow results look reasonable:

99% ──> Legacy
 1% ──> New
Enter fullscreen mode Exit fullscreen mode

Now 1% of real traffic actually sees the new ranking.

Stage 3 — Increase Ownership

If monitoring remains healthy:

95% ──> Legacy
 5% ──> New
Enter fullscreen mode Exit fullscreen mode

Then:

90% ──> Legacy
10% ──> New
Enter fullscreen mode Exit fullscreen mode

Then:

75% ──> Legacy
25% ──> New
Enter fullscreen mode Exit fullscreen mode

Then:

50% ──> Legacy
50% ──> New
Enter fullscreen mode Exit fullscreen mode

Eventually:

100% ──> New
Enter fullscreen mode Exit fullscreen mode

At every stage, the legacy system remains available as a fallback.

So the actual architecture is:

Reversible search migration

That’s what makes the migration reversible.

The percentages aren’t the architecture.

The ability to move traffic in either direction is.

Traffic migration stages


20. What Happens Inside the New Search Path

The final retrieval architecture now looks like:

New search architecture

The advantage isn’t that every query suddenly becomes perfect.

The advantage is that the system has more than one way of understanding the query.

Multiple retrieval signals


21. Why the Dataset Structure Matters

One thing I liked about using an actual government-scheme dataset is that retrieval isn’t happening over random paragraphs.

A scheme contains structured information.

For example:

Scheme Name
Benefits
Eligibility
Exclusions
Application Process
Documents
FAQs
Enter fullscreen mode Exit fullscreen mode

That means the same retrieval system can eventually support more specific questions.

For example:

"Which farmer schemes provide financial assistance?"
Enter fullscreen mode Exit fullscreen mode

might rely heavily on:

Benefits
Tags
Description
Enter fullscreen mode Exit fullscreen mode

While:

"Who is eligible for PM-KISAN?"
Enter fullscreen mode Exit fullscreen mode

can benefit from:

Eligibility Criteria
Eligibility (General)
Exclusions
FAQs
Enter fullscreen mode Exit fullscreen mode

And:

"How do I apply?"
Enter fullscreen mode Exit fullscreen mode

can retrieve:

Application Process
Documents Required
Official Link
Enter fullscreen mode Exit fullscreen mode

So the searchable document isn’t just:

Scheme Name + Description
Enter fullscreen mode Exit fullscreen mode

It can represent the entire information structure of the scheme.


22. The Bigger Picture

The interesting thing about vernacular e-governance isn’t simply that people speak different languages.

It’s that people also describe the same requirement differently.

Consider:

"scheme for women entrepreneurs"
Enter fullscreen mode Exit fullscreen mode
"mahila business ke liye government help"
Enter fullscreen mode Exit fullscreen mode
"महिलाओं को व्यवसाय शुरू करने के लिए सरकारी सहायता"
Enter fullscreen mode Exit fullscreen mode

Three surface forms.

Potentially one intent.

A search system that depends entirely on exact terminology puts part of the burden back on the citizen.

A retrieval system should instead try to bridge:

Citizen language
      ↓
Retrieval representation
      ↓
Government terminology
Enter fullscreen mode Exit fullscreen mode

And that’s where multiple retrieval representations become interesting.

Vernacular e-governance retrieval


23. What I Would Measure Before Calling This Production-Ready

A retrieval demo can return plausible results.

Production search needs more evidence.

I’d track at least:

metrics = {
    "recall_at_3": [],
    "recall_at_5": [],
    "precision_at_5": [],
    "mrr": [],
    "ndcg_at_5": [],
    "latency_ms": []
}
Enter fullscreen mode Exit fullscreen mode

And evaluate by query type:

English
Hindi
Hinglish
Misspellings
Vocabulary mismatch
Long natural-language queries
Enter fullscreen mode Exit fullscreen mode

Not just aggregate performance.

Because an aggregate score can hide a very important failure.

For example:

Overall Recall@5
        0.90
Enter fullscreen mode Exit fullscreen mode

sounds good.

But if:

English      0.96
Hindi        0.95
Hinglish     0.91
Misspellings 0.88
Enter fullscreen mode Exit fullscreen mode

and:

Hard mismatch
0.62
Enter fullscreen mode Exit fullscreen mode

then the system still has a specific retrieval weakness.

That’s much more useful information than one headline number.


24. The Final Architecture

After going through the entire experiment, the architecture is no longer:

Query
  ↓
Embedding
  ↓
Nearest Neighbours
Enter fullscreen mode Exit fullscreen mode

It becomes:

Final hybrid retrieval architecture

And underneath that:

Final system architecture

Qdrant’s documentation supports this pattern of multiple named vectors, prefetch, and rank fusion such as RRF for combining retrieval signals.


25. What I Actually Learned

The biggest takeaway from this project wasn’t:

“Qdrant is better than Elasticsearch.”

That isn’t what this experiment proves.

And it wasn’t:

“Dense vectors are always better.”

The results don’t support that either.

What the experiment actually showed me was something more interesting.

A search query can fail for completely different reasons.

Sometimes the user uses the same vocabulary as the document.

PM Kisan farmer income support
Enter fullscreen mode Exit fullscreen mode

BM25 can handle that.

Sometimes the vocabulary changes.

yearly cash transfer for cultivators
Enter fullscreen mode Exit fullscreen mode

That’s where learned sparse retrieval can provide a different signal.

Sometimes the language or script changes.

किसानों को हर साल छह हजार रुपये की सहायता
Enter fullscreen mode Exit fullscreen mode

That’s where multilingual dense retrieval becomes much more important.

And sometimes several of these things happen simultaneously.

That’s where hybrid retrieval becomes useful.

So instead of thinking:

BM25 vs Vector Search
Enter fullscreen mode Exit fullscreen mode

I started thinking:

Search Problem
                     |
       +-------------+-------------+
       |             |             |
   Vocabulary     Semantics     Language
       |             |             |
       v             v             v
    Sparse         Dense        Dense
       |             |             |
       +-------------+-------------+
                     |
                   Hybrid
Enter fullscreen mode Exit fullscreen mode

That’s a much more useful mental model.


Conclusion

I started this project with a fairly simple question:

Can keyword search handle how people actually search for government schemes?

The answer isn’t simply yes or no.

BM25 remains useful when the query and document share terminology.

But vernacular search introduces other problems:

Different vocabulary
Different language
Different script
Transliteration
Misspellings
Paraphrasing
Enter fullscreen mode Exit fullscreen mode

That changes the retrieval problem.

A learned sparse representation gives us another way to capture important lexical signals.

A multilingual dense representation gives us a way to model semantic relationships across languages.

And hybrid retrieval gives us a mechanism to combine those signals instead of forcing every query through a single representation.

The progression therefore becomes:

BM25
  ↓
Learned Sparse Retrieval
  ↓
Multilingual Dense Retrieval
  ↓
Hybrid Ranking
Enter fullscreen mode Exit fullscreen mode

But the bigger lesson for me was that search quality is less about choosing one representation and more about understanding why a query failed in the first place.

If the user uses the exact government terminology, lexical retrieval can work.

If they use different terminology, learned sparse retrieval can provide another signal.

If they switch languages or scripts, multilingual dense retrieval becomes more relevant.

And when we don’t know which of these situations we’re going to get, combining the signals gives us a retrieval system that is less dependent on one particular way of asking a question.

For e-governance, that’s important.

The citizen shouldn’t have to know whether the official document calls something:

income support
Enter fullscreen mode Exit fullscreen mode

or:

financial assistance
Enter fullscreen mode Exit fullscreen mode

or:

cash transfer
Enter fullscreen mode Exit fullscreen mode

They should be able to describe what they need in the language they naturally use.

The retrieval system should handle the translation between how people ask and how information is stored.

That’s the problem I wanted to explore with this project.


References

  1. Dataset — Indian Government MyScheme Dataset
    Aryan-Pardeshi/gov-myscheme-dataset — 2,066 structured government schemes derived from MyScheme.

  2. Government of India — MyScheme
    myscheme.gov.in

  3. Qdrant — Hybrid and Multi-Stage Queries
    Qdrant Hybrid Queries Documentation — multiple named vectors, prefetch and RRF.

  4. Qdrant — Hybrid Search
    Combining Semantic and Lexical Search

  5. Qdrant — FastEmbed
    FastEmbed Documentation

  6. BAAI — BGE-M3
    BAAI/bge-m3 on Hugging Face

  7. SPLADE PP en v1
    prithivida/Splade_PP_en_v1

  8. AI4Bharat — IndicQA
    IndicQA Dataset

  9. Project Repository — E-Governance
    github.com/AKR4PC/E-Governance

Top comments (0)