DEV Community

I Built an AI Saree Search Engine for My Uncle's Local Business

Hacktoberfest Weekend Challenge: Build for a Friend Submission 🀝

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

What I Built

My uncle runs a local saree business, and I wanted to build something that could actually help him rather than building another generic AI chatbot.

So I built Saree Saathi AI β€” an AI-powered natural-language saree search and recommendation system.

The idea is simple:

Instead of asking customers to understand product categories and manually select filters like:

Fabric: Silk
Occasion: Wedding
Price: < β‚Ή7000
Enter fullscreen mode Exit fullscreen mode

they can simply say:

"I need a traditional silk saree for a wedding under β‚Ή7000."

The system understands the request, searches the catalog semantically, applies exact business constraints, and generates a recommendation based only on the sarees that were actually retrieved.

The problem

A small local saree business can have a large and diverse catalog:

  • Mangalagiri
  • Uppada
  • Pochampally
  • Kanchipuram
  • Dharmavaram
  • Banarasi
  • Paithani
  • Cotton
  • Silk
  • Silk Cotton
  • Different colors, occasions and price ranges

Traditional search doesn't capture the way people actually describe what they want.

Someone might say:

"I want something elegant for my sister's wedding, preferably silk, but I don't want to spend more than β‚Ή7000."

That contains intent, semantics and hard constraints mixed together.

That's the problem I wanted to solve.


Demo

The project currently runs locally because the AI stack uses Ollama for local Gemma inference.

Example

Request:

curl -X POST http://localhost:3000/api/search \
  -H "Content-Type: application/json" \
  -d '{"query":"Show me a cotton saree for office under β‚Ή3000"}'
Enter fullscreen mode Exit fullscreen mode

The system understands:

{
  "semanticQuery": "cotton saree for office",
  "maxPrice": 3000
}
Enter fullscreen mode Exit fullscreen mode

It also extracts the deterministic constraint:

{
  "fabric": "Cotton"
}
Enter fullscreen mode Exit fullscreen mode

Then MongoDB Atlas Vector Search finds relevant products while enforcing the price and availability constraints.

Example results included:

Mangalagiri Cotton Saree       β‚Ή1800
Mangalagiri Cotton Saree       β‚Ή2200
Pochampally Ikat Cotton Saree  β‚Ή2400
Enter fullscreen mode Exit fullscreen mode

And the final recommendation is generated from those retrieved products.

The complete local setup and API instructions are available in the repository README.


Code

The project is open source:

GitHub: https://github.com/bhaskar359/saree-saathi-ai

The repository contains the complete backend implementation, sample catalog, embedding pipeline, MongoDB integration, AI orchestration and local setup instructions.

The main architecture is:

src/
β”œβ”€β”€ ai-search.js
β”œβ”€β”€ embedding.js
β”œβ”€β”€ query-understanding.js
β”œβ”€β”€ query-constraints.js
β”œβ”€β”€ server.js
β”‚
β”œβ”€β”€ controllers/
β”‚   └── search-controller.js
β”‚
β”œβ”€β”€ errors/
β”‚   └── AppError.js
β”‚
β”œβ”€β”€ middleware/
β”‚   β”œβ”€β”€ validate-search.js
β”‚   └── error-handler.js
β”‚
β”œβ”€β”€ routes/
β”‚   └── search.js
β”‚
β”œβ”€β”€ services/
β”‚   β”œβ”€β”€ saree-search.js
β”‚   └── recommendation.js
β”‚
└── utils/
    β”œβ”€β”€ format-search-result.js
    └── with-timeout.js
Enter fullscreen mode Exit fullscreen mode

How I Built It

The interesting part of this project isn't just "I used an LLM."

I wanted the system to demonstrate how several AI techniques can work together.

1. Natural-language understanding

I use Gemma 3 1B through Ollama to understand the user's request.

For example:

"I need a traditional silk saree for a wedding under β‚Ή7000"
Enter fullscreen mode Exit fullscreen mode

becomes:

{
  "semanticQuery": "traditional silk saree for a wedding",
  "maxPrice": 7000
}
Enter fullscreen mode Exit fullscreen mode

The LLM is used for language understanding, not for constructing arbitrary database queries.


2. Deterministic constraints

This was one of the most important architectural decisions.

I initially experimented with allowing the LLM to extract things such as:

fabric
state
price
Enter fullscreen mode Exit fullscreen mode

But LLMs can hallucinate structured values.

For example, if a user says:

"I want something beautiful from Andhra Pradesh"

we don't want the model randomly deciding:

{
  "fabric": "Silk"
}
Enter fullscreen mode Exit fullscreen mode

So the architecture became:

             User Query
                 β”‚
        β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”
        β–Ό                 β–Ό
      Gemma          Application Code
        β”‚                 β”‚
        β–Ό                 β–Ό
 Semantic meaning    Hard constraints
Enter fullscreen mode Exit fullscreen mode

In other words:

LLM for meaning. Application code for rules.

This makes the search pipeline considerably more predictable.


3. Semantic embeddings

For semantic search I use:

Xenova/all-MiniLM-L6-v2

The model converts text into 384-dimensional embeddings.

For example:

Traditional silk saree suitable for a wedding
Enter fullscreen mode Exit fullscreen mode

and:

Elegant silk saree for a marriage ceremony
Enter fullscreen mode Exit fullscreen mode

have very similar semantic representations even though they don't use exactly the same words.

This lets the system search by meaning rather than keyword matching.


4. MongoDB Atlas Vector Search

The embeddings are stored alongside the saree catalog in MongoDB Atlas.

Conceptually:

{
  "id": "SAR003",
  "name": "Uppada Jamdani Silk Saree",
  "state": "Andhra Pradesh",
  "fabric": "Silk",
  "price": 6500,
  "available": true,
  "embedding": [384 dimensions...]
}
Enter fullscreen mode Exit fullscreen mode

The MongoDB vector index uses:

384 dimensions
cosine similarity
Enter fullscreen mode Exit fullscreen mode

But this is where another important lesson appeared.


5. Why vector search alone wasn't enough

Suppose someone asks:

"Silk saree under β‚Ή7000."

Vector search understands:

silk
saree
Enter fullscreen mode Exit fullscreen mode

and the semantic intent.

But vector similarity isn't the right mechanism for enforcing:

price <= 7000
Enter fullscreen mode Exit fullscreen mode

A β‚Ή15,000 saree could still be semantically very similar to the query.

So I combined vector search with structured MongoDB filters.

             Query
               β”‚
       β”Œβ”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”
       β–Ό                β–Ό
Semantic Search     Hard Filters
       β”‚                β”‚
       β”‚        price <= β‚Ή7000
       β”‚        fabric = Silk
       β”‚        available = true
       β”‚
       β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”˜
                β–Ό
          Hybrid Search
                β”‚
                β–Ό
        Relevant Sarees
Enter fullscreen mode Exit fullscreen mode

This became one of the core architectural principles of Saree Saathi AI.


6. Grounded RAG recommendations

After retrieving the relevant sarees, I use Gemma again to generate the final recommendation.

But the model receives the retrieved catalog information as context.

It is explicitly instructed:

  • Don't invent prices.
  • Don't invent fabrics.
  • Don't invent colors.
  • Don't invent availability.
  • Don't invent reviews.
  • Don't recommend products that weren't retrieved.
  • Use the catalog as the source of truth.

So the flow is:

User Query
    +
Retrieved Catalog
    ↓
   Gemma
    ↓
Grounded Recommendation
Enter fullscreen mode Exit fullscreen mode

For example:

I recommend the Mangalagiri Cotton Saree (SAR001).

It's priced at β‚Ή1800 and is suitable for office wear.
Enter fullscreen mode Exit fullscreen mode

The recommendation is grounded in the actual catalog record.


7. REST API

I wrapped the AI pipeline behind an Express API:

POST /api/search
Enter fullscreen mode Exit fullscreen mode

and:

GET /health
Enter fullscreen mode Exit fullscreen mode

The request validation layer handles things such as:

missing query
wrong data type
empty query
excessively long query
Enter fullscreen mode Exit fullscreen mode

I also added:

  • response formatting
  • centralized error handling
  • AI timeout protection
  • environment-based configuration

The goal was to keep the AI logic from becoming one giant API handler.


The Architecture

The complete pipeline looks like this:

                     User
                      β”‚
                      β–Ό
              Natural Language
                      β”‚
                      β–Ό
               Express API
                      β”‚
                      β–Ό
              Request Validation
                      β”‚
          β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
          β”‚                       β”‚
          β–Ό                       β–Ό
       Gemma                Deterministic
          β”‚                   Constraints
          β–Ό                       β”‚
    Semantic Query                β”‚
          β”‚                       β”‚
          β–Ό                       β”‚
     MiniLM Embedding             β”‚
          β”‚                       β”‚
          β–Ό                       β”‚
    Query Vector                  β”‚
          β”‚                       β”‚
          β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                      β–Ό
             MongoDB Atlas
             Vector Search
                      β”‚
                      β–Ό
              Relevant Sarees
                      β”‚
                      β–Ό
                  RAG + Gemma
                      β”‚
                      β–Ό
             Grounded Recommendation
                      β”‚
                      β–Ό
                 API Response
Enter fullscreen mode Exit fullscreen mode

Why Does Open Innovation Matter?

This project is especially interesting to me because I didn't want to build it around a black-box AI API and stop there.

Using open and locally runnable AI components allowed me to experiment with the entire pipeline.

Local inference

With Ollama and Gemma, I can run the LLM locally.

That means I can experiment without sending every development query to a hosted API.

Open embedding model

The embedding pipeline uses:

Xenova/all-MiniLM-L6-v2
Enter fullscreen mode Exit fullscreen mode

This gives me direct control over the embedding process instead of treating embeddings as an invisible API operation.

Replaceable architecture

The project is intentionally designed so that the AI components can evolve.

For example:

Current:
Gemma + MiniLM
       ↓
Future:
Another open-weight LLM
       +
Another embedding model
Enter fullscreen mode Exit fullscreen mode

The application architecture doesn't need to be rewritten.

That's one of the biggest things I learned from building this:

Open AI isn't just about having access to a model. It's about being able to understand, experiment with, replace and control the pieces of the system.


My Agent Session

I decided not to include a DevRelay Agent Session for this submission.

The project was developed iteratively with ChatGPT, including the architecture decisions, model experiments, debugging, implementation and refinement.

Rather than presenting a different tool as the development agent, I'm keeping that part transparent.


Prize Categories

πŸƒ MongoDB Atlas

Saree Saathi AI uses MongoDB Atlas Vector Search as a core part of its architecture.

MongoDB stores:

  • Saree catalog data
  • Embeddings
  • Structured metadata

and performs the hybrid retrieval using vector similarity plus metadata filters.

The project would therefore not just be using MongoDB as a conventional databaseβ€”the vector search capability is fundamental to the AI search pipeline.


What I Learned

This project taught me that building an AI application isn't simply:

User β†’ LLM β†’ Answer
Enter fullscreen mode Exit fullscreen mode

A useful production-oriented AI system often looks more like:

User
 ↓
Validation
 ↓
Query Understanding
 ↓
Structured Constraints
 ↓
Embeddings
 ↓
Vector Retrieval
 ↓
Metadata Filtering
 ↓
Grounded Context
 ↓
LLM
 ↓
Validated Response
Enter fullscreen mode Exit fullscreen mode

The most valuable lesson was knowing where not to use an LLM.

Prices, availability, database operators and business rules don't need creativity.

They need determinism.


What's Next?

The current version is intentionally focused on the backend AI/search pipeline.

Future improvements include:

  • Automated Jest test suite
  • API documentation
  • Rate limiting
  • Structured logging
  • Search latency metrics
  • Multilingual/Telugu search
  • Image-based saree search
  • Saree image embeddings
  • Better retrieval evaluation
  • Re-ranking
  • Conversation-aware search
  • Admin catalog management
  • Frontend experience

Eventually, I'd like to turn this into something my uncle could actually use with his real inventory.


Try It Yourself

The project is designed to run locally with:

Node.js
MongoDB Atlas
Ollama
Gemma 3 1B
all-MiniLM-L6-v2
Enter fullscreen mode Exit fullscreen mode

The complete setup instructions, MongoDB Vector Search configuration, environment variables and API examples are available in the repository README.

GitHub: https://github.com/bhaskar359/saree-saathi-ai


Final Thoughts

I started this project with a simple question:

Can I build something useful for someone close to me instead of building another tutorial project?

That question turned into an exploration of:

  • embeddings
  • vector databases
  • semantic search
  • hybrid retrieval
  • LLM structured output
  • deterministic constraints
  • RAG
  • local inference
  • API architecture

And that became Saree Saathi AI.

Built for my uncle. Built to learn. Built to be useful. πŸ₯»πŸ€–


#Hacktoberfest #BuildForAFriend #AI #MongoDB #OpenSource #JavaScript

Top comments (0)