Originally published at satyam-mishra-two.vercel.app.
Shoppers don’t search the way product catalogs are written. They type “back pain”, “60 inch tv” or “something for standing all day”, and most store search boxes answer with nothing. Typesense, an open-source search engine, lets a small team build search that understands what people mean, fast, and on their own terms. This is why it matters and how we put it to work, guardrails included.
Search is where buying intent is loudest
Someone who types into your search box has already told you what they want. That makes search the most valuable box on the page, and the one most stores get wrong. Baymard Institute benchmarks ecommerce search every year across more than 170 sites and apps. Its 2026 numbers say 56% of sites fail to adequately support what users search for, and roughly half of shoppers prefer search over the menu to find products.
The failures are not random. Baymard splits searches into eight query types, and the gap shows up exactly where shoppers stop typing product names and start describing their problem.
Searching by name mostly works. Searching by need, symptom or use mostly doesn’t. Source:Baymard Institute, Ecommerce Search UX (2026)
A keyword engine can only match words that are in your product data. “Back pain” is not in the title of an ergonomic chair. “Footwear” is not in the title of an insole, and that is a good thing, because a shopper looking for shoes does not want insoles. Good search has to understand both what a query means and what it doesn’t.
What Typesense brings to the table
Typesense is a search engine written in C++ that keeps its index in memory. Its own description is a “fast, typo tolerant, in-memory fuzzy search engine”, and its features list targets searches under 50 milliseconds. For an online store, a handful of features carry most of the weight:
- Typo tolerance on by default. Up to two typos per word, with prefix matching so results update as the shopper types. “Insloe” still finds insoles.
- Synonyms. Tell the engine that two words mean the same thing in your shop, including local words your customers use and your catalog doesn’t.
- Curation. Pin chosen products to the top for a specific query. The query that brings the most traffic deserves a hand-picked answer.
-
Built-in vector search. Typesense can create embeddings itself with a bundled model such as
all-MiniLM-L12-v2, both when it indexes products and when a query arrives. No separate machine-learning service to run. - Hybrid search. One request runs a keyword search and a semantic search together and fuses the two rankings. You choose how much weight the semantic side gets, and a distance threshold that drops weak semantic matches.
- Natural-language search. Since version 29, an LLM can turn “a Honda or BMW with at least 200 hp” into real filters and sorts. More on that, and its risks, below.
Hybrid search: matching words and matching meaning
Keyword search is precise and fast. Semantic search understands meaning, but on its own it drifts: everything is a little bit similar to everything. Hybrid search uses each for what it is good at. Researchers describe the same idea in a 2024 paper on hybrid semantic search: combining keyword matching with semantic embeddings captures both the explicit and the implicit intent behind a query.
In practice that means the shopper who types a product name gets the exact product first, because the keyword side wins. The shopper who types a symptom or a use case still gets relevant products, because the semantic side finds them when no keyword matches. The tuning comes down to two decisions:
- How much to trust meaning over words. Typesense’s default fusion gives 70% weight to keyword rank and 30% to semantic rank. For a store where people often search by product name, keeping keywords in charge is usually right.
- How far is too far. A distance threshold cuts semantic matches that are only loosely related. Set it too loose and a search for “chair” starts showing cushions. Set it too tight and “back pain” finds nothing.
Why open source matters for a store
Typesense is licensed under GPL-3.0 and is free to self-host as a single binary. That changes the conversation in three ways.
| Question | Open source (Typesense) | Closed SaaS search |
|---|---|---|
| What do you pay for? | Your own server, or Typesense Cloud billed hourly by cluster size, with no per-search or per-record charges | Often per search request and per record, so cost grows with traffic and catalog size |
| A new sort order (price, rating)? | Sort fields are chosen at query time on one index | Some providers need a duplicate index per sort order, which adds billable records |
| Can you leave? | Yes. Same engine on your server or theirs, and you can read the code | Your ranking logic lives in someone else’s dashboard |
| Where does the data live? | Where you choose, which helps with EU data rules | Where the provider runs it |
For a growing store, the first row is the big one. A sale weekend multiplies searches. With per-request pricing, your best days cost the most. With a fixed-size cluster, they cost the same.
How we put it to work
At Frido we rebuilt product search on Typesense. I won’t share the internals, but the shape is one any store can copy, and the lessons were the useful part.
A generic shape for open-source product search. The catalog stays the source of truth; the search index is a copy you can rebuild at any time. Diagram made with Archify.
- The index is a copy, never the truth. Products live in the store platform. A sync job pushes changes into Typesense as they happen, and a full rebuild takes seconds. If the index is ever wrong or empty, we rebuild it instead of fixing it by hand.
- Each result carries its own card. We store what the product card needs next to the searchable fields, so a hit renders straight away with no second lookup.
- Typeahead stays keyword-only. Suggestions while typing need to be instant. In our measurements a keyword lookup returns in a few milliseconds, while embedding a new query adds roughly a tenth of a second. So the dropdown skips the semantic step and the full results page uses hybrid search.
- Zero results get a second chance. If the hybrid search finds nothing, we retry with semantic search alone and a looser threshold. Shoppers almost never hit a dead end, and the extra cost is only paid on misses.
What the data taught us
- Keep marketing copy out of keyword fields. Long benefit bullets mention every related word, so they make products match searches they have nothing to do with. We keep keyword matching on titles, categories and tags, and let the embedding read the descriptive text.
- Feed the embedding less, not more. Adding noisy fields pulled one product’s meaning away from what it actually was. Removing them made it find its own category again.
- Use synonyms only where meaning fails. Embeddings are not perfect. In ours, “table” sat closer to chairs than to desks. A one-way synonym fixed it. The rule we follow: only true equivalents, never two different categories.
- Your copy limits your intent search. Semantic search can only find “for standing all day” if somewhere in the product text it says what the product is for. Better product descriptions improve search more than any setting.
Guardrails: fast is good, reliable is better
A search box is a public input on every page. It gets typos, junk, bots and traffic spikes. Typesense is fast, but a good setup assumes that anything can fail and plans for it.
Each step either answers the query or hands it on safely. Diagram made with Archify.
- Never put an admin key in the browser. Typesense supports search-only keys limited to one action on one collection, and scoped keys with filters baked in that users “will not be able to override”. The browser only ever sees that.
- Debounce and cancel. Wait until the shopper pauses typing and cancel the previous request. It cuts load and stops old results flashing over new ones.
- Set limits. A minimum query length (show popular categories instead of searching for one letter), a ceiling on results per page and a rate limit per visitor on public endpoints.
- Normalise before you cache. Lowercase and trim the query, and cap the length of the cache key, so “Insoles” and “insoles ” share one entry and junk queries can’t fill the cache. Typesense also has its own result cache, which is off by default; switch it on.
- Use a hard timeout. A failed request throws an error you can catch. A hung request just waits. Without a deadline, your fallback never runs.
- Fall back to plain keyword search. If the engine is down, slow or rebuilding, use the store platform’s own search or a simple database match, and render the same product cards. Worse results beat a blank page.
- Keep merchandising rules few and visible. A handful of pinned results is easy to reason about. Hundreds of hidden rules become a second, undocumented ranking system.
If you add LLM natural-language search, treat the model’s output as untrusted input. Typesense’s own guide warns that LLMs can misread a query or produce invalid syntax. Check the generated filters against the fields you actually allow, cap the time it can take, and fall back to a normal search when the output does not validate. We have not needed it yet: hybrid search already handles most intent queries without a model call on every search.
A checklist to start
- Read your own search logs. Find the top queries and the ones that return nothing.
- Index only what helps matching: title, category, tags, price, stock. Keep long copy for the embedding.
- Turn on hybrid search with keywords in charge, then tune the distance threshold with real zero-result queries.
- Add synonyms for the gaps the embedding misses. Pin the answer to your top two or three queries.
- Add the guardrails: search-only key, debounce, limits, cache, timeout, fallback.
- Measure search speed with the rest of the page. A fast search box on a slow page still feels slow; see why your Shopify store is slow.
FAQ
Is Typesense free?
The engine is open source under GPL-3.0 and free to self-host. Typesense Cloud is the paid hosted option, billed by cluster size per hour rather than per search.
What is hybrid search in ecommerce?
It runs a keyword search and a semantic (meaning-based) search on the same query and merges the rankings. Exact product names still come first, and descriptive searches like “back pain” still find relevant products.
Do I need an LLM to understand search intent?
Not for most stores. Typesense can create embeddings itself with a small bundled model, and hybrid search covers most intent queries. An LLM helps when shoppers type filters in plain language, like a price range or brand, and it needs validation before its output reaches the engine.
Can Typesense work with Shopify?
Yes. Products stay in Shopify, and a sync process copies them into Typesense whenever they change. The storefront queries Typesense and can fall back to Shopify’s own search if Typesense is unavailable.
How is Typesense different from Algolia?
Both are fast, in-memory engines. Typesense is open source and can run on your own servers; Algolia is closed-source SaaS. Their pricing models differ too: per cluster for Typesense Cloud, by requests and records on Algolia’s usage-based plans.
Sources
- Ecommerce Search UX: The 8 Query Types, Baymard Institute (updated 2026)
- Typesense source code and feature list, GitHub (2026)
- Typesense vs Algolia vs Elasticsearch vs Meilisearch, Typesense (2026)
- Vector Search and Hybrid Search, Typesense docs v30.2 (2026)
- Search parameters (typo tolerance, synonyms, curation, caching), Typesense docs v30.2 (2026)
- API Keys and scoped search keys, Typesense docs v30.2 (2026)
- Natural Language Search guide, Typesense
- Hybrid Semantic Search: Unveiling User Intent Beyond Keywords, Ahluwalia et al., arXiv (2024)
I'm Satyam Mishra, a software developer in Pune working with stores and startups in Ireland, Europe and the US. More guides at satyam-mishra-two.vercel.app/guides.



Top comments (1)
tr.ee/dev-to