Your users learned to ask questions somewhere else, and now your agents are asking too. Your search bar is the last place that still wants keywords.
A citizen types "can I renew my passport if it expired two years ago" into a government services portal, and the search bar hands back every page that contains the words "passport," "renew," and "expired," ranked by how often those words appear. The one page that answers the question sits on the third results screen, if it surfaces at all. The citizen gives up and calls the contact center, which costs the agency real money on a query the website should have answered for free.
If you run search for an organization, you have watched some version of this play out. Your users arrive from ChatGPT, Perplexity, and Claude, where they ask full questions and get direct answers. Then they reach your search box and have to translate the question back into the two or three keywords they hope your index contains. The friction is not subtle, and it shows up in your analytics as abandoned searches, repeated queries, and support tickets for answers that already exist in your content.
The reason to move to vector search is not that the technology is new or interesting. The reason is that the cost of staying on keyword-only search now shows up on a budget line, and it is about to show up again as you put agents in front of your content. An agent answering a user's question runs its own searches, and it hits the same wall your users do when the index only matches exact terms. This guide covers when keyword search is actually failing you, what vector and hybrid search on Amazon OpenSearch Service change for both people and agents, and how to stage the migration without ripping out the system you already run.
When keyword search is actually failing you
Keyword search still does real work, and moving to Amazon OpenSearch Service does not mean giving it up. Exact product codes, known document titles, specific error strings, faceted navigation, and structured filters all run on BM25 ranking that returns the right result quickly and transparently, and OpenSearch Service runs that lexical path natively. The point is not to replace keyword search. It is to stop asking keyword search to answer the queries it was never built for.
The failure mode appears when queries stop looking like keywords and start looking like questions. A search for "vehicle hire" against a catalog that only says "car rental" returns nothing, because the two strings share no terms even though a person would treat them as the same request. Keyword search matches the characters the user typed against the characters in your index. When those two vocabularies diverge, the query fails silently, and the user rarely tells you. They just leave.
Pull your search logs and sort by zero-result and low-engagement queries, and the failures line up in a clear pattern. They cluster on natural-language phrasings, multi-part questions, and searches that use different words than your content authors used. Those are the queries bleeding engagement, and they are exactly the queries that keyword matching cannot rescue no matter how carefully you tune the ranking.
What vector and hybrid search change
Vector search matches on the company words keep rather than the exact characters a user typed. An embedding model maps text to coordinates so that phrases appearing in similar contexts across the training corpus land near each other. "Vehicle hire" and "car rental" end up close because they show up in overlapping surrounding text, not because a model comprehends either phrase. The practical effect is that a query retrieves relevant documents even when the author and the user chose different words for the same thing. Correlation over word usage, not understanding, is what closes the vocabulary gap.
Vector search alone overcorrects. Ask it for an exact product code and it may return a cluster of conceptually adjacent items while burying the precise match the user actually wanted. This is why production search combines both approaches. Hybrid search runs the keyword path and the vector path together and merges the results, so exact matches keep ranking high while related-but-differently-worded documents fill in the results a keyword-only system would have missed. OpenSearch Service runs hybrid search natively, and the balance between the two paths is a setting you tune to your traffic rather than a rewrite you commit to.
Increasingly this is not one-shot retrieval-augmented generation but agent-driven question answering. An agent decides how to answer a question: it can run several retrievals, refine them based on what came back, pull from more than one index, and compose an answer grounded in what it found. OpenSearch Service supports this directly with agentic search, where the agent chooses the lexical or vector path per query based on the query itself, using exact matching for a date filter or a product code and vector retrieval for a conceptual question. Grounding the agent in content it retrieves from your own indices is what keeps the answer built from your documents rather than from the model's training data alone. Retrieval quality sets the ceiling on answer quality, so the move to better search pays off twice: once in direct results and again in every agent you build on top of them.
How to stage the migration
Treat the move as an augmentation, not a rip-and-replace. Your keyword index keeps running and keeps serving the exact-match and filter traffic it already handles well. You add a vector path alongside it and route more query types through the hybrid pipeline as your confidence grows. Nothing about this requires a flag-day cutover, and the staged approach means you can measure each step against the system it replaces.
Start where keyword search fails hardest, which your zero-result logs already identified. Stand up an embedding pipeline in OpenSearch Service so documents get vectorized as they are indexed, point a hybrid query at the failing query set, and compare relevance against your current system on the same queries. The Search Relevance Workbench in OpenSearch Dashboards runs keyword and neural queries side by side against the same index, so you can see the difference on your own queries rather than take it on faith, and it is enough of a topic on its own that it deserves its own article. Amazon OpenSearch Service can host the embedding model inside the engine and vectorize both documents at index time and queries at search time, so you are not standing up a separate model-serving tier to get started. Once the hybrid pipeline wins on the queries that were failing, widen its coverage.
Measure the transition on user behavior, not just offline relevance scores. The catch is that most teams do not capture user behavior in any structured way, so the signal that would prove the migration worked never gets recorded. This is what User Behavior Insights (UBI) addresses: a standard schema that ties each query to the results shown and the actions the user took on them, so zero-result rate, search refinement, and click-through depth become data you can query rather than anecdotes. You can collect UBI-formatted data with Amazon OpenSearch Service and analyze it in the same engine that serves the search. Those metrics also make the business case for the next phase, because they translate search quality into the operational costs your organization already tracks, like the support contacts that duplicate content the site should have surfaced.
The cost of waiting
The gap between what your users expect and what your search bar delivers widens every day they spend in a conversational interface somewhere else. When one organization in your sector ships natural-language search, the expectation resets for everyone your users interact with next, including you. The migration stops being a project you might schedule and becomes a gap your users feel on every visit.
Amazon OpenSearch Service gives you a path that does not start with tearing down what you have. Keep the keyword search that works, add vector and hybrid search where questions have replaced keywords, and let agents run retrieval across your indices when you are ready to return answers instead of links. The technology is production-ready today. The queries your users and your agents are already typing are the ones telling you where to start.
Top comments (0)