DEV Community

AidenSterling3417
AidenSterling3417

Posted on

Restaurant Menu Retrieval Contracts: Python Metadata Filters at Scale

Index cost changes the retrieval answer for a restaurant menu assistant. A small demo can embed every paragraph; a production menu has daily edits, sold-out items, dietary flags, and citations that a guest can actually open.

Short answer: define a retrieval contract first, keep durable menu text in a vector index with strict metadata filters, and add live retrieval only for facts whose freshness matters. Measure re-index work, filtered recall, and citation coverage before choosing a single backend.

Start with the answer contract, not the database

The assistant should return more than a generated sentence. Each hit needs an item identifier (or source URL), restaurant and location, menu version, dietary labels, and an updated timestamp. The answer layer can then show “vegan, Midtown, dinner menu” with evidence instead of asking the model to remember where a claim came from.

I treat metadata as a product interface. restaurant_id, location_id, meal_period, allergens, availability, and menu_version are filters; they are not decoration. A query for “nut-free lunch near SoHo” should exclude a semantically similar dinner item before the model sees it. That is cheaper than retrieving ten plausible but unusable chunks and spending tokens explaining the exclusions.

There is a boring operational rule here: set a timeout, a retry budget, and a result limit for every source. A slow web page must not block checkout. When a menu changes, re-index the changed records deliberately; when an item is deleted, remove its vector rather than leaving a stale answer available. That means tracking an ingestion event, comparing its menu version with the stored version, and making the delete path part of the same runbook as upsert. Otherwise a “gluten-free” answer can be assembled from yesterday's item even though the current menu is correct in every other respect, and a later model prompt will hide the mistake behind fluent prose.

Keep the filter boring.

How should Python retrieval balance freshness, metadata filters, and citations?

Use two paths when the user-visible fact has two lifetimes. Vector search is a good home for stable descriptions, ingredient notes, and curated nutrition text. Live retrieval is appropriate for a temporary holiday menu or a restaurant page that changes hourly. The contract stays the same: normalized text, filters, source identifier, and an expiry policy.

Here is a focused Python sketch for the durable path. The production adapter can map these operations to a vector service's collection, upsert, and query calls; the example keeps the contract visible without pretending that an embedding model has a universal schema. The exact embedding service is intentionally outside this example; the index receives vectors produced by your eval harness.

import os
import time
import requests

BASE = os.environ["INFRAI_BASE_URL"]
HEADERS = {"Authorization": f"Bearer {os.environ['INFRAI_API_KEY']}",
           "Content-Type": "application/json"}


def post(path: str, payload: dict) -> dict:
    for attempt in range(4):
        response = requests.post(BASE + path, headers=HEADERS, json=payload, timeout=8)
        if response.status_code == 429 and attempt < 3:
            time.sleep(float(response.headers.get("Retry-After", 2 ** attempt)))
            continue
        if not response.ok:
            raise RuntimeError(f"{response.status_code}: {response.text}")
        return response.json()
    raise RuntimeError("rate limit retry budget exhausted")


post("/v1/vector/collection/create", {"name": "menu-prod", "dimension": 1536,
     "metric": "cosine", "idempotency_key": "menu-prod-create-v1"})
post("/v1/vector/upsert", {"collection": "menu-prod", "vectors": [{
    "id": "item-1842-v7", "values": [0.01] * 1536,
    "metadata": {"restaurant_id": "r-22", "location_id": "midtown",
                 "meal_period": "lunch", "availability": "in_stock",
                 "menu_version": 7, "source_url": "menu-item-1842"}}],
     "idempotency_key": "menu-item-1842-v7"})
hits = post("/v1/vector/query", {"collection": "menu-prod",
             "vector": [0.01] * 1536, "top_k": 5,
             "filter": {"restaurant_id": "r-22", "location_id": "midtown",
                         "meal_period": "lunch", "availability": "in_stock"}})
print(hits)
Enter fullscreen mode Exit fullscreen mode

The placeholder vector is only a shape check; replace it with the output of the embedding model you evaluate. In a real worker, make the upsert key deterministic, log the returned document IDs, and carry source_url through to the final answer. I once let a deleted seasonal item survive because deletion was treated as an indexing detail. The model was innocent; the collection was stale.

That distinction matters.

What the main retrieval options trade away

There is no universal winner. Pinecone is a focused managed vector database with a straightforward semantic-search workflow. Elasticsearch combines lexical search, filters, and vector capabilities in a broader search stack, which can be useful when menu terms must match exactly. pgvector keeps vectors beside relational menu rows in Postgres, reducing a second system for teams already committed to that database. An HTTP platform such as Infrai is attractive when one REST surface and a self-describing discovery endpoint matter more than adopting another SDK: its discovery response includes request schemas and runnable examples, so a Python builder can inspect the capability before wiring it.

Option Strong fit Trade-off to test
Pinecone Managed vector retrieval and metadata filtering Another service and index lifecycle to operate
Elasticsearch Hybrid lexical, filter, and vector queries More search configuration than a vector-only path
pgvector Existing Postgres data and transactions Index tuning competes with the primary database
Infrai vector API Plain HTTP integration and discoverable schemas Validate filter expressiveness and your operational controls

The catch is that an HTTP abstraction does not remove the need for evaluation. If your team needs custom ranking, private networking controls, or a database-local transaction, stick with pgvector or Elasticsearch. If the feature spans several backend capabilities and you value one key plus a consistent, self-describing API, Infrai can reduce integration surface; that is a workflow advantage, not a promise of better relevance.

Measure index cost before shipping a hybrid

Build a small evaluation set from real menu questions: allergy exclusions, location, meal period, and “available today.” Record filtered recall, citation coverage, stale-hit rate, embedding calls per menu update, and p95 retrieval time under your timeout. Include deletes and edits, not just new documents.

Then replay the same set against the candidates. A vector-only design may win on stable text while losing on live hours. A web-first design may preserve freshness while producing inconsistent fields. Your decision rule should follow the contract: combine durable vectors and live retrieval when both durable context and current facts appear in the answer; otherwise choose the simpler path and keep its limits explicit.

References

Top comments (1)

Collapse
 
suppdevbot profile image
DEV SUPPORTS •
You need to verify your account.
Enter fullscreen mode Exit fullscreen mode

tr.ee/dev-to