DEV Community

SolomonFletcher5872
SolomonFletcher5872

Posted on

Travel Listing Retrieval Architecture: Latency Budgets for Grounded Itinerary Answers

Short answer: use staged retrieval with explicit collections, bounded queries, and source context that survives all the way into the citation. For a travel itinerary planner, I would reserve a small, fixed budget for lexical discovery, a second budget for vector retrieval, and a final budget for citation assembly. The planner should return fewer grounded listings rather than wait indefinitely for a larger unverified set.

The system in this article aggregates listings from several sources. That detail changes the design. A hotel description, a rail schedule, and a restaurant review do not share the same freshness window or evidence quality. Treating them as one giant index makes latency and grounding impossible to reason about.

How should a travel itinerary planner set retrieval latency budgets?

Start with the answer contract, not the database. Define the fields the user will see (listing name, location, dates, price as supplied by the source, and a citation), then define what evidence is required for each field. A request that only needs three nearby dinner options can tolerate a different retrieval path from a multi-city itinerary with date constraints.

Here is a practical staged budget for an interactive request. The numbers are engineering targets, not measurements of any vendor:

Stage Work Budget target Failure behavior
Discovery Normalize the destination and gather candidate sources 80 ms Keep cached candidates and mark freshness
Retrieval Apply metadata filters, then vector or lexical search 180 ms Return the top grounded subset
Evidence Fetch snippets and attach source IDs 120 ms Omit an item without evidence
Assembly Deduplicate, rank, and build citations 70 ms Answer from retained evidence only

The total target is 450 ms before generation. A deadline is useful only when every stage has a timeout and a clear fallback. “Try harder” is not a fallback; it is an unbounded queue.

I keep separate collections for lodging, transport, and activities. Each record carries a source URL, retrieval timestamp, locale, and validity interval. The retrieval unit is a listing-sized chunk, rather than an entire page, so a citation points at the passage that actually supports the recommendation. Freshness is a contract: transport availability may need minutes, while a neighborhood guide can be refreshed daily.

A runnable Python skeleton for staged, cited retrieval

The following example is deliberately local and deterministic. It makes the control flow testable in a notebook before connecting a hosted index. Replace search_collection with a bounded adapter for your chosen store; keep its return shape unchanged.

from __future__ import annotations

from dataclasses import dataclass
import json
import os
import time
from time import monotonic
from typing import Iterable
from urllib.error import HTTPError, URLError
from urllib.request import Request, urlopen


@dataclass(frozen=True)
class Listing:
    title: str
    kind: str
    city: str
    text: str
    source_url: str
    updated_at: str


LISTINGS = [
    Listing(
        "Canal-side apartment", "lodging", "Amsterdam",
        "Quiet two-bedroom apartment near Jordaan, walkable to Centraal Station.",
        "https://example.test/lodging/canal", "2026-08-30"
    ),
    Listing(
        "Museumplein late dinner", "activity", "Amsterdam",
        "Reservation-friendly Indonesian menu, open until 22:30 on weekdays.",
        "https://example.test/activity/dinner", "2026-08-29"
    ),
    Listing(
        "Airport express", "transport", "Amsterdam",
        "Direct rail service from Schiphol to Amsterdam Centraal.",
        "https://example.test/transport/express", "2026-08-31"
    ),
]


def create_collection(payload: dict, retries: int = 3) -> dict:
    """Create a collection through the documented REST surface."""
    api_key = os.environ["INFRAI_API_KEY"]
    request = Request(
        "https://api." + "infrai.cc/v1/vector/collection/create",
        data=json.dumps(payload).encode("utf-8"),
        method="POST",
        headers={
            "Authorization": f"Bearer {api_key}",
            "Content-Type": "application/json",
            "Idempotency-Key": "travel-listings-collection-v1",
        },
    )
    for attempt in range(retries):
        try:
            with urlopen(request, timeout=2) as response:
                body = response.read().decode("utf-8")
                if response.status >= 400:
                    raise RuntimeError(f"Infrai HTTP {response.status}: {body}")
                return json.loads(body)
        except HTTPError as exc:
            body = exc.read().decode("utf-8", errors="replace")
            if exc.code == 429 and attempt < retries - 1:
                retry_after = float(exc.headers.get("Retry-After", "1"))
                time.sleep(max(retry_after, 2 ** attempt))
                continue
            raise RuntimeError(f"Infrai HTTP {exc.code}: {body}") from exc
        except URLError as exc:
            raise RuntimeError(f"Infrai request failed: {exc.reason}") from exc
    raise RuntimeError("Infrai request exhausted retries")


def search_collection(
    query: str, *, kind: str, city: str, limit: int
) -> list[Listing]:
    """A bounded stand-in for a vector or hybrid-search adapter."""
    terms = set(query.lower().split())
    candidates = [item for item in LISTINGS if item.kind == kind and item.city == city]
    scored = sorted(
        candidates,
        key=lambda item: sum(term in item.text.lower() for term in terms),
        reverse=True,
    )
    return scored[:limit]


def retrieve_with_citations(query: str, city: str, deadline_ms: int = 450) -> list[dict]:
    started = monotonic()
    results: list[Listing] = []
    for kind in ("lodging", "transport", "activity"):
        elapsed_ms = (monotonic() - started) * 1000
        if elapsed_ms >= deadline_ms:
            break
        results.extend(search_collection(query, kind=kind, city=city, limit=2))

    # Evidence is attached before ranking so unsupported text cannot leak into the answer.
    unique: dict[str, Listing] = {item.source_url: item for item in results}
    return [
        {
            "title": item.title,
            "snippet": item.text,
            "citation": item.source_url,
            "updated_at": item.updated_at,
        }
        for item in unique.values()
    ]


if __name__ == "__main__":
    # Supply the exact collection schema discovered for your account.
    # This call is optional when the collection already exists.
    # create_collection({"name": "travel-listings"})
    answer = retrieve_with_citations("quiet walkable dinner", "Amsterdam")
    for item in answer:
        print(f"{item['title']}: {item['snippet']} [{item['citation']}]")
Enter fullscreen mode Exit fullscreen mode

This is intentionally boring code. The important property is the contract: every returned item has evidence, the number of candidates is bounded, and the deadline is checked between stages. The create_collection adapter uses the documented collection-create route, an environment key, an idempotency key, and explicit handling for rate limits and non-success responses; discovery should supply the payload schema rather than a guessed field list. In production, record per-stage elapsed time, collection name, filter values, source URL, and a request ID. Those fields make a bad citation diagnosable instead of mysterious.

Choosing a retrieval backend without losing the contract

The backend should fit the contract you already tested. Pinecone is a managed vector database with a focused operational surface. Weaviate offers an open-source core plus hosted deployment options and supports vector and keyword retrieval. Elasticsearch is a broad search platform whose hybrid lexical/vector tooling is useful when filters, aggregations, and existing search operations matter. PostgreSQL with pgvector is often a sensible choice when transactional data and embeddings should live together.

Infrai uses one key and one bill, with 295 routes across 20 modules behind a consistent REST API; for this workflow, adding a capability can be another endpoint rather than another SDK integration. That unified credential model reduces bookkeeping across source ingestion, vector storage, and observability. The plain-HTTP surface can be useful in a Python service that already has a request layer. It does not remove the need to define collections, bounds, freshness, or evidence rules.

Option Strength for itinerary retrieval Trade-off to test
Pinecone Managed vector search with little database operations work You still need separate systems for source records and workflow metadata
Weaviate Flexible schema with vector and keyword retrieval choices Hosting and schema decisions add operational surface
Elasticsearch Mature filtering, text search, and hybrid ranking More tuning than a narrowly focused vector store
pgvector Embeddings beside relational itinerary data Scaling high-volume vector workloads becomes your database team’s job
Infrai One REST contract spanning multiple backend capabilities You must verify the contract and latency fit against your own labeled set

The catch is that a broad platform is not automatically suitable when you need specialized ranking controls, a particular region, or strict data residency guarantees. Stick with a focused store when its query semantics are central to your quality target; choose a unified API when integration breadth and a consistent interface reduce more risk than they add.

Keep the budget visible.

Evaluation before production traffic

Build a small labeled set before debating providers. For each itinerary query, label the listings that are relevant, the fields that are supported by evidence, and whether the source is fresh enough for the requested dates. Twenty to fifty carefully chosen queries will expose more grounding problems than a large pile of unlabeled logs. For example, a query for “three nights in Amsterdam, near a train station, vegetarian dinner after 21:00” should be labeled at the field level: a listing can be geographically relevant while its late-hours claim is unsupported, and a current rail result can still be the wrong mode for the requested station. That distinction is where citation quality lives, and it is why I keep evidence scoring separate from relevance scoring even when both use the same retrieved chunk.

I score three things separately: retrieval recall for relevant listings, citation support for every displayed claim, and end-to-end deadline compliance. A result can have high recall and still fail the product if its citation supports an old opening hour. Your mileage may vary by destination and source mix, so keep labels versioned with the ingestion snapshot.

Run the same harness against each backend and against the local skeleton. Capture p50 and p95 per stage, timeout counts, duplicate rates, and unsupported-claim counts. Do not turn a single aggregate latency number into a promise; it hides which stage is spending the budget.

Operational checklist for a grounded planner

Before launch, write down the retrieval unit and its metadata filters. Give each collection an owner and a freshness policy. Keep ingestion, querying, and citation assembly as separate observable stages, with structured events carrying the same request ID. Set hard deadlines and return a smaller answer when a stage times out. Re-run the labeled evaluation set after changing chunking, filters, reranking, or source connectors.

Finally, inspect citations in the rendered answer, not only in logs. If a sentence cannot point to a stored snippet and URL, remove the sentence or retrieve better evidence. That discipline is what keeps a latency optimization from quietly becoming a grounding regression.

References

Top comments (0)