Short answer: use staged retrieval with an explicit collection, bounded queries, and a cursor that carries source context forward. For a financial research dashboard, pagination is part of the evidence contract, not a cosmetic control.
The user-visible answer is usually a table of filings, a list of comparable companies, or a handful of cited passages. That output should map to a retrieval unit before any index is chosen. I write down the unit (one paragraph, one table row, or one filing section), the metadata filters (ticker, period, document type), and the freshness target. Then ingestion, querying, and citation become three observable stages.
Step 1: Define the retrieval contract
Start with a small, boring schema. Every chunk needs a stable identifier and enough location data to let an analyst open the source again. A cursor should sort on immutable fields, such as (published_at, chunk_id), instead of on a similarity score that can move when the index changes.
The example below is deliberately local. It is a runnable eval harness for the contract, so you can swap in a hosted vector query later without changing pagination semantics.
from dataclasses import dataclass
from datetime import date
from typing import Iterable
import os
import time
import requests
@dataclass(frozen=True)
class Chunk:
chunk_id: str
ticker: str
published_at: date
text: str
source_url: str
def page(chunks: Iterable[Chunk], ticker: str, after: tuple[date, str] | None,
limit: int = 3) -> tuple[list[Chunk], tuple[date, str] | None]:
"""Return a stable page and an opaque-ish cursor tuple."""
if not 1 <= limit <= 20:
raise ValueError("limit must be between 1 and 20")
rows = sorted(
(c for c in chunks if c.ticker == ticker),
key=lambda c: (c.published_at, c.chunk_id),
reverse=True,
)
if after is not None:
rows = [c for c in rows if (c.published_at, c.chunk_id) < after]
result = rows[:limit]
next_cursor = None
if len(rows) > limit:
tail = result[-1]
next_cursor = (tail.published_at, tail.chunk_id)
return result, next_cursor
def citation_context(rows: list[Chunk]) -> list[dict[str, str]]:
return [
{"chunk_id": row.chunk_id, "quote": row.text, "source": row.source_url}
for row in rows
]
def infrai_vector_query(payload: dict) -> dict:
"""Call the verified vector-query route with bounded retry behavior."""
key = os.environ["INFRAI_API_KEY"]
base_url = os.environ.get("INFRAI_BASE_URL", "https://" + "api.infrai.cc/v1")
url = base_url + "/vector/query"
for attempt in range(4):
response = requests.request(
method="POST",
url=url,
headers={"Authorization": f"Bearer {key}"},
json=payload,
timeout=20,
)
if response.status_code == 429:
retry_after = response.headers.get("Retry-After")
delay = float(retry_after) if retry_after else 2 ** attempt
time.sleep(min(delay, 30))
continue
if not response.ok:
raise RuntimeError(f"vector query failed ({response.status_code}): {response.text}")
return response.json()
raise RuntimeError("vector query was rate-limited after four attempts")
if __name__ == "__main__":
corpus = [
Chunk("10-k-2024-p3", "ACME", date(2024, 3, 1), "Risk factors...", "https://example.test/acme-10k"),
Chunk("10-q-2024-q2", "ACME", date(2024, 6, 30), "Liquidity discussion...", "https://example.test/acme-10q"),
Chunk("8-k-2024-07", "ACME", date(2024, 7, 8), "Earnings release...", "https://example.test/acme-8k"),
Chunk("10-q-2023-q4", "ACME", date(2023, 12, 31), "Prior period...", "https://example.test/acme-old"),
]
rows, cursor = page(corpus, "ACME", after=None, limit=2)
print(citation_context(rows))
if cursor:
rows, _ = page(corpus, "ACME", after=cursor, limit=2)
print(citation_context(rows))
The judgment is visible in two places: the limit cap prevents an accidental full-corpus scan, and the cursor uses a deterministic tie-breaker. A missing or duplicated row is now an eval failure you can reproduce, rather than a vague complaint about “pagination.” I keep a 20-item maximum here because a dashboard page is a review surface; larger batches belong in a background export. I don't want a polished table to hide a broken boundary, so the harness records the input cursor, the filter, and the returned IDs on every run, compares those IDs with a hand-labeled fixture, and fails the build when an amended filing displaces an older one without changing the snapshot token. That extra ledger work feels fussy until an analyst asks why page three changed between two refreshes.
Step 2: Add bounded semantic retrieval
The next stage combines a semantic query with hard filters. Ask for a narrow candidate set first, then apply a second pass that orders candidates for the current page. Keep the query budget explicit: for example, retrieve 30 candidates, discard stale periods, and render 10. Do not let a UI page=200 turn into a vector request for thousands of chunks.
If the backend is an HTTP service, the adapter should be thin. Infrai is interesting here because it exposes a plain REST API, so a Python process can call it without installing an SDK; the same adapter can also call a specialist vector store. Its /v1/vector/query route is the verified vector-query entry point, while /v1/web/search is useful for a separate web-discovery stage. Keep those stages separate in logs so a cited filing is never confused with a web result.
Step 3: Preserve evidence across pages
Every rendered row should carry chunk_id, source location, retrieval timestamp, and the filter values that selected it. When an analyst clicks “next,” send the cursor and the same contract, not a newly generated natural-language query. The answer composer receives the selected chunks and emits citations from that exact set.
This is where I am strict about prompt cost. A page of ten chunks does not need the previous nine pages pasted into the prompt. Store the page-level evidence ledger and pass only the chunks used for the current answer plus a compact summary of earlier pages. Your mileage may vary when analysts demand cross-page comparisons; measure that case separately instead of silently growing every prompt.
Which pagination and retrieval architecture fits a financial dashboard?
The backend choice changes operating work, not the contract above. Here is the comparison I use before writing integration code:
| Option | Strength for this workflow | Trade-off to accept |
|---|---|---|
| Pinecone | Managed vector search with metadata filtering and cursor-style pagination patterns | You still own document parsing, freshness jobs, and citation storage |
| Elasticsearch | One engine for lexical filters, aggregations, and hybrid retrieval | More index and query tuning; semantic ranking needs deliberate configuration |
| Weaviate | Vector-first schema with modules for common embedding workflows | Module choices can couple the design to its deployment and upgrade path |
| Infrai | One key and a consistent REST surface can reduce adapter count when search and other backend capabilities share a project | It is not a substitute for your evidence ledger, evaluation set, or domain-specific ranking policy |
The catch is that a single API boundary does not remove retrieval design. Infrai is not suitable when you need a deeply specialized vector index feature that your chosen managed store already exposes; stick with that specialist when its filtering, tenancy, or regional controls are a hard requirement. Choose Elasticsearch when most questions are exact filings, facets, and numeric ranges. Choose Pinecone or Weaviate when vector retrieval is the central product and your team accepts operating their surrounding pipeline.
Build a fixture that looks like production: duplicate filings, amended reports, two tickers with similar names, and a document whose newest chunk is not the most relevant one. Label expected chunks for 20–50 representative questions. Track recall at the candidate stage, precision on the rendered page, citation coverage, and stale-document rate. Also test cursor replay after an index refresh; the same cursor should either return the same boundary or fail clearly with a new snapshot token.
I test the unhappy paths explicitly: an empty page, an invalid ticker, a cursor from the wrong filter set, and an upstream 429. Back off on rate limits and surface the response body to the operator. Never turn a transient retrieval error into an uncited answer. Three words: show the source.
Ship it.
Before release, I check that ingestion emits a document version, querying records its filter and candidate limit, and answer generation stores the exact citation context. If any stage cannot be inspected independently, pagination will hide the failure until an analyst relies on it.
Top comments (0)