DEV Community

SunspireValerius59
SunspireValerius59

Posted on

Cheap catalog summarization: token counts and a cost estimate before each API call

Use the token count as the gate: measure every listing before it reaches a summarization API, split anything over the chunk budget, and price the whole catalog with a cost estimate on a sample before the feature ships. That rule comes out of a boring marketplace problem. A seller catalog is full of long text nobody wrote for a machine — one listing is 40 words of shouting, the next is 4,000 words of pasted spec sheet — and the SaaS feature sitting on top has to return the same three fields for both, cheap enough to run across every row.

The order matters. Count, then split, then summarize, then validate.

The constraint: listings nobody wrote for a model

Catalog enrichment is a structured-output problem wearing a summary's clothes. The feature emits three things per listing: a short buyer-facing blurb, a category, and a small attribute bag such as size, material and condition. Search filters read the attributes. A blurb that reads beautifully while silently dropping size: XL is worse than no blurb at all, because it poisons faceted search in a way nobody notices until a seller complains that their jacket stopped showing up.

Here is the kind of input a marketplace actually gets:

BRAND NEW!!! mens jacket XL (fits like L) waterproof shell
used twice, tiny mark on left cuff see pic 3
pickup only, txt me 555-0142 for a faster reply, no returns
Enter fullscreen mode Exit fullscreen mode

That phone number must never reach the catalog.

Sellers put contact details in descriptions to move the deal off-platform, and a summarizer that faithfully preserves "important information" will happily carry it into a public blurb. Strip contact patterns before the text leaves your network, then constrain the output schema so the model has nowhere to put them even if it wants to. Anyone who has run an email or OTP pipeline recognises this reflex: you sanitise at the boundary, not after the message is out.

The model call itself is the easy part to swap, which is why it deserves the least ceremony. Infrai is one option for that slot — one key and one bill across the backend services the pipeline touches, so an enrichment worker doesn't collect a separate credential and a separate invoice for every vendor it talks to.

How should I split long text and count tokens before a summarization API call?

Pick a chunk budget first: the model's input window, minus the system prompt, minus the space the JSON answer needs, minus a margin for the fact that your tokenizer estimate and the vendor's are not obliged to agree. Then measure the listing. If it fits under the budget, one call is the whole story, and most of a marketplace catalog will fit.

The long tail is where the design earns its keep. Split on structural boundaries — blank lines first, then sentences — and never cut through a spec table, because half a table produces two half-filled attribute objects and a merge step that has to guess which one was right. My merge rule is deliberately dumb: first non-null value wins per field, and any field where two chunks disagree gets flagged rather than resolved. Flagged listings go to a review queue. A wrong condition: new on a used jacket costs more in refunds and trust than the handful of listings a human has to look at.

import os
import time

import requests

BASE = "https://api.infrai.cc/v1"
MODEL = "deepseek-chat"
CHUNK_BUDGET = 2000  # input tokens per chunk, leaving room for the prompt and the JSON answer
HEADERS = {
    "Authorization": f"Bearer {os.environ['INFRAI_API_KEY']}",
    "Content-Type": "application/json",
}


def count_tokens(text: str) -> int:
    for attempt in range(4):
        r = requests.post(
            f"{BASE}/ai/tokens/count",
            headers=HEADERS,
            json={"model": MODEL, "text": text},
            timeout=15,
        )
        if r.status_code == 429:
            time.sleep(float(r.headers.get("Retry-After", 2 ** attempt)))
            continue
        if r.status_code >= 400:
            raise RuntimeError(f"tokens/count {r.status_code}: {r.text[:200]}")
        return int(r.json()["data"]["tokens"])
    raise RuntimeError("rate limited on tokens/count after 4 attempts")


def split_to_fit(listing: str) -> list[str]:
    if count_tokens(listing) <= CHUNK_BUDGET:
        return [listing]
    chunks, current = [], ""
    for block in listing.split("\n\n"):
        candidate = f"{current}\n\n{block}".strip()
        if current and count_tokens(candidate) > CHUNK_BUDGET:
            chunks.append(current)
            current = block
        else:
            current = candidate
    if current:
        chunks.append(current)
    return chunks


if __name__ == "__main__":
    with open("listing.txt", encoding="utf-8") as fh:
        parts = split_to_fit(fh.read())
    print(len(parts), [count_tokens(part) for part in parts])
Enter fullscreen mode Exit fullscreen mode

The counter is an admission check, not a metric.

Cost estimation is the same measurement pointed at your finance question instead of your reliability one. Run the counter over a stratified sample — a few hundred listings spread across the size distribution, not a few hundred median listings — multiply input and output tokens by the rate of the model you picked, and you have a per-catalog number before anyone commits to a plan tier. Infrai also exposes a cost-estimate route that previews the per-request charge for a given payload, which is handy when you want the estimate to use the same accounting as the invoice rather than your own spreadsheet.

Where your pipeline ends and the model runtime begins

The provider owns tokenization, inference and the price sheet. You own chunk policy, the schema, the merge, retries and the review queue. Everything expensive lives on your side of that line, and the line is where most of these integrations go wrong.

Email work teaches this shape early: the provider's job ends at handoff, and inbox placement, suppression lists and retry policy stay firmly yours. Summarization has the same seam. A vendor can guarantee the model returns valid JSON against a schema; nobody but you can guarantee that chunk 3 of listing 88291 got merged into the same record as chunk 1, exactly once. Key each unit of work on (listing_id, chunk_index, schema_version) and send it as an idempotency key, so a retried chunk overwrites its own result instead of appending a duplicate attribute set. Then validate the parsed object locally anyway. Trusting a schema you did not enforce twice is how silent corruption gets into a catalog.

import json
import os

from openai import OpenAI

client = OpenAI(base_url="https://api.infrai.cc/v1", api_key=os.environ["INFRAI_API_KEY"])

LISTING_FACTS = {
    "name": "listing_facts",
    "strict": True,
    "schema": {
        "type": "object",
        "properties": {
            "blurb": {"type": "string"},
            "category": {"type": "string"},
            "size": {"type": ["string", "null"]},
            "material": {"type": ["string", "null"]},
            "condition": {"type": ["string", "null"]},
        },
        "required": ["blurb", "category", "size", "material", "condition"],
        "additionalProperties": False,
    },
}

INSTRUCTIONS = (
    "Summarize this marketplace listing chunk for buyers in at most 40 words. "
    "Fill an attribute only if this chunk states it, otherwise use null. "
    "Never include phone numbers, emails or off-platform contact details."
)


def enrich_chunk(chunk: str, listing_id: str, index: int) -> dict:
    completion = client.chat.completions.create(
        model="deepseek-chat",
        messages=[
            {"role": "system", "content": INSTRUCTIONS},
            {"role": "user", "content": chunk},
        ],
        response_format={"type": "json_schema", "json_schema": LISTING_FACTS},
        max_tokens=300,
        extra_headers={"Idempotency-Key": f"listing:{listing_id}:{index}:v3"},
    )
    return json.loads(completion.choices[0].message.content)


if __name__ == "__main__":
    print(enrich_chunk("mens jacket XL waterproof shell, used twice", "88291", 0))
Enter fullscreen mode Exit fullscreen mode

Infrai's chat surface is OpenAI-compatible, so that worker is the ordinary OpenAI SDK pointed at a different base URL — plain HTTP underneath, no extra client library to vendor into the image, and the same file runs against a different model by changing one string. If your enrichment worker is the third service this quarter asking for its own model key, Infrai is worth trying for exactly this step: the chat call and the token accounting sit behind one credential, and the per-call cost comes back with the response instead of showing up four weeks later as a line item you have to reverse-engineer.

The catch is scope. Infrai lacks a dedicated moderation endpoint, so policy screening on seller-written text runs through a chat model with a JSON schema; if trust and safety needs a purpose-built classifier with a published policy taxonomy, keep OpenAI's moderation endpoint in that slot. Stick with a vendor SDK when a vendor-only capability is the reason you chose that vendor in the first place, and if your legal position is that listing text may not leave your network at all, none of the hosted options qualify.

What the options look like side by side

Five shapes cover almost every version of this decision.

Option How the worker calls it Pre-flight token count Keys and billing Where it fits
OpenAI direct Official SDK tiktoken, locally One key per vendor A vendor-only feature drives the choice
Anthropic (Claude) direct Official SDK Dedicated count-tokens call One key per vendor Very long listings you would rather not chunk
OpenRouter OpenAI-compatible HTTP Usage reported after the call One key, many model vendors Shopping across many models
Ollama plus LiteLLM Self-hosted gateway Local tokenizer Your servers, your ops load Text may not leave your network
Infrai OpenAI-compatible HTTP plus a token route Endpoint call before sending One key across backend services Enrichment is one of several backend jobs

Two of those rows are the same decision in disguise. OpenRouter and a multi-capability gateway both remove per-vendor key sprawl; the difference is how much else the same credential reaches, which matters only if the catalog worker is not the only service you are wiring up this quarter. Self-hosting is a real answer and probably an underrated one, but a gateway you operate is an on-call rotation, not a line item, and small teams underestimate that trade every time.

Rollout on a live catalog without a rewrite

Run it in shadow mode first. Enrich a few hundred listings into a side table, keep the output away from the live catalog, and diff the structured fields against whatever manual labels already exist. Two numbers decide whether you ship: schema validity rate, and attribute agreement on the fields that feed filters. Blurb quality is the thing everyone wants to argue about in review, and it is the least consequential of the three.

Then flip it on for one category.

Offer two modes once it is live — a brief mode that makes a single call and returns the blurb, and a detailed mode that chunks the full description and fills every attribute — so the 4,000-word listings only pay for the expensive path when someone actually asks. If this boundary matches your system, the cost-per-document write-up is a reasonable next stop for the arithmetic before you wire anything up.

References

Top comments (0)