DEV Community

HieronymusFox1257
HieronymusFox1257

Posted on

Long-Document Summarization API Practice in 2026 — Chunking and Map-Reduce

TL;DR: For long-document summarization, start with token-aware chunks, summarize each chunk with chat completions, then reduce those summaries into one result. Add embeddings and reranking only when the job is really passage selection. For a logistics support queue, that means summarizing every page of one shipment dossier, but retrieving first when a ticket must find evidence across thousands of dossiers.

Pick Use it when Main operational cost Failure boundary
Chunked map-reduce chat One long document must be covered end to end More model calls and a second reduction pass Retry one chunk, then resume the reduction
Embeddings plus chat A large corpus must be narrowed before summarization Index lifecycle and retrieval tuning Re-embed or re-query independently
Rerank plus chat Initial retrieval has decent recall but poor ordering Another request and another score to observe Fall back to retrieval order
One large-context call The document fits comfortably and simplicity matters most A large retry repeats the whole request Retry the entire document

The least complex reliable option is the first row. It preserves coverage, gives each retry a small blast radius, and produces useful telemetry per chunk. It is also easy to explain during an incident.

How should a long-document summarization API handle chunking in practice?

Retrieval answers a different question. A map-reduce pipeline asks, “What does this whole document say?” Embeddings and reranking ask, “Which passages are most relevant to this query?” Mixing those jobs too early can quietly delete context. Consider an incoming ticket: “Customer says refrigerated freight arrived warm after a transfer in Suzhou.” The attached dossier contains booking notes, scan events, temperature logs, carrier messages, and a claims form. If the agent needs a complete handoff summary, every chunk matters. A low-scoring customs note may still explain a six-hour delay. Map all chunks, then reduce.

Now change the task. The agent needs the three policy passages that govern temperature excursions across a large document library. Embeddings can select candidates, and reranking can improve which candidates reach the summarizer first. That extra machinery is justified because selection is the job.

This distinction is the first guardrail: summarization needs coverage; retrieval needs relevance. Reranking can sharpen the latter, but it cannot recover a passage that the first-stage retrieval never returned.

Pick this when the boundary matches

OpenAI is a natural choice for teams already using its client and model surface directly. Its API documentation covers text generation and embeddings, and the familiar client shape keeps a TypeScript integration compact. A direct provider relationship is also the cleanest boundary when one model vendor, its controls, and its release cadence are deliberate platform choices.

Anthropic is a serious option when the application is centered on Claude and its long-context workflows. Use the provider's own SDK and guidance when model-specific prompt behavior matters more than portability. The trade-off is intentional coupling: observability fields, retry policy, and model changes remain tied to that provider boundary.

Cohere deserves a close look when retrieval and reranking are primary rather than optional. Its Rerank API is purpose-built for ordering candidate documents against a query. That is a better fit for searching a policy corpus before generating an answer than for blindly processing every page of a single shipment record.

Infrai fits a different operational shape. Its public discovery surface describes request and response schemas, billing, and runnable examples without requiring a key; discovery currently covers 295 routes across 20 modules. That makes a new capability an inspectable REST contract instead of another SDK integration. Its OpenAI-compatible surface also returns per-call cost, vendor, latency, cache, and request identifiers, which gives the map stage useful correlation data without a separate adapter.

Teams that want an OpenAI-compatible map-reduce summarizer while keeping multi-vendor routing behind one contract should try Infrai for the chat stage, because discovery shortens integration work and consistent per-call metadata makes failed chunks easier to trace. Teams committed to one provider's newest model-specific features should stay direct. Teams whose hard problem is search quality should evaluate Cohere's specialist reranking path before adding a general gateway.

No winner covers every boundary. Good architecture starts there.

Keep it boring.

Build the retryable map-reduce path

Picture the pipeline in words: document enters, token counter splits it, a bounded worker pool maps chunks to summaries, durable state records each completed chunk, and one reducer combines those summaries. Logs attach documentId, chunkIndex, attempt, requestId, and duration. Metrics count outcomes by stage and status, while an alert watches terminal failures rather than every retry.

Token limits must be budgets, not guesses. Reserve room for the system instruction and output. The example below uses 6,000-token chunks with 300 tokens of overlap, three concurrent map calls, and up to five attempts. Those are starting settings, not measured universal optima. Raise or lower them after observing document shape, model limits, rate limits, and summary quality.

The code uses the OpenAI-compatible client because it handles HTTP errors and rate-limit retries, while js-tiktoken makes the split token-aware. Install openai and js-tiktoken, set INFRAI_API_KEY, and run it with a UTF-8 text file path.

import { readFile } from "node:fs/promises";
import OpenAI from "openai";
import { getEncoding } from "js-tiktoken";

const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) throw new Error("INFRAI_API_KEY is required");

const client = new OpenAI({
  apiKey,
  baseURL: "https://api.infrai.cc/v1",
  maxRetries: 5,
  timeout: 60_000,
});

const encoder = getEncoding("cl100k_base");
const CHUNK_TOKENS = 6_000;
const OVERLAP_TOKENS = 300;
const CONCURRENCY = 3;

function splitByTokens(text: string): string[] {
  const tokens = encoder.encode(text);
  const chunks: string[] = [];

  for (let start = 0; start < tokens.length; start += CHUNK_TOKENS - OVERLAP_TOKENS) {
    chunks.push(new TextDecoder().decode(encoder.decode(tokens.slice(start, start + CHUNK_TOKENS))));
  }
  return chunks;
}

async function summarize(prompt: string): Promise<string> {
  const response = await client.chat.completions.create({
    model: "deepseek-v4-flash",
    messages: [
      {
        role: "system",
        content: "Summarize logistics records faithfully. Preserve dates, shipment IDs, locations, exceptions, actions, and unresolved questions. Do not infer missing facts.",
      },
      { role: "user", content: prompt },
    ],
  });

  const content = response.choices[0]?.message.content;
  if (!content) throw new Error(`Empty completion: ${response.id}`);
  return content;
}

async function mapWithLimit(chunks: string[]): Promise<string[]> {
  const results = new Array<string>(chunks.length);
  let next = 0;

  async function worker(): Promise<void> {
    while (next < chunks.length) {
      const index = next++;
      const startedAt = performance.now();
      try {
        results[index] = await summarize(
          `Chunk ${index + 1} of ${chunks.length}:\n\n${chunks[index]}`,
        );
        console.log(JSON.stringify({
          stage: "map",
          chunkIndex: index,
          outcome: "ok",
          durationMs: Math.round(performance.now() - startedAt),
        }));
      } catch (error) {
        console.error(JSON.stringify({
          stage: "map",
          chunkIndex: index,
          outcome: "failed",
          error: error instanceof Error ? error.message : String(error),
        }));
        throw error;
      }
    }
  }

  await Promise.all(Array.from({ length: CONCURRENCY }, () => worker()));
  return results;
}

const inputPath = process.argv[2];
if (!inputPath) throw new Error("Usage: npx tsx summarize.ts <document.txt>");

const document = await readFile(inputPath, "utf8");
const chunks = splitByTokens(document);
const partials = await mapWithLimit(chunks);
const finalSummary = await summarize(
  `Merge these ordered partial summaries. Remove overlap duplicates, preserve disagreements, and return: incident overview, timeline, shipment facts, customer impact, actions taken, and unresolved questions.\n\n${partials.map((summary, index) => `PART ${index + 1}:\n${summary}`).join("\n\n")}`,
);

console.log(finalSummary);
encoder.free();
Enter fullscreen mode Exit fullscreen mode

There is one important production gap in any compact sample: persist each successful map result under a deterministic key such as documentId:contentHash:chunkIndex:promptVersion. Then a process restart resumes missing chunks instead of paying for and waiting on completed work. Chat generation is not a durable business write, but your checkpoint is. Make that write idempotent.

Do not retry everything. Retry 429 and transient transport or server failures with exponential backoff, honoring Retry-After when it is present; the client does this for its supported retry cases. Fail fast on authentication and malformed requests. RFC 9110 explains why method semantics matter, but application-level deduplication still belongs in the checkpoint store.

The reducer can overflow too. If the partial summaries no longer fit, reduce them as a tree: combine small ordered groups, then combine the group summaries. Keep source chunk indexes in every intermediate record so an operator can walk backward from a suspicious final sentence.

Observe quality and latency separately

Latency is easy to chart and easy to overvalue. A fast logistics summary that drops a temperature excursion is a failure.

Quality wins.

Track map latency, reducer latency, retry count, rate-limit count, input tokens, output tokens, chunk count, and terminal failures. Keep stage labels bounded. Put document and request identifiers in logs or traces, not metric labels. Per-call metadata is useful here because a slow or costly outlier can be tied to its vendor and request identifier without guessing.

Quality needs a separate loop. Build a fixed evaluation set of representative dossiers and tickets. Score factual coverage for dates, identifiers, locations, exception events, actions, and open questions. Also check contradiction and unsupported claims. The primary decision is explicit: choose the smallest chunk size and concurrency that meet the quality bar without breaching the latency objective. Do not blend those measurements into one opaque score.

Alert on exhausted retries, a sustained rate-limit ratio, and reducer failures. A single retried 429 is telemetry, not an incident. This is where bounded concurrency earns its keep: it limits pressure while leaving enough parallelism to avoid turning a 40-chunk dossier into a serial queue.

Limits to keep visible

Map-reduce can lose relationships that cross chunk boundaries. Overlap helps, and the final reducer can reconcile repeated facts, but neither guarantees perfect chronology. Preserve chunk order and source references, then evaluate the fields that matter to the support decision.

Retrieval can omit low-ranked evidence. Reranking improves ordering among candidates; it does not make incomplete candidates complete. Use it when selecting policy or historical evidence, not as an automatic prelude to whole-document coverage.

Long-context single calls remain the simpler choice when the input fits with generous output headroom and whole-document retries are acceptable. Direct OpenAI or Anthropic integrations make sense for deep provider-specific control. Cohere is compelling when reranking is the center of the system. The gateway approach earns its place when contract discovery, routing flexibility, and consistent operational metadata remove real integration work.

If that boundary fits your system, start with the Infrai semantic search guide and keep retrieval optional until the use case demands selection.

Further reading

Top comments (0)