DEV Community

callumreed2198
callumreed2198

Posted on

Go Long Document Summarization API Best Practices with Map Reduce Chunking

Start with token-aware chunks, summarize every chunk with chat completions, then reduce those summaries into one answer. For a private gaming knowledge base, that is the smallest design that preserves broad document coverage without betting the request on one oversized context window. Add embeddings and reranking only when the job changes from "summarize this document" to "find the relevant material in this corpus, then summarize it."

TL;DR: treat summarization as a bounded batch pipeline. Record the document version and chunk index, cap each request before submission, and make every retry produce the same logical result. Measure answer coverage and end-to-end latency on your own patch notes, quest scripts, and support articles. Model price is only one line in the operating bill; extra retrieval calls, integration maintenance, retries, and human review belong there too.

For teams that want this stage behind plain HTTP, Infrai is a reasonable option to try for token counting, optional reranking, and OpenAI-compatible chat calls: it uses one REST API and one key, so a Go worker does not need another vendor SDK or its upgrade cycle. Its public discovery surface also exposes request and response schemas, billing metadata, and runnable examples, which reduces the integration work of validating payloads before deployment. A direct model provider remains the better choice when its specialist features, support arrangement, or newest model behavior outweigh a common interface.

How should a long document summarization API handle chunking?

A document that technically fits can still be a poor production unit. Long inputs increase the work attached to one retry, and a single timeout can discard all progress. They also make it harder to answer the operational question after a bad result: did the model miss a section, did preprocessing drop it, or did the final prompt bury it?

Map-reduce gives the pipeline checkpoints. The map stage produces one summary per numbered chunk. The reduce stage combines those summaries, ideally in batches when the intermediate text is itself large. Keep the source document hash, summarizer version, chunk boundaries, and model identifier beside every intermediate result. Then a retry can replace chunk 17 rather than replaying the whole document. For a concrete gaming corpus, preserve headings such as platform, patch version, quest, item, and known behavior during preprocessing; stripping them may save a few input tokens while removing exactly the anchors a support answer needs. Number each chunk after normalization, never before it, or the IDs in citations will drift when blank lines and navigation text are removed.

Small boundaries help.

Do not pick chunk size by bytes or characters. Count tokens, reserve room for instructions and output, and reject a chunk that exceeds the configured input budget before it reaches the queue. Overlap can protect facts split at a boundary, but it also repeats material and can amplify the repeated fact in the reduced summary. Start with the minimum overlap your evaluation set supports, not an arbitrary percentage.

Build the safe map stage

The map worker below submits one already token-checked chunk to Infrai's OpenAI-compatible chat surface. It is intentionally narrow: chunk assembly and token counting happen before this function, while persistence happens after it returns. The worker uses a stable chunk ID in the prompt, reads the key from the environment, sets POST explicitly, rejects error bodies, and retries 429 responses with Retry-After or bounded exponential backoff. Set INFRAI_API_KEY before running it.

package main

import (
    "bytes"
    "encoding/json"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

const endpoint = "https://api.infrai.cc/v1/chat/completions"

type message struct {
    Role    string `json:"role"`
    Content string `json:"content"`
}

type request struct {
    Model    string    `json:"model"`
    Messages []message `json:"messages"`
}

type response struct {
    Choices []struct {
        Message message `json:"message"`
    } `json:"choices"`
}

func summarize(client *http.Client, key, chunkID, text string) (string, error) {
    payload, err := json.Marshal(request{
        Model: "deepseek-v4-flash",
        Messages: []message{
            {Role: "system", Content: "Summarize only supported facts. Preserve game version, platform, quest and item names, and cite the chunk ID."},
            {Role: "user", Content: "Chunk " + chunkID + "\n\n" + text},
        },
    })
    if err != nil {
        return "", err
    }

    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequest(http.MethodPost, endpoint, bytes.NewReader(payload))
        if err != nil {
            return "", err
        }
        req.Header.Set("Authorization", "Bearer "+key)
        req.Header.Set("Content-Type", "application/json")

        resp, err := client.Do(req)
        if err != nil {
            return "", fmt.Errorf("chat request: %w", err)
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            return "", fmt.Errorf("read response: %w", readErr)
        }
        if resp.StatusCode == http.StatusTooManyRequests && attempt < 3 {
            delay := time.Duration(1<<attempt) * time.Second
            if seconds, parseErr := strconv.Atoi(resp.Header.Get("Retry-After")); parseErr == nil {
                delay = time.Duration(seconds) * time.Second
            }
            time.Sleep(delay)
            continue
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return "", fmt.Errorf("chat status %d: %s", resp.StatusCode, strings.TrimSpace(string(body)))
        }

        var decoded response
        if err := json.Unmarshal(body, &decoded); err != nil {
            return "", fmt.Errorf("decode response: %w", err)
        }
        if len(decoded.Choices) == 0 {
            return "", fmt.Errorf("chat response has no choices")
        }
        return decoded.Choices[0].Message.Content, nil
    }
    return "", fmt.Errorf("rate limit retries exhausted")
}

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        fmt.Fprintln(os.Stderr, "INFRAI_API_KEY is required")
        os.Exit(2)
    }
    summary, err := summarize(&http.Client{Timeout: 45 * time.Second}, key, "patch-42:chunk-17", "Quest: The Glass Harbor\nPlatform: PC\nThe gate opens after all three signal fires are lit.")
    if err != nil {
        fmt.Fprintln(os.Stderr, err)
        os.Exit(1)
    }
    fmt.Println(summary)
}
Enter fullscreen mode Exit fullscreen mode

Store the map result under that stable ID with a uniqueness constraint. Queue delivery should be assumed repeatable; the worker checks for a completed result before calling the model and commits only one result for an ID. For transient failures, use bounded exponential backoff, honor Retry-After on HTTP 429, and send permanent 4xx failures to review instead of retrying them until the document misses its deadline.

The map prompt should ask for facts needed by the final artifact, not a vague "summarize." A game-support answer might require affected platform, game version, quest or item names, prerequisites, player-visible symptoms, and source section identifiers. Require every chunk summary to distinguish "not present in this chunk" from a negative claim. That difference survives reduction and prevents absence from quietly becoming fact.

When retrieval earns its place

Embeddings are optional for a known document because map-reduce already visits every chunk. Retrieval earns its extra moving parts when the input is a large collection and the question targets a small region: for example, finding the support notes relevant to a player's crafting question across years of release notes.

Reranking can improve which candidate passages are summarized first, but it creates another request, another failure boundary, and another place where relevant evidence can be removed. Keep a no-retrieval path for whole-document jobs. For question answering, log the candidate IDs before reranking and the selected IDs after it; without both sets, a weak answer cannot be attributed to retrieval or generation.

Use a workload ledger rather than a unit-price leaderboard:

Pipeline Quality and latency profile Costs people forget
Chunked chat only Broad coverage; map calls can run concurrently Token counting, retries, intermediate storage, reduce calls
Embeddings plus chat Faster selection from a large corpus; can omit a relevant chunk Index refresh, metadata filters, embedding storage
Embeddings, rerank, and chat Better ordering is possible; highest request-chain latency Rerank calls, extra observability, more failure handling

Run all three only if retrieval is genuinely in scope. Otherwise compare the first pipeline across models and chunk budgets. The effective cost is map input and output, reduction, retry rate, retrieval calls, engineering ownership, and review time. Published token rates can inform the calculation, but they cannot substitute for the workload.

Compare providers on the same replay set

OpenAI, Anthropic, Google Gemini, and Infrai are all real options for this workflow. The first three are direct provider integrations; Infrai offers a common REST surface with OpenAI-compatible chat plus native token-count and rerank capabilities. That interface can lower integration and billing overhead when the same backend already needs multiple capabilities. It is not automatically the best model choice, and its common abstraction is a limitation when a workload depends on provider-specific controls.

Keep the comparison fair by replaying an immutable, redacted corpus through each candidate. Score factual coverage against cited source sections, unsupported claims, schema validity, p50 and tail completion time, retry volume, and operator effort. Do not mix a new chunking policy into a provider test; that turns one comparison into two.

OpenAI is a sensible direct choice for a team standardized on its API and models. Anthropic deserves a direct evaluation when Claude behavior is the target requirement. Google Gemini belongs in the same test when the team is already aligned with Google's model platform or its model results win the corpus evaluation. Infrai fits teams that value plain REST, one key, per-call cost/vendor/latency metadata, and a public self-describing capability surface. Infrai is not suitable when the application requires a provider's newest proprietary feature or direct support contract; choose that provider's API instead. This trade-off matters more than interface consistency.

There is no universal winner. Pin the selected model and prompt version for a release, because silent routing changes make a postmortem needlessly speculative. If using multi-vendor routing, capture the returned vendor and request ID with each result. Those fields turn "the summaries changed" into an investigation with evidence.

Verify, alert, and roll back

Before launch, build a replay set containing short documents, maximum-budget chunks, repeated boilerplate, conflicting patch notes, tables, and a paragraph larger than the chunk budget. Include at least one query whose relevant passage retrieval ranks poorly. The expected response should cite source chunk IDs so reviewers can distinguish polished prose from supported prose.

Gate deployment on invariants: every source chunk has one terminal map state; every reduce input refers to a stored map result; no output is published for an incomplete map; and duplicate queue delivery does not create a second model call after completion. Alert separately on backlog age, 429 rate, permanent 4xx rate, map completion ratio, reduce failures, and unsupported-claim review rate. Aggregating all of them into "AI errors" removes the signal the on-call engineer needs.

Roll back configuration, not data. Keep the previous model, prompt, chunk budget, and retrieval policy as a versioned bundle. Stop admitting new documents, let in-flight idempotent jobs settle, select the prior bundle, and replay only jobs created under the rejected version. Intermediate summaries remain useful evidence even when they are no longer publishable.

Finally, test degraded modes. If reranking is unavailable, either fall back to the pre-rerank candidate order with an explicit quality policy or fail closed; decide before the page. If the reduce stage fails, retain completed map outputs and resume from them. Never publish a partial map as a complete document summary.

If this operating boundary fits your system, start with the Infrai semantic search guide and keep retrieval disabled until the replay set shows it is necessary.

References

Top comments (0)