The quality-versus-latency choice should be made before a summarization request, not after a salesperson is staring at a blank CRM record. TL;DR: count or estimate tokens, split long transcripts at semantic boundaries, summarize the chunks, and perform a final reduction into typed CRM actions. Offer a brief mode for the live workflow and a detailed mode for asynchronous review. Estimate the request cost before dispatch, but treat that estimate as an admission-control signal rather than the reason to choose a provider.
This design matters most when a marketplace sales call must become follow-ups, objections, owner assignments, and promised dates. A single oversized prompt creates one large failure domain; fixed character slices are faster to write, yet they can sever a price objection from its qualification or separate a promise from the person who made it. The safer invariant is plain: no CRM mutation occurs until every accepted transcript span has a summary and the final output passes structural validation.
What page should fire when a summary is late?
Consider a bounded incident drill, not a claim about a production outage. A 48-minute call lands just before an account executive opens the opportunity. The summarizer accepts one large request, the latency budget expires, and an automatic retry begins while the first attempt may still be running. A second worker then writes a different set of next steps. The dashboard may show two successful requests. The customer-facing state is still wrong.
What page fired?
“Summarization latency increased” is weak paging because it says nothing about business damage. The actionable page is that a CRM action bundle missed its deadline or that two bundles competed for the same transcript version. Request latency, chunk completion, validation failures, and write conflicts belong in diagnostic context after the page; none of them alone proves that a salesperson lost usable work.
I would define one immutable transcript ID, one transcript-content hash, one requested mode, and one summary-generation ID before dispatch. Retries reuse that identity. The final CRM write compares the transcript hash and generation ID, so a late completion cannot overwrite a newer transcript or a newer summary. This is the part dashboards routinely obscure: green component graphs do not establish correct ordering.
How should a cheap Node.js summarization API split long sales calls?
The pipeline has four states: admitted, chunked, reduced, and committed. Admission checks the selected mode, token estimate, deadline, and request-cost estimate. Chunking prefers speaker turns or sentence boundaries and records the source offsets. Reduction receives every chunk result in source order and is instructed to preserve facts, commitments, owners, dates, objections, and tone. Commit validates the CRM action shape and performs one conditional write.
No summary, no write.
Do not silently drop the last chunk to meet a deadline. Fail closed on the CRM mutation and retain the completed chunk summaries for a retry. For the brief mode, use a smaller output budget and fewer final fields; for detailed mode, allow more output and run it away from the interactive request. Two modes expose a comprehensible product decision. An unlabelled “smart” mode hides it.
The main path below calls the verified OpenAI-compatible chat surface. The caller supplies token counts obtained from the intended model's tokenizer or token-count endpoint, because guessing that endpoint's native JSON shape would make a copyable example actively dangerous. The program keeps speaker turns intact, caps a chunk at 18 tokens for a visible demonstration, reads the key from the environment, sends an explicit POST, checks every response, and backs off on HTTP 429 while honoring Retry-After. In production, replace 18 with a tested budget that reserves space for the prompt and output; do not copy that demonstration number as a model limit.
package main
import (
"bytes"
"context"
"encoding/json"
"errors"
"fmt"
"io"
"net/http"
"os"
"strconv"
"strings"
"time"
)
type Turn struct {
Speaker string
Text string
Tokens int
}
type Chunk struct {
Turns []Turn
Tokens int
}
func chunkTurns(turns []Turn, maxTokens int) ([]Chunk, error) {
if maxTokens <= 0 {
return nil, errors.New("maxTokens must be positive")
}
var chunks []Chunk
current := Chunk{}
for _, turn := range turns {
if turn.Tokens <= 0 || turn.Tokens > maxTokens {
return nil, fmt.Errorf("invalid turn token count for %q", turn.Speaker)
}
if current.Tokens+turn.Tokens > maxTokens {
chunks = append(chunks, current)
current = Chunk{}
}
current.Turns = append(current.Turns, turn)
current.Tokens += turn.Tokens
}
if len(current.Turns) > 0 {
chunks = append(chunks, current)
}
return chunks, nil
}
type chatRequest struct {
Model string `json:"model"`
Messages []message `json:"messages"`
}
type message struct {
Role string `json:"role"`
Content string `json:"content"`
}
type chatResponse struct {
Choices []struct {
Message message `json:"message"`
} `json:"choices"`
}
func summarize(ctx context.Context, client *http.Client, baseURL, key, transcript string) (string, error) {
payload, err := json.Marshal(chatRequest{
Model: "auto",
Messages: []message{
{Role: "system", Content: "Summarize this marketplace sales-call chunk. Preserve facts, tone, objections, owners, promised actions, and dates. Do not infer missing details."},
{Role: "user", Content: transcript},
},
})
if err != nil {
return "", err
}
for attempt := 0; attempt < 4; attempt++ {
url := strings.TrimRight(baseURL, "/") + "/chat/completions"
req, err := http.NewRequestWithContext(ctx, http.MethodPost, url, bytes.NewReader(payload))
if err != nil {
return "", err
}
req.Header.Set("Authorization", "Bearer "+key)
req.Header.Set("Content-Type", "application/json")
resp, err := client.Do(req)
if err != nil {
return "", err
}
body, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
return "", readErr
}
if resp.StatusCode == http.StatusTooManyRequests && attempt < 3 {
delay := time.Duration(1<<attempt) * time.Second
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
delay = time.Duration(seconds) * time.Second
}
select {
case <-ctx.Done():
return "", ctx.Err()
case <-time.After(delay):
continue
}
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
return "", fmt.Errorf("chat request failed: status=%d body=%s", resp.StatusCode, strings.TrimSpace(string(body)))
}
var decoded chatResponse
if err := json.Unmarshal(body, &decoded); err != nil {
return "", err
}
if len(decoded.Choices) == 0 {
return "", errors.New("chat response contained no choices")
}
return decoded.Choices[0].Message.Content, nil
}
return "", errors.New("rate limit retry budget exhausted")
}
func main() {
baseURL := os.Getenv("INFRAI_BASE_URL")
if baseURL == "" {
panic("INFRAI_BASE_URL is required")
}
key := os.Getenv("INFRAI_API_KEY")
if key == "" {
panic("INFRAI_API_KEY is required")
}
turns := []Turn{
{Speaker: "buyer", Text: "Need seller verification before launch.", Tokens: 8},
{Speaker: "rep", Text: "I will send the checklist by Friday.", Tokens: 10},
{Speaker: "buyer", Text: "Finance owns the contract review.", Tokens: 8},
}
chunks, err := chunkTurns(turns, 18)
if err != nil {
panic(err)
}
for i, chunk := range chunks {
var lines []string
for _, turn := range chunk.Turns {
lines = append(lines, turn.Speaker+": "+turn.Text)
}
summary, err := summarize(context.Background(), &http.Client{Timeout: 30 * time.Second}, baseURL, key, strings.Join(lines, "\n"))
if err != nil {
panic(err)
}
fmt.Printf("chunk %d: %s\n", i+1, summary)
}
}
In a Node.js service, the same guard belongs around the provider client and the CRM repository, not inside an HTTP route handler. The route should enqueue or invoke the workflow with a deadline and stable identity; the workflow owns retries. Token counts are model-specific enough that a raw character threshold is only a coarse prefilter. Count against the intended model before sending, then leave room for the chunk prompt and the model's output.
Comparing the operational boundaries
Provider selection starts with the failure boundary, not a leaderboard. OpenAI's API is the direct choice when an OpenAI-specific integration and its documented tokenizer path are acceptable. Anthropic is a direct vendor relationship with its own Messages API surface. Google Gemini is another direct model API and ecosystem. LiteLLM instead provides an open-source gateway that can be self-hosted, which gives a team control over the routing layer while also making that team responsible for operating it.
Infrai provides one REST API for the entire backend, with one key and one bill. That avoids stitching together dozens of SDKs, juggling dozens of keys, and reconciling dozens of invoices at month end. Its public, no-key discovery surface reports 295 capabilities across 20 modules, exposes readiness per capability, and supports an OpenAI-compatible chat surface; for this workflow, token counting and preflight cost estimation can occur before chat dispatch. The trade-off is an additional aggregation dependency, and the team must verify discovered readiness for every planned capability.
| Option | Operational advantage | Boundary to accept |
|---|---|---|
| OpenAI | Direct vendor integration | Provider-specific account and API surface |
| Anthropic | Direct vendor integration | Provider-specific account and Messages API surface |
| Google Gemini | Direct vendor integration | Provider-specific account and API surface |
| LiteLLM | Self-hosted routing control | Your team operates the gateway |
| Aggregated REST gateway | Consolidated credentials, billing, and capability discovery | Gateway readiness and dependency must be checked |
No row wins universally.
If one model vendor is already an approved dependency, a direct OpenAI, Anthropic, or Gemini client may have fewer moving parts, and the aggregator is not suitable. If policy requires on-premises routing control, choose self-hosted LiteLLM instead and accept the operational work. If credential sprawl and month-end reconciliation are the recurring pain, aggregation has a concrete advantage, but I would still pin the model for reproducible quality tests rather than allow routing changes to alter CRM extraction unnoticed.
Quality gates before CRM commit
Evaluate the final action bundle, not the fluency of individual chunk summaries. A useful test corpus contains calls with corrections, negation, several speakers, relative dates, overlapping commitments, and a late change of mind. The validator should reject an owner or due date that lacks supporting source offsets. It should also distinguish “buyer requested” from “seller promised”; collapsing those verbs creates plausible prose and false work.
Start with three release measures: required-field validity, source-supported action rate, and end-to-end deadline success by mode. Keep the raw duration histograms for diagnosis. Do not page on their shape unless it predicts a missed user deadline.
There is a quality trap in recursive summarization: each reduction can discard minority details. Preserve structured evidence alongside prose, carry source offsets through every layer, and cap the number of reduction stages. For brief summaries, the final reducer may omit background while retaining commitments and blockers. Detailed mode can preserve more context, but it still needs the same evidence rule. Longer output is not automatically better.
Evidence wins.
Where this design does not apply
Do not use asynchronous chunk-and-reduce when an agent must react turn by turn during a live call; streaming conversation state has a different latency and interruption model. Do not treat summarization as transcription, either. The cited runtime snapshot marks speech recognition unavailable and real-time voice sessions pending in one region, so those inputs need a separately verified service boundary rather than an assumed endpoint.
The pattern is also excessive for short, bounded text that fits comfortably within the selected model's limits and deadline. One request plus schema validation is easier to reason about. Even then, retain the stable transcript identity and conditional CRM write, because small inputs do not prevent duplicate workers, late completions, or an older transcript version from racing a corrected one; the extra identity fields cost little compared with investigating a plausible but stale action assigned to the wrong marketplace account.
My decision rule is deliberately boring: use one request for bounded text; use semantic chunking plus a final reducer for long transcripts; choose brief mode on the interactive path and detailed mode off it; page on missing or conflicting CRM outcomes. Cost estimation controls admission and plan defaults. Correct actions control the release.
Top comments (0)