DEV Community

AshwhisperTorvin64
AshwhisperTorvin64

Posted on

Sales Call Text API: Summarization via Chat Completions JSON

Put a versioned, provider-neutral result contract between a Node.js text summarization API, its chat completions adapter, and every CRM write; reject any JSON output that cannot satisfy it. Provider portability matters here, but the deciding operational constraint is simpler: code that must summarize a long sales-call transcript must never turn plausible prose into an untraceable customer-record mutation.

TL;DR: accept long transcripts or article-like source text through a bounded chunk-and-reduce pipeline, request JSON that matches one small schema, validate it outside the model, attach every proposed action to source evidence, and make the final CRM write idempotent. Keep the provider adapter thin enough to replace, while treating prompts, schemas, and evaluation fixtures as application code.

This is the page I would want to fire: "validated CRM-action rate fell below its normal band." A chart showing token volume, latency, or successful HTTP responses is supporting evidence, not the incident. Green requests can still produce invented follow-ups, omit the decision-maker, or assign an action to the wrong owner.

How should a Node.js text summarization API handle chat completions?

The useful signal sits at the application boundary. For each completed sales call, record whether the pipeline produced a schema-valid result, whether every action carried a source quote or transcript span, whether the CRM write was accepted exactly once, and whether a human later corrected a material field. The first three can gate automation immediately; the last one is a delayed quality signal.

Do not page on one malformed response. Models, networks, and upstream transcript feeds all have transient failures, and a retry queue is built for that. Page when the system is losing its ability to create verified actions: for example, a sustained collapse in validated results while transcript intake remains normal, or a growing queue whose oldest item threatens the business deadline. The exact threshold belongs to your traffic pattern and service objective; inventing a universal percentage would make the runbook look precise while making the alert worse.

Silence is dangerous too. A missing-call detector should compare expected completed calls with summarization jobs received. Otherwise a broken webhook can leave the processing dashboard perfectly green because nothing reached it. Ask what page fired, then ask what population vanished from the denominator.

The output contract should distinguish absence from uncertainty. An empty next_actions array means the call contained no supported action. A failed validation means automation has no result. Those states must not collapse into the same blank CRM field.

That's the boundary.

Build a contract that survives adapter changes

Keep the domain object smaller than the provider response. A useful record for this job contains a concise summary, explicit actions, owners only when stated, due dates only when stated, and evidence copied from the transcript. Provider-specific request IDs and usage counters belong in execution metadata, where they help investigation without leaking into the CRM contract.

The application may be written in Node.js even though the example below is Go; the contract is an HTTP and JSON boundary, not an SDK abstraction. A Node.js adapter should perform the same steps: send the system instruction and source text, receive the chat completion, decode one JSON value, reject unknown fields, and return a domain result rather than the provider's response object. This separation is deliberate. It lets a team change its runtime or model endpoint without changing what a CRM action means, while keeping the operational example in one language throughout.

Here is the core boundary in Go. The HTTP details vary by adapter, but validation and commit rules do not.

package summary

import (
    "context"
    "errors"
    "strings"
)

type Action struct {
    Task     string `json:"task"`
    Owner    string `json:"owner,omitempty"`
    DueDate  string `json:"due_date,omitempty"`
    Evidence string `json:"evidence"`
}

type Result struct {
    SchemaVersion string   `json:"schema_version"`
    Summary       string   `json:"summary"`
    NextActions   []Action `json:"next_actions"`
}

type GenerateRequest struct {
    SystemPrompt string
    Input        string
    SchemaName   string
}

type Generator interface {
    GenerateJSON(context.Context, GenerateRequest) ([]byte, error)
}

func Validate(r Result, transcript string) error {
    if r.SchemaVersion != "1" || strings.TrimSpace(r.Summary) == "" {
        return errors.New("invalid result envelope")
    }
    for _, a := range r.NextActions {
        if strings.TrimSpace(a.Task) == "" || strings.TrimSpace(a.Evidence) == "" {
            return errors.New("action lacks task or evidence")
        }
        if !strings.Contains(transcript, a.Evidence) {
            return errors.New("action evidence is not present in transcript")
        }
    }
    return nil
}
Enter fullscreen mode Exit fullscreen mode

The substring check is intentionally modest. It proves that evidence was copied from the submitted transcript; it does not prove that the action is a faithful interpretation. Normalizing whitespace before comparison can handle transcript formatting, but fuzzy matching deserves caution because it can make fabricated evidence appear valid. Semantic correctness still needs an evaluation set and, for consequential fields, human approval. This is a real limitation, not a parser detail.

Long input changes the implementation. Split on speaker turns or sentence boundaries, preserve stable span identifiers, summarize each bounded segment into candidate facts, then reduce those candidates into the final contract. Never cut blindly at a byte offset: UTF-8 code points, words, and conversational turns are different boundaries. The provider publishes a context limit for a model, but that limit is not an input budget; instructions, schema, candidates, and output all consume the same context window.

A map-reduce pipeline can also create duplicate actions. Use a deterministic action key derived from the call ID plus normalized task, owner, and due date, then send that key through the writer. The CRM side should enforce uniqueness or expose an equivalent conditional-write mechanism. A retry after a timeout must not create a second follow-up.

The model proposes; deterministic code authorizes. That division is the portability boundary worth defending.

The trade-off is extra application code and a narrower output than free-form summarization. This pattern is not suitable when exploratory prose is the actual deliverable and no downstream system will act on individual fields; in that case, keep the result in a review surface instead of pretending it is a CRM command. For automated actions, the narrower contract is worth the maintenance because it creates a place to stop unsafe writes.

Safe execution and error handling

Parse with a strict decoder, reject trailing data, validate the complete object, and only then enqueue a CRM command. The worker should classify failures rather than putting every error through the same retry loop. Transport timeouts and rate limits may be retryable. Invalid JSON can be retried a small, bounded number of times with the same immutable source, but repeated schema failure belongs in a dead-letter path for inspection. Missing evidence is a quality failure, not a network failure.

package summary

import (
    "bytes"
    "context"
    "encoding/json"
    "errors"
    "io"
)

var ErrInvalidOutput = errors.New("invalid model output")

func Generate(ctx context.Context, g Generator, transcript string) (Result, error) {
    raw, err := g.GenerateJSON(ctx, GenerateRequest{
        SystemPrompt: "Return only the requested sales-call summary object. Copy evidence exactly. Do not infer owners or dates.",
        Input:        transcript,
        SchemaName:   "sales_call_summary_v1",
    })
    if err != nil {
        return Result{}, err
    }

    dec := json.NewDecoder(bytes.NewReader(raw))
    dec.DisallowUnknownFields()

    var result Result
    if err := dec.Decode(&result); err != nil {
        return Result{}, errors.Join(ErrInvalidOutput, err)
    }
    if err := ensureEOF(dec); err != nil {
        return Result{}, errors.Join(ErrInvalidOutput, err)
    }
    if err := Validate(result, transcript); err != nil {
        return Result{}, errors.Join(ErrInvalidOutput, err)
    }
    return result, nil
}

func ensureEOF(dec *json.Decoder) error {
    var extra any
    if err := dec.Decode(&extra); !errors.Is(err, io.EOF) {
        if err == nil {
            return errors.New("trailing JSON value")
        }
        return err
    }
    return nil
}
Enter fullscreen mode Exit fullscreen mode

There is a trap in strict parsing: changing the schema is a deployment. If one producer starts emitting version 2 while old consumers only understand version 1, a superficially harmless prompt release becomes an outage. Version the schema, deploy readers before writers, and retain fixtures for both sides of the transition. JSON Schema Draft 2020-12 can document and validate richer contracts, but introducing it also adds validator compatibility and schema-lifecycle work; a small hand-checked structure can be the better choice while the contract remains this compact.

Avoid logging raw transcripts or complete model outputs by default. Sales calls can contain personal data, commercial terms, credentials spoken aloud, and material that does not belong in an incident channel. Logs need call identifiers, adapter and model identifiers, prompt and schema versions, timing, input and output size, validation class, retry count, and a content hash. Give responders a controlled path to retrieve source data under the organization's access and retention policy.

For offline or non-urgent backfills, asynchronous batch processing can separate bulk load from the interactive queue. That changes delivery timing and cancellation semantics, so it should be a distinct execution mode rather than an invisible adapter optimization. Do not let a historical backfill consume the capacity reserved for calls that just ended.

Verification before CRM writes

A unit test proving that JSON decodes is necessary and unimpressive. Build a fixed evaluation corpus from appropriately governed, de-identified transcripts and include the ugly cases: no action at all, a date mentioned as history rather than a deadline, two people with the same first name, corrections late in the call, contradictory speakers, and a transcript longer than one model request can accept. Expected results should permit legitimate wording variation while remaining exact about actions, owners, dates, and evidence.

Run that corpus against every prompt, schema, chunking, model, or adapter change. Compare field-level precision and recall for CRM actions, unsupported-action rate, evidence-match rate, validation failures, and end-to-end latency. Token counts and request cost help capacity planning, but they do not decide whether an action is safe to write.

Then shadow production traffic without writing. The candidate path receives the same immutable transcript, emits its proposed contract to a restricted comparison store, and is judged against the current path and later human edits. Shadowing catches distribution differences that a curated corpus misses, yet it must obey the same data controls as production; duplicating sensitive content into an experimental log is not a harmless test.

A compact release sequence is enough:

  1. Run deterministic parser and validator tests.
  2. Pass the governed offline evaluation thresholds.
  3. Shadow a representative traffic slice with CRM writes disabled.
  4. Canary writes for a bounded cohort, with the old path still available.
  5. Expand only while validation, queue age, duplicate prevention, and human-correction signals remain inside their established bands.

No dashboard earns trust merely by being green. During the canary, sample source transcript spans beside proposed actions and confirm that the alert routes to someone who can stop writes. Test the page. A runbook nobody has exercised is documentation, not a control.

Roll back the writer, not the evidence

The fastest safe rollback is a kill switch at the CRM command boundary. Stop new automated writes, keep accepting transcripts, and retain validated proposals in a bounded queue for replay. This preserves evidence while preventing uncertain mutations. If the queue approaches its retention or age limit, degrade to human review or pause intake according to the business runbook; do not silently discard calls.

Roll back prompt, model, adapter, and schema versions as one recorded release unit. Replaying with an older unit is safe only if the source transcript is immutable and the idempotency key remains stable. Never reuse an idempotency key for a materially different action set, because the writer could mistake changed work for a duplicate.

Recovery needs a reconciliation pass. Compare accepted command IDs with CRM records, locate calls with validated output but no confirmed write, and replay only those commands. The incident timeline should include the first bad release, the first failed validation or correction signal, the page that fired, the kill-switch time, and the population reconciled.

Portability is proven during rollback. If changing providers requires rewriting validation, action semantics, observability, or reconciliation, the adapter boundary was cosmetic. Keep those controls owned by the application, test each adapter against the same contract, and choose a provider using observed quality and operational fit on your governed workload rather than a feature checklist.

References

Top comments (0)