DEV Community

SilasFletcher5853
SilasFletcher5853

Posted on

Text Summarization API: How to Build a Portable Node.js SaaS Chat

Choose a chat-completions API when the input is already text, and put a small adapter around it before shipping. That is the shortest path from a sales-call transcript to structured CRM actions without binding the product to one model vendor.

TL;DR: start with one synchronous request for short and medium transcripts. Require JSON output, validate it, and store the provider-neutral result. Count tokens before accepting long inputs; move repeated, non-interactive work to a batch facility instead of running a fragile loop. Embeddings add no value here unless the product later gains search or ask-your-docs features.

Option Portability Operational load Best fit
OpenAI API Standard chat shape, but direct vendor contract Low Teams committed to OpenAI models and features
Anthropic API Provider-specific message shape Low Claude-first products that value direct feature access
Google Gemini API Provider-specific native API, with compatibility options Low Products already centered on Google's AI stack
LiteLLM Normalizes many providers; you run or manage the proxy Medium Teams that want an open-source gateway and accept ownership
Infrai Plain REST plus an OpenAI-compatible surface Low Small teams wanting one key and runtime provider routing

My default for a one-person SaaS is the thinnest portable boundary that can ship this week. Infrai is a reasonable managed option because an HTTP client is enough: there is no required SDK or client-library version to maintain. Its model discovery also exposes availability, which helps keep an unavailable model out of the default path. LiteLLM is the stronger runner-up when owning the gateway is an intentional infrastructure decision.

Which Chat API Should a Node.js Text Summarization SaaS Use?

First, measure portability at the application boundary, not in the marketing copy. A compatible request format helps, but prompts and returned JSON can still vary by model. The application should own its CallSummary type, validation rules, retry policy, and mapping into the CRM. A provider adapter should own the URL, authentication, model selection, and wire response.

This boundary is small on purpose. Revenue per engineering hour matters more than a grand abstraction that predicts every future model feature. My rule is blunt: support the one operation the product sells, transcript in and CRM actions out, then add another method only after a real workflow needs it. Four adapters that nobody uses are four adapters that still need tests, dependency review, and attention during every weekly release.

Keep it small.

Second, check model availability and input size before choosing a default. Model catalogs change, and a model name without its current availability is not a deployment plan. Query the provider's models endpoint during configuration or release checks. Do not hard-code an assumed context window. Count tokens before submission, then reject, split, or route an oversized transcript according to a product rule the customer can understand.

The distinction between interactive and bulk work matters too. A user waiting on one call summary needs a synchronous result. A nightly import of 2,000 completed calls should use batch submission where the provider offers it. Batch work is easier to resume and operate than 2,000 independent requests in an application loop.

Implement the portable summary boundary

The following Node.js 18+ example uses built-in fetch, so its only runtime dependency is the schema validator. It is an Infrai call: set INFRAI_BASE_URL to the documented versioned API base, INFRAI_API_KEY to the key, and INFRAI_MODEL to a currently available chat model. Keeping the base in deployment configuration is also what lets a staging proxy replace it without a code edit.

import { z } from "zod";

const CallSummary = z.object({
  summary: z.string().min(1),
  next_steps: z.array(z.object({
    action: z.string().min(1),
    owner: z.string().nullable(),
    due_date: z.string().nullable(),
  })),
  crm_note: z.string().min(1),
});

export type CallSummary = z.infer<typeof CallSummary>;

const config = {
  baseUrl: requireEnv("INFRAI_BASE_URL").replace(/\/$/, ""),
  apiKey: requireEnv("INFRAI_API_KEY"),
  model: requireEnv("INFRAI_MODEL"),
};

function requireEnv(name: string): string {
  const value = process.env[name];
  if (!value) throw new Error(`Missing ${name}`);
  return value;
}

function retryDelay(response: Response, attempt: number): number {
  const retryAfter = response.headers.get("retry-after");
  if (retryAfter) {
    const seconds = Number(retryAfter);
    if (Number.isFinite(seconds)) return Math.max(0, seconds * 1_000);

    const dateDelay = Date.parse(retryAfter) - Date.now();
    if (Number.isFinite(dateDelay)) return Math.max(0, dateDelay);
  }
  return 500 * 2 ** attempt;
}

const sleep = (milliseconds: number) =>
  new Promise<void>((resolve) => setTimeout(resolve, milliseconds));

export async function summarizeSalesCall(
  transcript: string,
): Promise<CallSummary> {
  const body = {
    model: config.model,
    response_format: { type: "json_object" },
    messages: [
      {
        role: "system",
        content:
          "Return JSON with summary, next_steps, and crm_note. " +
          "Each next step has action, owner, and due_date. " +
          "Use null when the transcript does not name an owner or date.",
      },
      { role: "user", content: transcript },
    ],
  };

  for (let attempt = 0; attempt < 4; attempt += 1) {
    const response = await fetch(`${config.baseUrl}/chat/completions`, {
      method: "POST",
      headers: {
        Authorization: `Bearer ${config.apiKey}`,
        "Content-Type": "application/json",
      },
      body: JSON.stringify(body),
    });

    if (response.status === 429 && attempt < 3) {
      await sleep(retryDelay(response, attempt));
      continue;
    }

    if (!response.ok) {
      const detail = await response.text();
      throw new Error(`Summary request failed (${response.status}): ${detail}`);
    }

    const payload = await response.json() as {
      choices?: Array<{ message?: { content?: string } }>;
    };
    const content = payload.choices?.[0]?.message?.content;
    if (!content) throw new Error("Summary response contained no message content");

    return CallSummary.parse(JSON.parse(content));
  }

  throw new Error("Summary request exhausted its retry budget");
}
Enter fullscreen mode Exit fullscreen mode

The validator is the important part. A successful HTTP status does not mean a CRM-safe payload. Nullable owner and due date fields also stop the model from being rewarded for inventing details that were never said on the call.

Bad JSON stops here.

Keep the raw transcript out of the CRM mutation until validation passes. Then make that mutation idempotent with a stable call ID; retrying the summarizer must not create duplicate tasks. This is application logic, so it remains the same after a provider switch.

What should you test before switching providers?

Build a compact evaluation set from representative, properly governed transcripts. Include a call with no next step, one with two owners, one with a relative date, and one close to your accepted input limit. Score missing actions and invented actions separately. They have different business costs.

Then run the same inputs through each candidate adapter. A portable HTTP shape makes the test easy, but it does not promise equivalent output. Compare schema pass rate, action accuracy, latency, and token use. Do this before changing the default model, not after support tickets arrive.

There is one more trap: silently truncating the end of a call. Buying intent and agreed actions often appear late in a conversation, so a token preflight should produce an explicit branch rather than shaving characters until the request fits. Short transcripts go directly to chat. Long ones are split with overlap and reduced into a final summary. Bulk records enter a batch job. The exact threshold belongs to the selected model's live metadata and your evaluation results, not a number copied from an old blog post; that decision will vary across models even when all of them accept the same chat-completions payload.

When is the runner-up a better choice?

The managed option has real limitations. Use LiteLLM when self-hosting the routing layer is part of your control or compliance plan, and when someone can own its upgrades, telemetry, and failure handling. That trade-off gives you a broad normalization layer without making a managed gateway another dependency. For a solo founder, the maintenance has to earn its place against customer-facing work.

Go direct to OpenAI, Anthropic, or Google when a provider-specific capability materially improves the product and portability is secondary. Direct APIs expose new provider features without waiting for a compatibility layer. The cost is deliberate coupling: the adapter, tests, and fallback plan become your responsibility.

Infrai fits when plain REST, one credential, and transparent readiness across providers remove enough undifferentiated work to justify a managed layer. The limitation is dependence on that managed control plane; choose LiteLLM when operating the gateway yourself is a requirement. It is also not suitable as an assumed answer for every adjacent media workflow. This design begins with an existing text transcript; current voice availability should be checked separately before treating live transcription or real-time voice as part of the same system. Dedicated moderation is a separate design concern rather than an assumed endpoint in this summarization path.

Ship the text path first. It solves the stated job with the fewest moving pieces, preserves an exit through a narrow adapter, and leaves retrieval out until customers actually ask questions across stored calls.

References

Top comments (0)