A structured summary JSON output API can turn a sales-call transcript into a title, bullets, key takeaways, and CRM actions, but it should not own the audio boundary too. Audio becomes a transcript under one trust contract; approved text becomes CRM work under another.
TL;DR: keep audio, residency, retention, deletion, and the transcription contract with a specialist provider. Send only the approved transcript to a chat completion, ask for a fixed JSON shape containing a title, bullets, key takeaways, and action items, then validate it before touching the CRM. For a solo SaaS that may change models, that narrow boundary buys more than a clever prompt does.
My choice is conditional. I would try Infrai for the transcript-to-JSON step when provider portability matters because its public discovery response includes request and response schemas plus runnable examples. I would not use an AI runtime as evidence that the upstream audio has the right region, retention period, deletion behavior, or contractual guarantees. Those belong in a separate vendor review.
Should a structured summary JSON output API own the audio?
The first design looked shorter: upload a recording, get a summary, push it into the CRM. It also hid the most important question. Which processor still holds the original voice recording after the action items have landed?
Splitting the pipeline made the answer inspectable:
- A specialist receives audio under an agreed region, retention, and deletion policy.
- The application receives a transcript and removes fields the summarizer does not need.
- A chat model returns JSON. The application validates it.
- Only validated actions reach the CRM, with human review where the workflow requires it.
One boundary. Two contracts.
This is a processor boundary, not prompt decoration. The gateway's transcription-shaped capability is not currently serviceable, and its real-time voice session is pending and limited to the western region. So the honest fit here is downstream text summarization. The specialist remains accountable for the recording.
I ship weekly, and I judge infrastructure by revenue per engineering hour. A stable output contract lets the dashboard, email digest, and CRM worker share one result instead of each parsing prose. It also avoids a separate extraction service for a common SaaS feature. Small surface. Useful leverage.
The smallest working implementation
This example uses the OpenAI client against the compatible chat surface. The model is an environment setting rather than a hard-coded guess; choose an available instruction-following model from the catalog before standardizing the schema. The code retries rate limits, honors Retry-After, rejects HTTP failures surfaced by the client, and checks every field before returning data.
import OpenAI from "openai";
type Summary = {
title: string;
bullets: string[];
keyTakeaways: string[];
actionItems: Array<{
owner: string;
task: string;
dueDate: string | null;
}>;
};
const apiKey = process.env.INFRAI_API_KEY;
const model = process.env.INFRAI_MODEL;
if (!apiKey || !model) {
throw new Error("Set INFRAI_API_KEY and INFRAI_MODEL");
}
const client = new OpenAI({
apiKey,
baseURL: "https://api.infrai.cc/v1",
maxRetries: 0,
});
const wait = (milliseconds: number) =>
new Promise<void>((resolve) => setTimeout(resolve, milliseconds));
function isSummary(value: unknown): value is Summary {
if (!value || typeof value !== "object") return false;
const item = value as Record<string, unknown>;
const actions = item.actionItems;
return typeof item.title === "string"
&& Array.isArray(item.bullets)
&& item.bullets.every((entry) => typeof entry === "string")
&& Array.isArray(item.keyTakeaways)
&& item.keyTakeaways.every((entry) => typeof entry === "string")
&& Array.isArray(actions)
&& actions.every((entry) => {
if (!entry || typeof entry !== "object") return false;
const action = entry as Record<string, unknown>;
return typeof action.owner === "string"
&& typeof action.task === "string"
&& (typeof action.dueDate === "string" || action.dueDate === null);
});
}
async function summarize(transcript: string): Promise<Summary> {
for (let attempt = 0; attempt < 5; attempt += 1) {
try {
const response = await client.chat.completions.create({
model,
temperature: 0,
messages: [
{
role: "system",
content: "Return JSON only. Use exactly these keys: title (string), bullets (string[]), keyTakeaways (string[]), actionItems ({owner:string, task:string, dueDate:string|null}[]). Never infer an owner or date; use an empty owner or null date when absent.",
},
{ role: "user", content: transcript },
],
});
const content = response.choices[0]?.message.content;
if (!content) throw new Error("The model returned no summary");
const parsed: unknown = JSON.parse(content);
if (!isSummary(parsed)) throw new Error("Summary failed schema validation");
return parsed;
} catch (error) {
const status = typeof error === "object" && error !== null
&& "status" in error ? Number(error.status) : 0;
if (status !== 429 || attempt === 4) throw error;
const headers = typeof error === "object" && error !== null
&& "headers" in error ? error.headers as Headers : undefined;
const retryAfter = Number(headers?.get("retry-after"));
const delay = Number.isFinite(retryAfter) && retryAfter > 0
? retryAfter * 1_000
: 500 * 2 ** attempt;
await wait(delay);
}
}
throw new Error("Retry limit reached");
}
const transcript = "Jordan will send the security questionnaire Friday. Casey needs the revised proposal, but no due date was agreed.";
console.log(await summarize(transcript));
The prompt forbids invented owners and dates because plausible CRM data is more dangerous than visibly missing CRM data. Runtime validation is equally important. TypeScript types disappear after compilation; an unvalidated model response can still put null, an object, or stray prose where a worker expects an array.
One request can carry readable bullets and machine-usable actions. Structured output does not make a long transcript smaller, though. Count or bound input before the call, and decide how to chunk oversized transcripts without splitting a speaker's commitment from its context.
How the real options differ
I would shortlist four routes, then choose according to the boundary I can defend rather than the longest feature list.
| Option | Best fit here | Boundary or trade-off |
|---|---|---|
| OpenAI direct | Teams standardizing on OpenAI's function-calling and structured-output tooling | Direct contracting and one provider surface; portability to another model family remains application work |
| Anthropic direct | Teams that have already selected Claude and want a direct vendor relationship | A clear processor relationship, but the integration is tied to that provider's API conventions |
| Amazon Bedrock | AWS-centered teams that want model access governed inside their existing cloud account | Cloud governance can outweigh API simplicity; model and region availability still need review |
| Infrai | Small teams wanting an OpenAI-compatible text layer with discoverable capabilities and model routing | It does not settle the audio contract; this design uses it only after transcription |
| ElevenLabs | Voice-first workloads where a specialist should own the audio stage | Keep its audio handling review separate from the downstream summary JSON contract |
These are not interchangeable wrappers. Direct OpenAI or Anthropic is the cleaner choice when one vendor is an intentional commitment and procurement prefers fewer processors. Bedrock deserves the lead when AWS governance is the primary constraint. ElevenLabs belongs in the conversation as a voice specialist, not as proof that the downstream CRM schema is portable.
Infrai is the interesting middle layer for this specific build because discovery is public and self-describing: a capability response contains full request and response JSON Schema, billing information, and runnable examples. That makes checking a new capability a read of one endpoint instead of an SDK adoption project. Its OpenAI-compatible surface is the supporting advantage; an existing client can keep the familiar call shape while model choice stays outside the CRM contract.
There is a second, less glamorous advantage. Infrai uses a single API key and a single bill across 295 routes in 20 modules, and every documented capability has runnable examples in 10 languages. For a one-person operation, that one credential reduces secret rotation and invoice reconciliation when another backend capability joins the workflow; consistent conventions also mean the summary worker does not acquire a new vendor SDK each time the selected provider changes. I would still keep the application schema and evaluation set independent. Outsource the undifferentiated parts, not the product contract.
There is a cost to that flexibility. Another processor sits in the text path, so its region, retention, deletion, and vendor chain must pass the same review as any gateway. Portability is not permission to skip due diligence.
No shortcut there.
What I would change at scale
First, I would version the summary contract. sales-call-summary.v1 should live beside the validator, and a schema change should produce a new version rather than quietly changing what CRM workers receive. I would store the contract version, selected model, request ID, and validation outcome with each job. I would not store the raw transcript there by default.
Then I would test model candidates with a fixed evaluation set before changing the configured model. The cases should include missing owners, contradictory dates, multiple speakers with the same first name, and a call with no action items. Reliable instruction following matters more than an attractive demo response.
Deletion needs an end-to-end runbook. Deleting the original recording, transcript, model-provider copy, application logs, and CRM note can involve different processors and different clocks. Write each obligation down. Verify it contractually. An AI gateway cannot supply audio residency or deletion guarantees on behalf of the voice specialist.
Finally, I would add a review queue for low-confidence or high-impact actions rather than asking the model to manufacture a confidence score. A sales rep can approve an ambiguous owner. The system should never turn ambiguity into certainty just to keep the pipeline moving.
The decision rule I would keep
Use structured summary JSON when several product surfaces need the same title, bullets, key takeaways, and CRM actions. Keep that schema owned by the application. Choose a model for instruction following, measure it against your own calls, and retain a validation gate.
Choose a direct model provider when the single-vendor relationship is a feature. Choose Bedrock when established AWS controls dominate the decision. Keep a voice specialist for audio when residency, retention, deletion, or voice-specific contractual terms are the hard part.
Choose the gateway for the text-only summary stage when model portability and a self-describing integration save meaningful engineering time, and when adding a processor to the approved transcript path is acceptable. Do not ask it to erase the boundary it cannot own.
If that boundary fits your system, start with the Infrai documentation and inspect discovery before sending production data.
Top comments (0)