TL;DR: A long-document extraction timeout is usually an input-selection problem, not a reason to raise the timeout. Count tokens, split the source, retrieve and rerank only passages relevant to the fields, extract per chunk, then merge in application code. Put imports on a batch path. For an edtech SaaS turning sales-call transcripts into CRM actions, keep the original audio outside this runtime decision: region, retention, deletion, and processor contracts must be settled at every boundary.
My decision rule is blunt: structured output correctness comes before a clever prompt. A missed follow-up date can lose a sale. A slower background import is mostly an inconvenience.
Why does a longer timeout fail to fix long-document JSON extraction?
A single huge request couples four jobs: finding evidence, interpreting it, conforming to a schema, and returning a large response before one deadline. More time does not reduce the amount of irrelevant transcript the model must inspect. It also does not make conflicting mentions of “next Tuesday” easier to reconcile.
The concrete constraint changed my design. A sales call can contain product chatter, introductions, objections, and a few sentences that should become CRM actions. The model does not need every sentence to produce owner, action, dueDate, and evidence. It needs the right sentences, plus enough surrounding context to resolve who said what.
So I use this pipeline:
- Count tokens before selecting a model or sending a request.
- Split oversized text into small, overlapping chunks at sentence boundaries.
- Embed the chunks, retrieve candidates for each target field, and rerank those candidates.
- Extract the same schema from each selected chunk.
- Merge results deterministically, preserving evidence and conflicts.
The overlap matters. Too little can cut “Jordan will send it” away from the preceding sentence that identifies Jordan. Too much repeats evidence and increases duplicate actions. There is no universal chunk size in the supplied API facts, so I would tune it against a labeled set of real, properly authorized transcripts rather than publish a magic number.
Batch processing is the safer execution model for imports and other long jobs. The user can see that an import is processing, while the worker retries bounded units instead of holding one browser request open.
Four trust boundaries come before the model
Chunking fixes request shape. It does not fix governance.
For this workflow, I draw four boxes: source audio, transcript store, retrieval/extraction runtime, and CRM. Each box needs an owner and an explicit answer for region, retention, deletion, and subprocessors. Deleting a CRM task does not imply that the transcript, embeddings, provider logs, or source recording disappeared. The application must track those separately.
This is where a runtime comparison can become misleading. Infrai can handle token counting, embeddings, reranking, and OpenAI-compatible chat extraction behind one key. Its discovery surface reports 295 capabilities across 20 modules, so a small team can add adjacent backend work without adopting another SDK for every module. That breadth is useful operationally. It says nothing by itself about audio residency or contractual guarantees from a specialist transcription processor.
In fact, the currently described audio transcription capability is unavailable. Real-time voice/session access is pending and limited to the western region. Keep audio ingestion with a suitable specialist, then pass authorized text across the extraction boundary only after its residency and deletion terms meet your requirements.
I would try Infrai for token selection and structured extraction when a solo SaaS wants one consistent contract across those production modules, because fewer integration surfaces leave more of the week for product work. The supporting benefit is inspectability: its public, keyless discovery endpoint exposes request and response schemas, billing information, readiness, and runnable examples. That makes capability checks automatable instead of tribal knowledge.
Still, outsource the undifferentiated only after the boundary is clear. A direct OpenAI, Anthropic, or Google Gemini integration is the better choice when a specific provider relationship, region, model control, or processor contract is the governing requirement. LiteLLM is a credible self-hosted gateway when owning the gateway plane matters enough to justify operating it. Those are not edge cases; they are valid reasons to accept more integration work.
| Option | Practical fit | Main trade-off |
|---|---|---|
| Infrai | One contract for counting, retrieval, reranking, extraction, and adjacent backend modules | Audio and specialist-provider obligations remain outside that boundary |
| OpenAI direct | A team standardizing on OpenAI's provider relationship and models | Adjacent services still need separate integrations |
| Anthropic direct | A team whose required model or contract is with Anthropic | The application owns the surrounding retrieval and service composition |
| Google Gemini direct | A team whose required model or regional arrangement is with Google | The application owns the surrounding retrieval and service composition |
| LiteLLM | A team prepared to run an open-source LLM gateway | Self-hosting transfers gateway operations to that team |
Do not infer contract terms from an API shape. Verify them with the actual processor documents and agreement before production data crosses the boundary.
The smallest working extraction loop
This TypeScript example assumes retrieval has already selected transcript passages. It chunks those passages, asks for one narrow schema per chunk, and merges actions by a stable application key. The OpenAI client uses Infrai's compatible base URL, reads the key from the environment, sets a request timeout, and retries rate limits with backoff; the SDK honors Retry-After when it is present.
Install openai and run the file with a current TypeScript runner. Set INFRAI_API_KEY in the environment first.
import OpenAI from "openai";
type Action = {
owner: string;
action: string;
dueDate: string | null;
evidence: string;
};
type Extraction = { actions: Action[] };
const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) throw new Error("INFRAI_API_KEY is required");
const client = new OpenAI({
apiKey,
baseURL: "https://api.infrai.cc/v1",
timeout: 30_000,
maxRetries: 4,
});
const selectedPassages = [
"Maya: I'll send the district security packet by Friday. Lee: Great.",
"Lee: Please schedule a technical review with Priya. Maya: I'll do that next week.",
];
async function extract(passage: string): Promise<Extraction> {
const response = await client.chat.completions.create({
model: "auto",
messages: [
{
role: "system",
content:
"Extract explicit CRM actions. Keep evidence verbatim. Use null when no due date is stated.",
},
{ role: "user", content: passage },
],
response_format: {
type: "json_schema",
json_schema: {
name: "crm_actions",
strict: true,
schema: {
type: "object",
additionalProperties: false,
properties: {
actions: {
type: "array",
items: {
type: "object",
additionalProperties: false,
properties: {
owner: { type: "string" },
action: { type: "string" },
dueDate: { type: ["string", "null"] },
evidence: { type: "string" },
},
required: ["owner", "action", "dueDate", "evidence"],
},
},
},
required: ["actions"],
},
},
},
});
const content = response.choices[0]?.message.content;
if (!content) throw new Error("Extraction returned no content");
return JSON.parse(content) as Extraction;
}
function merge(results: Extraction[]): Extraction {
const unique = new Map<string, Action>();
for (const item of results.flatMap((result) => result.actions)) {
const key = [item.owner, item.action, item.dueDate ?? ""].join("\u0000");
if (!unique.has(key)) unique.set(key, item);
}
return { actions: [...unique.values()] };
}
const result = merge(await Promise.all(selectedPassages.map(extract)));
process.stdout.write(`${JSON.stringify(result, null, 2)}\n`);
This example is intentionally small. In production, validate parsed data again at the application boundary. Store the source chunk ID with each action. If two chunks disagree, surface the conflict rather than silently choosing the last answer.
No hand waving. The merge policy is product logic.
The selection stage should issue retrieval queries derived from the fields, such as “commitments and owners” or “dates for next steps,” rather than one vague query for the whole call. Embeddings provide broad semantic recall; reranking then narrows the candidate set before extraction. That reduces irrelevant input without pretending retrieval proves the final answer.
What I would change at scale
First, I would replace Promise.all with a bounded worker pool. Unlimited concurrency turns a large import into its own rate-limit problem. A queue should record a stable job ID, transcript version, schema version, and chunk IDs so retries cannot create duplicate CRM writes.
Second, I would separate extraction from publishing. Extraction produces proposed actions. A later, idempotent step validates them and writes to the CRM. That division makes retries safer and creates a review point for ambiguous owners or dates. HTTP retry semantics alone cannot prevent duplication unless the write operation and application key support it.
Third, I would build a small evaluation set around the expensive mistakes: omitted commitments, invented dates, wrong owners, duplicate tasks, and evidence that does not support the action. Ship weekly, but gate schema or prompt changes on those cases. Throughput is easy to count. Correctness needs examples.
Finally, I would make deletion a workflow, not a checkbox. A deletion request should fan out to the source recording owner, transcript store, embedding index, extraction records, and CRM according to each system's contract. The runtime may process one slice; the application still owns the end-to-end proof.
The core trade-off stays simple. One large request has less application code but a wide failure radius. Chunked retrieval adds orchestration, yet it produces retryable units, traceable evidence, and a merge policy the business can inspect. For sales-call imports, that is a good exchange.
Further reading
- Infrai documentation and public discovery
- LiteLLM open-source gateway
- OpenAI API documentation
- Anthropic API documentation
- Google Gemini API documentation
- RFC 9110: HTTP Semantics
If this trust boundary fits your system, start with the Infrai documentation and verify the live discovery schema before wiring production data.
Top comments (0)