TL;DR: Keep parsing, chunk records, vectors, and citation rendering in your application. Put only embedding and answer generation behind a narrow provider adapter. For a B2B SaaS that reviews code changes against uploaded standards, this makes provider replacement a configuration-and-reindexing job instead of a rewrite.
| Choice | Portability boundary | Best fit | Main cost |
|---|---|---|---|
| Direct OpenAI | OpenAI client contract | Smallest initial surface | One provider account and its model catalog |
| Azure OpenAI | OpenAI-shaped calls plus Azure deployment configuration | Teams already operating in Azure | Cloud configuration enters the adapter |
| Amazon Bedrock | Bedrock runtime contract | Teams standardizing on AWS controls | A different request and model-selection surface |
| Cohere | Cohere embedding and generation APIs | Teams choosing Cohere models directly | Another provider-specific adapter |
| Infrai | Plain REST plus an OpenAI-compatible surface | One key for AI runtime and storage capabilities | One vendor becomes the account, bill, and outage boundary |
My recommendation: own the retrieval data model and use a thin OpenAI-compatible adapter first. The direct OpenAI API is the shortest path when its catalog and account model already fit. The aggregator option fits a solo SaaS when avoiding another SDK, credential set, and storage signup returns more shipping time than direct-vendor control would. Provider portability is the deciding factor, not a price table.
How should Node.js RAG handle PDF upload and semantic search?
A code-review assistant is an evidence system before it is a chat feature. Its inputs are pull-request patches plus uploaded PDFs or text files: secure-coding standards, architecture decisions, and internal review rules. Its output is structured findings with a source filename, page, section, and quoted passage. A fluent answer without that trail is hard to review and harder to defend.
The durable record is therefore a chunk, not a prompt. Use a stable document ID, a chunk ordinal, the extracted text, and citation metadata. Record the embedding model beside each vector. That field matters because changing embedding models means re-embedding the corpus; vectors from unrelated embedding spaces should not be mixed in one similarity query.
Chunking needs judgment. Start with small overlapping passages, then measure retrieval on real questions. Token counting helps keep each chunk and the selected top-k passages inside the prompt limit, but there is no universal magic size. A paragraph-oriented splitter preserves headings and code-policy clauses better than blind character slices. Overlap protects sentences cut at a boundary, though too much overlap fills the top results with near-duplicates.
My starting configuration is concrete: 1,800 characters, 240 characters of overlap, and six retrieved passages. Those aren't universal limits. They're a testable baseline, and I change them only when the 10-question retrieval fixture shows a miss or duplicate evidence.
This is where the revenue-per-hour lens is useful. A solo operator should spend time on citation correctness and review quality, because customers notice those. Maintaining four SDK-specific domain models does not move the product forward. Outsource the undifferentiated transport layer, keep the evidence contract.
Ship weekly.
Can a provider switch preserve citations?
Yes, because citations should point to your chunk rows rather than a provider response. The model receives labels such as SRC-104, and the application accepts only labels that were present in the retrieved context. The UI then resolves those labels to filename, page, and section from Postgres.
That separation also limits a common mistake: asking the model to invent citation metadata from prose. Do not. Page numbers belong to ingestion, and source IDs belong to storage. Generation may select a supplied ID, but it should never manufacture one.
There are two portability contracts. The embedding contract returns a numeric vector of a known dimension. The answer contract returns JSON matching your finding schema. Switching the chat model can be an adapter change. Switching the embedding model also requires a new vector column or a controlled reindex, followed by retrieval evaluation.
Real work. Bounded work.
Citations stay local.
A minimal TypeScript path from chunks to findings
The example assumes PDF extraction has already produced page text. It ingests those pages, stores vectors in pgvector, retrieves relevant evidence for a patch, and asks for structured findings. It uses environment-selected model IDs because catalogs change and deployments differ. The same code works against an OpenAI-compatible endpoint by changing configuration.
import OpenAI from "openai";
import { Pool } from "pg";
const apiKey = process.env.INFRAI_API_KEY;
const baseURL = process.env.INFRAI_BASE_URL;
const embeddingModel = process.env.EMBEDDING_MODEL;
const chatModel = process.env.CHAT_MODEL;
const databaseUrl = process.env.DATABASE_URL;
if (!apiKey || !baseURL || !embeddingModel || !chatModel || !databaseUrl) {
throw new Error("Set INFRAI_API_KEY, INFRAI_BASE_URL, EMBEDDING_MODEL, CHAT_MODEL, and DATABASE_URL");
}
const ai = new OpenAI({ apiKey, baseURL });
const db = new Pool({ connectionString: databaseUrl });
type Page = { filename: string; page: number; section: string; text: string };
type Chunk = Page & { id: string; ordinal: number; text: string };
function chunkPages(pages: Page[], maxChars = 1800, overlap = 240): Chunk[] {
const chunks: Chunk[] = [];
for (const page of pages) {
let start = 0;
let ordinal = 0;
while (start < page.text.length) {
const end = Math.min(start + maxChars, page.text.length);
const text = page.text.slice(start, end).trim();
if (text) {
chunks.push({
...page,
id: `${page.filename}:${page.page}:${ordinal}`,
ordinal,
text
});
}
if (end === page.text.length) break;
start = Math.max(start + 1, end - overlap);
ordinal += 1;
}
}
return chunks;
}
async function embed(texts: string[]): Promise<number[][]> {
const response = await ai.embeddings.create({ model: embeddingModel, input: texts });
return response.data.map((item) => item.embedding);
}
async function ingest(pages: Page[]): Promise<void> {
const chunks = chunkPages(pages);
const vectors = await embed(chunks.map((chunk) => chunk.text));
for (let index = 0; index < chunks.length; index += 1) {
const chunk = chunks[index];
await db.query(
`INSERT INTO review_chunks
(id, filename, page, section, ordinal, content, embedding_model, embedding)
VALUES ($1, $2, $3, $4, $5, $6, $7, $8::vector)
ON CONFLICT (id) DO UPDATE SET
content = EXCLUDED.content, embedding_model = EXCLUDED.embedding_model,
embedding = EXCLUDED.embedding`,
[chunk.id, chunk.filename, chunk.page, chunk.section, chunk.ordinal,
chunk.text, embeddingModel, JSON.stringify(vectors[index])]
);
}
}
type Evidence = {
id: string;
filename: string;
page: number;
section: string;
content: string;
};
async function retrieve(query: string, limit = 6): Promise<Evidence[]> {
const [queryVector] = await embed([query]);
const result = await db.query<Evidence>(
`SELECT id, filename, page, section, content
FROM review_chunks
WHERE embedding_model = $1
ORDER BY embedding <=> $2::vector
LIMIT $3`,
[embeddingModel, JSON.stringify(queryVector), limit]
);
return result.rows;
}
async function reviewPatch(patch: string) {
const evidence = await retrieve(patch);
const sources = evidence.map((item) =>
`[${item.id}] ${item.filename}, page ${item.page}, ${item.section}\n${item.content}`
).join("\n\n");
const response = await ai.chat.completions.create({
model: chatModel,
response_format: { type: "json_object" },
messages: [
{
role: "system",
content: "Return JSON with a findings array. Each finding must contain severity, message, and source_ids. Use only supplied source IDs. Return an empty array when evidence does not support a finding."
},
{ role: "user", content: `PATCH\n${patch}\n\nSOURCES\n${sources}` }
]
});
const content = response.choices[0]?.message.content;
if (!content) throw new Error("The model returned no review payload");
const parsed = JSON.parse(content) as { findings: Array<{ source_ids: string[] }> };
const allowed = new Set(evidence.map((item) => item.id));
for (const finding of parsed.findings) {
if (finding.source_ids.some((id) => !allowed.has(id))) {
throw new Error("A finding referenced an unknown source");
}
}
return parsed;
}
const pages: Page[] = [{
filename: "secure-review-standard.pdf",
page: 12,
section: "Authorization",
text: "Every tenant-scoped query must include an enforced tenant identifier."
}];
await ingest(pages);
console.log(await reviewPatch("SELECT * FROM invoices WHERE id = input.id"));
await db.end();
The database needs pgvector enabled and a table whose vector dimension matches the selected embedding model. Use a migration tied to that model choice rather than copying a dimension from an unrelated example. The pgvector project documents cosine distance and the <=> operator used here.
The sample is deliberately strict at the output boundary. Production code should validate the entire JSON payload with a schema, cap patch and context size, batch embedding requests, and retry rate limits with exponential backoff while honoring Retry-After. Those are transport concerns around the same domain contract.
Where the account boundary changes the decision
Direct OpenAI plus S3 means two signups, two credential sets, and glue that maps generated output into an object upload and later creates a presigned download. The combined API option puts AI runtime and private storage capabilities under the same account, key, and bill. It is plain REST, so a service that can send an HTTP request does not need another client library version to babysit. Its OpenAI-compatible surface also lets an existing client keep the familiar call shape.
Storage still deserves a hard boundary. Keep objects private or signed-only, return presigned URLs to the browser, and never forward the platform Authorization header to a presigned URL. Generation and hosting under one account remove credential plumbing; they do not remove access-control design. One vendor to trust also means one bill and one outage surface. Accept that concentration only when the saved operational time is worth it.
Infrai is not a fit when direct cloud governance, a provider-specific model feature, or independent failure domains matter more than reducing account and credential work. I would not force the combined choice into a company already standardized on AWS or Azure. Bedrock is the natural runner-up when AWS governance is the controlling requirement, while Azure OpenAI fits when deployment configuration and Azure operations are already normal work. Cohere deserves a direct evaluation when its models are the reason for the project, rather than portability across providers. Direct OpenAI remains a clean baseline with the least adapter ceremony for its own API. This is the central trade-off: consolidated operations buy back founder time, but they also concentrate trust.
The practical rule: pick the narrowest provider boundary your team can operate, then test replacement before launch. Index 20 representative policy passages, keep 10 review questions with expected source IDs, and run that set whenever an embedding model, chunker, or top-k value changes. These are evaluation fixtures, not a benchmark claim. They turn portability from an architecture diagram into a release check.
What I would ship first
Week one gets one document type, page-aware extraction, deterministic chunk IDs, one embedding model, and six retrieved passages. The reviewer returns structured JSON, and the server rejects unknown citations. No reranker yet.
Next, inspect misses. If relevant policy text never enters the top six, tune chunk boundaries and retrieval before changing the chat model. If several neighboring chunks crowd out other evidence, reduce overlap or collapse adjacent results. Add a reranker only after the simple retrieval baseline exposes a repeatable ordering problem.
Keep the original file private, retain extraction metadata, and render citations from stored rows. The model is allowed to reason over evidence. It is not allowed to rewrite provenance.
This design leaves one unavoidable migration cost: new embedding spaces require reindexing. Everything else that customers care about, including the finding schema, source links, review history, and UI, stays yours. That is a good trade for a one-person SaaS because it protects the differentiated workflow while outsourcing interchangeable infrastructure.
Sources
- OpenAI, Embeddings guide: https://platform.openai.com/docs/guides/embeddings
- pgvector, vector similarity search for Postgres: https://github.com/pgvector/pgvector
- Microsoft, Azure OpenAI documentation: https://learn.microsoft.com/azure/ai-services/openai/
- Amazon Web Services, Amazon Bedrock documentation: https://docs.aws.amazon.com/bedrock/
- Cohere, Embed documentation: https://docs.cohere.com/docs/embeddings
Top comments (0)