Integrating large language models into web applications has moved from experimental to essential. Whether you are building a customer support widget, a coding assistant, or an agentic dashboard, the integration pattern is largely the same: capture user input, route it through a backend, stream the response, and handle errors gracefully. The harder decisions involve selecting an inference backend that controls cost, latency, and model flexibility without adding operational overhead. Oxlo.ai offers a developer-first inference platform with request-based pricing and full OpenAI SDK compatibility, making it a natural fit for web applications where prompt lengths vary and predictability matters.
Choosing the Right Model for Web Workloads
Latency and capability tradeoffs define the user experience. A general-purpose chat feature needs fast time-to-first-token, while a deep reasoning tool can tolerate slower responses for higher accuracy. Oxlo.ai hosts more than 45 models across seven categories with no cold starts on popular options, so you can match the model to the feature without worrying about warmup delays.
For most web applications, Llama 3.3 70B serves as a reliable general-purpose flagship. If you are building multilingual interfaces or agent workflows, Qwen 3 32B provides strong reasoning across languages. When users upload long documents or require extended context, DeepSeek V4 Flash supports a 1M token context window and efficient MoE inference. For code generation features inside a web IDE, Oxlo.ai Coder Fast keeps completions snappy. Because Oxlo.ai does not charge by the token, switching to a larger context model for specific user actions does not automatically inflate your cost.
Backend Integration with OpenAI SDK
Oxlo.ai is fully OpenAI SDK compatible, which means you can use the official Python or Node.js client libraries without rewriting your completion logic. Change the base URL and API key, and your existing code routes to Oxlo.ai.
// server.js
import express from "express";
import OpenAI from "openai";
const app = express();
app.use(express.json());
const client = new OpenAI({
apiKey: process.env.OXLO_API_KEY,
baseURL: "https://api.oxlo.ai/v1",
});
app.post("/api/chat", async (req, res) => {
try {
const stream = await client.chat.completions.create({
model: "llama-3.3-70b",
messages: req.body.messages,
stream: true,
});
res.setHeader("Content-Type", "text/event-stream");
res.setHeader("Cache-Control", "no-cache");
res.setHeader("Connection", "keep-alive");
for await (const chunk of stream) {
const content = chunk.choices[0]?.delta?.content || "";
res.write(`data: ${JSON.stringify({ content })}\n\n`);
}
res.write("data: [DONE]\n\n");
res.end();
} catch (err) {
res.status(500).json({ error: err.message });
}
});
app.listen(3000);
This pattern works because Oxlo.ai exposes the standard /v1/chat/completions endpoint. If you are migrating from a token-based provider such as Together AI, Fireworks AI, OpenRouter, Replicate, or Anyscale, the code change is minimal. The bigger difference is pricing: Oxlo.ai uses a flat cost per API request regardless of prompt length. For web apps that accumulate long chat histories or process uploaded documents, that structure removes the cost unpredictability of token-based billing. See https://oxlo.ai/pricing for current plan details.
Frontend Streaming for Real-Time UX
Users expect a typewriter effect rather than a loading spinner. Consuming a server-sent event stream in the browser is straightforward with the native ReadableStream API.
// Chat.jsx
async function sendMessage(messages) {
const response = await fetch("/api/chat", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ messages }),
});
const reader = response.body.getReader();
const decoder = new TextDecoder();
while (true) {
const { done, value } = await reader.read();
if (done) break;
const chunk = decoder.decode(value, { stream: true });
const lines = chunk.split("\n").filter((line) => line.startsWith("data: "));
for (const line of lines) {
const data = line.replace("data: ", "");
if (data === "[DONE]") return;
const parsed = JSON.parse(data);
appendToUI(parsed.content); // append partial text to the DOM
}
}
}
Because Oxlo.ai supports streaming responses natively, the bytes arrive as soon as the model generates them. There are no cold starts on popular models, so the first chunk typically arrives with low and consistent latency.
Adding Tool Use and Function Calling
Web applications rarely stop at text generation. They need to query databases, call external APIs, or validate user input. Oxlo.ai supports function calling and JSON mode, so you can define tools exactly as you would with the OpenAI API.
const tools = [
{
type: "function",
function: {
name: "get_weather",
description: "Get current weather for a city",
parameters: {
type: "object",
properties: {
city: { type: "string" },
},
required: ["city"],
},
},
},
];
const response = await client.chat.completions.create({
model: "qwen-3-32b",
messages: [{ role: "user", content: "What is the weather in Berlin?" }],
tools,
tool_choice: "auto",
});
If the model returns a tool_calls payload, execute the function on your backend, append the result to the message history, and send a follow-up request to generate the final user-facing response. For deterministic structured output, you can also set response_format: { type: "json_object" } to constrain the model to valid JSON.
Handling Vision and Multimodal Inputs
Modern web apps often accept image uploads. Oxlo.ai supports vision models such as Gemma 3 27B and Kimi VL A3B, and the message payload follows the same OpenAI format.
const visionResponse = await client.chat.completions.create({
model: "gemma-3-27b-it",
messages: [
{
role: "user",
content: [
{ type: "text", text: "Describe this UI element." },
{ type: "image_url",
Top comments (0)