Building a production-grade chatbot requires more than calling a completions endpoint. You need to manage conversation state, handle tool use, and control costs as context windows grow. This tutorial walks through a complete implementation using the OpenAI SDK with Oxlo.ai as the inference backend, covering streaming, memory management, function calling, and vision inputs.
Architecture of an LLM Chatbot
A minimal chatbot has three layers: the interface, the state manager, and the inference backend. The interface accepts user input. The state manager maintains conversation history and metadata. The inference backend generates responses. For this tutorial, we will use Python with the OpenAI SDK, storing history in memory and routing requests to Oxlo.ai.
Selecting a Model
Your model choice determines reasoning quality, latency, and cost structure. Oxlo.ai hosts 45+ open-source and proprietary models across seven categories, all accessible through a single OpenAI-compatible endpoint. For general conversational agents, Llama 3.3 70B offers strong instruction following. For multilingual or agentic workflows, Qwen 3 32B is a solid candidate. If your chatbot handles deep reasoning or complex coding, DeepSeek R1 671B MoE or Kimi K2.6 provide advanced chain-of-thought capabilities. Because Oxlo.ai uses request-based pricing, long system prompts and extended conversation history do not inflate your bill the way token-based metering does. See https://oxlo.ai/pricing for plan details.
Environment Setup
Install the OpenAI SDK and set your Oxlo.ai API key.
pip install openai
import openai
client = openai.OpenAI(
base_url="https://api.oxlo.ai/v1",
api_key="YOUR_OXLO_API_KEY"
)
No custom client logic is required. Oxlo.ai is a drop-in replacement for the standard OpenAI client.
The Chat Loop with Streaming
A robust chat loop streams tokens to the user while appending them to the conversation history.
def chat_loop():
messages = [
{"role": "system", "content": "You are a helpful assistant. Answer concisely."}
]
while True:
user_input = input("User: ")
if user_input.lower() in ["exit", "quit"]:
break
messages.append({"role": "user", "content": user_input})
stream = client.chat.completions.create(
model="llama-3.3-70b",
messages=messages,
stream=True
)
print("Assistant: ", end="", flush=True)
response_text = ""
for chunk in stream:
if chunk.choices[0].delta.content:
token = chunk.choices[0].delta.content
print(token, end="", flush=True)
response_text += token
print()
messages.append({"role": "assistant", "content": response_text})
if __name__ == "__main__":
chat_loop()
Streaming improves perceived latency and keeps the interface responsive.
Managing Conversation Memory
As conversations grow, you need a strategy to stay within context limits and manage cost. On token-based platforms, every additional message increases the price. With Oxlo.ai, request-based pricing means the cost per turn stays flat regardless of how much history you include, so you can prioritize quality over aggressive truncation. Still, you should implement sliding-window or summarization memory to respect model context windows. Here is a sliding-window approach:
MAX_HISTORY = 20
def trim_messages(messages, max_turns=MAX_HISTORY):
# Keep system prompt and last N turns
if len(messages) <= max_turns + 1:
return messages
return [messages[0]] + messages[-max_turns:]
Call trim_messages before each API request.
Adding Tool Use
Modern chatbots call external APIs. Oxlo.ai supports function calling on compatible models. Define your tools in OpenAI schema format.
tools = [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get current weather for a location",
"parameters": {
"type": "object",
"properties": {
"location": {"type": "string"}
},
"required": ["location"]
}
}
}
]
response = client.chat.completions.create(
model="qwen3-32b",
messages=messages,
tools=tools,
tool_choice="auto"
)
Inspect response.choices[0].message.tool_calls, execute the function, and append the result back into messages with role "tool". This pattern creates agentic workflows that remain cost-predictable on Oxlo.ai because each round-trip is one request, not a variable token charge.
Vision-Enabled Chatbots
If your chatbot processes images, models like Kimi K2.6 or Gemma 3 27B accept vision inputs through the same chat/completions endpoint.
messages.append({
"role": "user",
"content": [
{"type": "text", "text": "Describe this image."},
{"type": "image_url", "image_url": {"url": "https://example.com/image.png"}}
]
})
The integration path is identical to OpenAI, so existing vision code ports without changes.
Deployment and Cost Considerations
When moving from prototype to production, cost predictability matters. Token-based providers scale charges with prompt length, so a chatbot with long system instructions or extensive tool context becomes expensive quickly. Oxlo.ai charges one flat cost per request, making it significantly cheaper for long-context and agentic workloads where prompts grow over time. You can run large contexts through Llama 3.3 70B or DeepSeek V4 Flash with 1M context without watching token meters accumulate. For pricing and request limits, visit https://oxlo.ai/pricing.
Conclusion
Building an LLM chatbot comes down to managing state, streaming output, and selecting the right model for the task. With Oxlo.ai, you get OpenAI SDK compatibility across 45+ models, function calling, vision support, and request-based pricing that stays flat as your context grows. Swap your base_url to https://api.oxlo.ai/v1, pick a model, and deploy.
Top comments (0)