We are building an adaptive coding agent that classifies a programming task, routes it to the best Oxlo.ai model for the job, and validates the generated code before returning it. This saves money on simple tasks while guaranteeing deep reasoning for complex algorithms, all through a single OpenAI-compatible client.
What you'll need
- Python 3.10 or newer.
- An Oxlo.ai API key from https://portal.oxlo.ai.
- The OpenAI SDK installed with
pip install openai.
Step 1: Define the model registry
First, I set up the Oxlo.ai client and a registry that maps complexity tiers to model IDs. Using a registry keeps the routing logic clean and makes it easy to swap models later.
import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.oxlo.ai/v1",
api_key=os.environ["OXLO_API_KEY"]
)
MODEL_REGISTRY = {
"quick": "deepseek-v3.2",
"standard": "qwen-3-32b",
"deep": "deepseek-r1-671b",
"agentic": "kimi-k2.6",
}
Step 2: Classify task complexity
Next, I use a fast, general-purpose model to classify the incoming task into a tier. This only costs a single request, and on Oxlo.ai that price is flat regardless of how long the prompt is.
def classify_task(prompt: str) -> str:
"""Classify the coding task into a complexity tier."""
classifier_prompt = (
"Classify the following coding task into exactly one category: "
"quick, standard, deep, or agentic. "
"Respond with only the single word.\n\nTask:\n" + prompt
)
response = client.chat.completions.create(
model="llama-3.3-70b",
messages=[{"role": "user", "content": classifier_prompt}],
max_tokens=10,
)
tier = response.choices[0].message.content.strip().lower()
return tier if tier in MODEL_REGISTRY else "standard"
Step 3: Write the system prompt
Here is the system prompt I use for every code generation call. It enforces reasoning, clean formatting, and a single fenced code block so extraction stays reliable.
SYSTEM_PROMPT = """You are an expert software engineer. Solve the user's coding task with advanced reasoning.
Follow these rules:
1. Analyze the problem and outline your reasoning before writing code.
2. Write clean, production-ready Python code.
3. Include docstrings and type hints where helpful.
4. Explain how you handle edge cases.
5. Output the code inside a single fenced python block.
Do not include installation instructions or external resource links."""
Step 4: Route and generate
Now I wire the router together. The function picks the model ID from the registry, calls the Oxlo.ai chat completions endpoint, and returns the raw markdown response.
import re
def generate_solution(task_prompt: str, tier: str) -> tuple[str, str]:
"""Route to the selected Oxlo.ai model and return the response."""
model_id = MODEL_REGISTRY.get(tier, "llama-3.3-70b")
response = client.chat.completions.create(
model=model_id,
messages=[
{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": task_prompt},
],
temperature=0.2,
)
return response.choices[0].message.content, model_id
Step 5: Critique and validate
Finally, I add a critique layer with a reasoning model and a syntax validator. If the critique finds issues, it returns corrected code; otherwise we keep the original.
def extract_python_code(text: str) -> str:
"""Extract the first fenced Python block."""
match = re.search(r"
```python\n(.*?)```
", text, re.DOTALL)
if match:
return match.group(1).strip()
return text.strip()
def critique_and_fix(raw_output: str, task_prompt: str) -> str:
"""Use a reasoning model to critique and fix the generated code."""
critique_prompt = f"""Review this code for the following task.
If it is correct and optimal, reply with the word PASS.
If changes are needed, reply with the complete corrected code in a python block.
Task: {task_prompt}
Code:
{raw_output}
"""
response = client.chat.completions.create(
model="deepseek-r1-671b",
messages=[
{"role": "system", "content": "You are a senior code reviewer."},
{"role": "user", "content": critique_prompt},
],
temperature=0.1,
)
critique = response.choices[0].message.content
if "PASS" in critique:
return raw_output
return critique
def syntax_check(code: str) -> bool:
import ast
try:
ast.parse(code)
return True
except SyntaxError as e:
print(f"Syntax error: {e}")
return False
def run_pipeline(task: str):
print(f"Task: {task}")
tier = classify_task(task)
print(f"Classified as: {tier}")
raw_output, model_id = generate_solution(task, tier)
print(f"Generated by: {model_id}")
final_output = critique_and_fix(raw_output, task)
code = extract_python_code(final_output)
if syntax_check(code):
print("Syntax OK")
print("\n--- Final Code ---\n")
print(code)
return code
else:
print("Syntax check failed. Review output manually.")
return final_output
Run it
Save the complete script as coding_agent.py, export your OXLO_API_KEY, and run it. Below is the expected output for a linear-time palindrome task.
if __name__ == "__main__":
task = (
"Write a Python function that finds the longest palindromic substring "
"in linear time using Manacher's algorithm. Include a test assertion."
)
run_pipeline(task)
Expected output:
Task: Write a Python function that finds the longest palindromic substring in linear time using Manacher's algorithm. Include a test assertion.
Classified as: deep
Generated by: deepseek-r1-671b
Syntax OK
--- Final Code ---
def longest_palindromic_substring(s: str) -> str:
"""
Returns the longest palindromic substring using Manacher's algorithm.
"""
if not s:
return ""
t = '#'.join('^{}$'.format(s))
n = len(t)
p = [0] * n
center = right = 0
for i in range(1, n - 1):
mirror = 2 * center - i
if i < right:
p[i] = min(right - i, p[mirror])
while t[i + p[i] + 1] == t[i - p[i] - 1]:
p[i] += 1
if i + p[i] > right:
center, right = i, i + p[i]
max_len = max(p)
center_index = p.index(max_len)
start = (center_index - max_len) // 2
return s[start:start + max_len]
if __name__ == "__main__":
assert longest_palindromic_substring("babad") in ("bab", "aba")
print("Test passed.")
Next steps
Try swapping the critique model to kimi-k2.6 for agentic review loops, or add an embedding retrieval step with Oxlo.ai's BGE-Large model to inject library documentation into the system prompt for domain-specific tasks.
Top comments (0)