DEV Community

shashank ms
shashank ms

Posted on

Choosing the Best LLM Model for Coding Tasks with Advanced Reasoning

We are building an adaptive coding agent that classifies a programming task, routes it to the best Oxlo.ai model for the job, and validates the generated code before returning it. This saves money on simple tasks while guaranteeing deep reasoning for complex algorithms, all through a single OpenAI-compatible client.

What you'll need

  • Python 3.10 or newer.
  • An Oxlo.ai API key from https://portal.oxlo.ai.
  • The OpenAI SDK installed with pip install openai.

Step 1: Define the model registry

First, I set up the Oxlo.ai client and a registry that maps complexity tiers to model IDs. Using a registry keeps the routing logic clean and makes it easy to swap models later.

import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.oxlo.ai/v1",
    api_key=os.environ["OXLO_API_KEY"]
)

MODEL_REGISTRY = {
    "quick": "deepseek-v3.2",
    "standard": "qwen-3-32b",
    "deep": "deepseek-r1-671b",
    "agentic": "kimi-k2.6",
}

Step 2: Classify task complexity

Next, I use a fast, general-purpose model to classify the incoming task into a tier. This only costs a single request, and on Oxlo.ai that price is flat regardless of how long the prompt is.

def classify_task(prompt: str) -> str:
    """Classify the coding task into a complexity tier."""
    classifier_prompt = (
        "Classify the following coding task into exactly one category: "
        "quick, standard, deep, or agentic. "
        "Respond with only the single word.\n\nTask:\n" + prompt
    )
    response = client.chat.completions.create(
        model="llama-3.3-70b",
        messages=[{"role": "user", "content": classifier_prompt}],
        max_tokens=10,
    )
    tier = response.choices[0].message.content.strip().lower()
    return tier if tier in MODEL_REGISTRY else "standard"

Step 3: Write the system prompt

Here is the system prompt I use for every code generation call. It enforces reasoning, clean formatting, and a single fenced code block so extraction stays reliable.

SYSTEM_PROMPT = """You are an expert software engineer. Solve the user's coding task with advanced reasoning.

Follow these rules:
1. Analyze the problem and outline your reasoning before writing code.
2. Write clean, production-ready Python code.
3. Include docstrings and type hints where helpful.
4. Explain how you handle edge cases.
5. Output the code inside a single fenced python block.

Do not include installation instructions or external resource links."""

Step 4: Route and generate

Now I wire the router together. The function picks the model ID from the registry, calls the Oxlo.ai chat completions endpoint, and returns the raw markdown response.

import re

def generate_solution(task_prompt: str, tier: str) -> tuple[str, str]:
    """Route to the selected Oxlo.ai model and return the response."""
    model_id = MODEL_REGISTRY.get(tier, "llama-3.3-70b")
    response = client.chat.completions.create(
        model=model_id,
        messages=[
            {"role": "system", "content": SYSTEM_PROMPT},
            {"role": "user", "content": task_prompt},
        ],
        temperature=0.2,
    )
    return response.choices[0].message.content, model_id

Step 5: Critique and validate

Finally, I add a critique layer with a reasoning model and a syntax validator. If the critique finds issues, it returns corrected code; otherwise we keep the original.

def extract_python_code(text: str) -> str:
    """Extract the first fenced Python block."""
    match = re.search(r"

```python\n(.*?)```

", text, re.DOTALL)
    if match:
        return match.group(1).strip()
    return text.strip()

def critique_and_fix(raw_output: str, task_prompt: str) -> str:
    """Use a reasoning model to critique and fix the generated code."""
    critique_prompt = f"""Review this code for the following task.
If it is correct and optimal, reply with the word PASS.
If changes are needed, reply with the complete corrected code in a python block.

Task: {task_prompt}

Code:
{raw_output}
"""
    response = client.chat.completions.create(
        model="deepseek-r1-671b",
        messages=[
            {"role": "system", "content": "You are a senior code reviewer."},
            {"role": "user", "content": critique_prompt},
        ],
        temperature=0.1,
    )
    critique = response.choices[0].message.content
    if "PASS" in critique:
        return raw_output
    return critique

def syntax_check(code: str) -> bool:
    import ast
    try:
        ast.parse(code)
        return True
    except SyntaxError as e:
        print(f"Syntax error: {e}")
        return False

def run_pipeline(task: str):
    print(f"Task: {task}")
    tier = classify_task(task)
    print(f"Classified as: {tier}")
    
    raw_output, model_id = generate_solution(task, tier)
    print(f"Generated by: {model_id}")
    
    final_output = critique_and_fix(raw_output, task)
    code = extract_python_code(final_output)
    
    if syntax_check(code):
        print("Syntax OK")
        print("\n--- Final Code ---\n")
        print(code)
        return code
    else:
        print("Syntax check failed. Review output manually.")
        return final_output

Run it

Save the complete script as coding_agent.py, export your OXLO_API_KEY, and run it. Below is the expected output for a linear-time palindrome task.

if __name__ == "__main__":
    task = (
        "Write a Python function that finds the longest palindromic substring "
        "in linear time using Manacher's algorithm. Include a test assertion."
    )
    run_pipeline(task)

Expected output:

Task: Write a Python function that finds the longest palindromic substring in linear time using Manacher's algorithm. Include a test assertion.
Classified as: deep
Generated by: deepseek-r1-671b
Syntax OK

--- Final Code ---

def longest_palindromic_substring(s: str) -> str:
    """
    Returns the longest palindromic substring using Manacher's algorithm.
    """
    if not s:
        return ""
    
    t = '#'.join('^{}$'.format(s))
    n = len(t)
    p = [0] * n
    center = right = 0
    
    for i in range(1, n - 1):
        mirror = 2 * center - i
        if i < right:
            p[i] = min(right - i, p[mirror])
        
        while t[i + p[i] + 1] == t[i - p[i] - 1]:
            p[i] += 1
        
        if i + p[i] > right:
            center, right = i, i + p[i]
    
    max_len = max(p)
    center_index = p.index(max_len)
    start = (center_index - max_len) // 2
    return s[start:start + max_len]


if __name__ == "__main__":
    assert longest_palindromic_substring("babad") in ("bab", "aba")
    print("Test passed.")

Next steps

Try swapping the critique model to kimi-k2.6 for agentic review loops, or add an embedding retrieval step with Oxlo.ai's BGE-Large model to inject library documentation into the system prompt for domain-specific tasks.

Top comments (0)