DEV Community

Sergey Inozemtsev
Sergey Inozemtsev Subscriber

Posted on

LLM function calling in Python across 9 providers without rewriting your tool loop

I built llm-api-adapter as a multi-provider Python interface for 57 registered model IDs from 9 LLM providers: OpenAI, Anthropic, Google, Mistral, xAI, Qwen, Kimi, DeepSeek, and Z.ai.

For function calling, the goal is simple: define the tool once and keep the same Python tool loop when the provider changes.

Here is a minimal one-round example using OpenAI.

Install

For OpenAI, Anthropic, and Google:

pip install llm-api-adapter
Enter fullscreen mode Exit fullscreen mode

The other provider integrations are installed through optional extras, for example:

pip install "llm-api-adapter[mistral,xai,qwen,kimi,deepseek,zai]"
Enter fullscreen mode Exit fullscreen mode

Define the tool

Let's use an inventory lookup. The inventory value lives in application state, so the model cannot know the answer without calling the function.

from llm_api_adapter.models.tools import ToolSpec


inventory_tool = ToolSpec(
    name="lookup_inventory",
    description="Return the current inventory for a SKU in a warehouse region.",
    json_schema={
        "type": "object",
        "properties": {
            "sku": {"type": "string"},
            "region": {
                "type": "string",
                "enum": ["us-east", "eu-central"],
            },
        },
        "required": ["sku", "region"],
        "additionalProperties": False,
    },
)
Enter fullscreen mode Exit fullscreen mode

The function itself belongs to the application:

INVENTORY = {
    ("SKU-1042", "eu-central"): 37,
    ("SKU-1042", "us-east"): 12,
}


def run_tool(name: str, arguments: dict) -> dict:
    if name != "lookup_inventory":
        raise ValueError(f"Unknown tool: {name}")

    key = (arguments["sku"], arguments["region"])

    return {
        "sku": arguments["sku"],
        "region": arguments["region"],
        "available_units": INVENTORY.get(key, 0),
    }
Enter fullscreen mode Exit fullscreen mode

Run the function-calling loop

import json
import os

from llm_api_adapter.models.messages.chat_message import (
    AIMessage,
    ToolMessage,
    UserMessage,
)
from llm_api_adapter.universal_adapter import UniversalLLMAPIAdapter


adapter = UniversalLLMAPIAdapter(
    organization="openai",
    model="gpt-5.6-sol",
    api_key=os.environ["OPENAI_API_KEY"],
)

messages = [
    UserMessage(
        "How many units of SKU-1042 are available in eu-central? "
        "Use lookup_inventory and do not guess."
    )
]

first = adapter.chat(
    messages=messages,
    max_tokens=4096,
    tools=[inventory_tool],
    tool_choice="auto",
)

if not first.tool_calls:
    print(first.content)

else:
    messages.append(
        AIMessage(
            content=first.content or "",
            tool_calls=first.tool_calls,
        )
    )

    for tool_call in first.tool_calls:
        result = run_tool(
            tool_call.name,
            tool_call.arguments,
        )

        messages.append(
            ToolMessage(
                tool_call_id=tool_call.call_id,
                content=json.dumps(result),
            )
        )

    final = adapter.chat(
        messages=messages,
        max_tokens=4096,
        previous_response=first,
    )

    print(final.content)
Enter fullscreen mode Exit fullscreen mode

When the tool is called, the final answer should report that SKU-1042 has 37 units available in eu-central.

The application-level flow is:

ToolSpec
   ↓
model
   ↓
ToolCall
   ↓
Python function
   ↓
ToolMessage
   ↓
model
   ↓
final response
Enter fullscreen mode Exit fullscreen mode

With tool_choice="auto", the model is allowed to answer without calling a tool. That is why the example handles both cases: if no ToolCall is returned, the application simply uses the model response; otherwise it executes the tool loop.

Switch providers without rewriting the loop

The ToolSpec, ToolCall, ToolMessage, and execution logic above do not depend on OpenAI.

For Anthropic, for example, the adapter configuration becomes:

adapter = UniversalLLMAPIAdapter(
    organization="anthropic",
    model="claude-sonnet-5",
    api_key=os.environ["ANTHROPIC_API_KEY"],
)
Enter fullscreen mode Exit fullscreen mode

For Google:

adapter = UniversalLLMAPIAdapter(
    organization="google",
    model="gemini-3.8-flash",
    api_key=os.environ["GOOGLE_API_KEY"],
)
Enter fullscreen mode Exit fullscreen mode

The same application-level tool loop is used for Mistral, xAI, Qwen, Kimi, DeepSeek, and Z.ai as well.

The provider setup itself is not always identical. External integrations require their corresponding package extra, and some APIs have additional connection parameters. Qwen Model Studio, for example, requires a workspace_id on every request.

Those differences stay in provider configuration rather than changing the tool contract.

Where tool_choice stops being portable

All 57 model IDs currently registered in the adapter support:

tool_choice="auto"
Enter fullscreen mode Exit fullscreen mode

But auto does not guarantee that a tool will be called. It means the model decides whether a tool is needed.

Forced tool selection is where provider and model differences start to appear.

Provider auto any Named tool
OpenAI Yes Yes Yes
Anthropic Yes Yes* Yes*
Google Yes Yes Yes
Mistral Yes Yes Yes
xAI Yes Yes Yes
Qwen Yes Yes* Yes*
Kimi Yes Yes* No
DeepSeek Yes Yes Yes*
Z.ai Yes No No

* Model-specific details:

  • Anthropic: claude-fable-5-1 and claude-opus-5-5 currently support only auto and none; the other registered Claude models support forced tool selection.
  • Qwen: Qwen 3.7 and 3.8 support forced any and named tools, but thinking must be disabled for those calls. The adapter handles this and warns unless reasoning_level="none" was already explicit.
  • Kimi: kimi-k3 supports any but not a named tool. kimi-k2.6 supports neither forced any nor a named tool.
  • DeepSeek: deepseek-flash supports named tool selection, but named selection and tool-result continuation require reasoning to be disabled for that tool loop.

Z.ai's currently registered glm-5.3-flash supports application tools with tool_choice="auto" only.

That is why I use auto as the common baseline.

What stays portable

The shared application contract is deliberately small:

Part of the tool flow Support
Define an application tool All 57 registered models
Tool name and description All 57 registered models
JSON Schema tool arguments All 57 registered models*
Receive a normalized ToolCall All 57 registered models
Parsed argument dictionary All 57 registered models
Tool call ID All 57 registered models
Execute the tool in application code Application code
Return a ToolMessage All 57 registered models
Continue after a tool result All 57 registered models
tool_choice="auto" All 57 registered models

* All 57 registered models accept JSON Schema for application-tool arguments. The exact schema vocabulary accepted by the underlying APIs can still differ; this does not remove any registered model from the basic tool-calling contract.

A normalized tool request gives the application the same fields:

tool_call.name
tool_call.arguments
tool_call.call_id
Enter fullscreen mode Exit fullscreen mode

tool_call.arguments is already a Python dictionary.

The adapter also deliberately does not execute the tool:

result = run_tool(
    tool_call.name,
    tool_call.arguments,
)
Enter fullscreen mode Exit fullscreen mode

That stays under application control.

The point of the abstraction is not to make 9 APIs appear identical. It is to keep the part the application actually depends on stable.

Continuation differs underneath

After a tool runs, application code adds the result in the same form:

messages.append(
    ToolMessage(
        tool_call_id=tool_call.call_id,
        content=json.dumps(result),
    )
)
Enter fullscreen mode Exit fullscreen mode

What happens underneath can differ.

For OpenAI models using the Responses API, previous_response can map to a provider response ID and use server-side continuation.

Other providers continue from explicit message history.

DeepSeek keeps conversation history explicit as well, while previous_response can carry matching reasoning-replay metadata rather than a server-side conversation ID.

The application still asks for the same thing:

continue after this tool result
Enter fullscreen mode Exit fullscreen mode

The provider-specific continuation mechanism stays below the adapter boundary.

Why the adapter does not execute tools

I deliberately stop the abstraction before tool execution.

Consider these tools:

get_weather()
search_documents()
create_support_ticket()
charge_customer()
delete_resource()
Enter fullscreen mode Exit fullscreen mode

Receiving a tool request is an LLM API concern.

Actually running that function is an application concern.

The distinction becomes important when a tool has side effects.

Imagine:

model requests create_support_ticket
        ↓
application creates the ticket
        ↓
LLM continuation request fails
Enter fullscreen mode Exit fullscreen mode

Blindly retrying the whole operation could create the ticket twice.

At that point the problem is no longer just function calling. It involves idempotency, retries, failover, checkpoints, and recovery after side effects.

I keep those concerns outside llm-api-adapter. I built a separate llm-api-resilience layer for retries, failover, circuit breakers, and checkpoint-based recovery instead of turning the provider adapter into an agent runtime.

The repository also includes cross-provider tests for the shared tool-calling contract.

The same contract is available asynchronously

Async does not require a different tool abstraction.

Install the async extra:

pip install "llm-api-adapter[async]"
Enter fullscreen mode Exit fullscreen mode

The async example below starts with a fresh message history rather than reusing the messages list modified by the synchronous example:

async_messages = [
    UserMessage(
        "How many units of SKU-1042 are available in eu-central? "
        "Use lookup_inventory and do not guess."
    )
]

first = await adapter.achat(
    messages=async_messages,
    max_tokens=4096,
    tools=[inventory_tool],
    tool_choice="auto",
)

if not first.tool_calls:
    print(first.content)

else:
    async_messages.append(
        AIMessage(
            content=first.content or "",
            tool_calls=first.tool_calls,
        )
    )

    for tool_call in first.tool_calls:
        result = run_tool(
            tool_call.name,
            tool_call.arguments,
        )

        async_messages.append(
            ToolMessage(
                tool_call_id=tool_call.call_id,
                content=json.dumps(result),
            )
        )

    final = await adapter.achat(
        messages=async_messages,
        max_tokens=4096,
        previous_response=first,
    )

    print(final.content)
Enter fullscreen mode Exit fullscreen mode

ToolSpec, ToolCall, and ToolMessage stay unchanged. astream_chat() follows the same contract for streaming.

The boundary that worked for me

The stable part of function calling ended up being:

ToolSpec
    ↓
ToolCall
    ↓
application execution
    ↓
ToolMessage
    ↓
continuation
Enter fullscreen mode Exit fullscreen mode

The provider adapter handles the different API formats and continuation mechanisms underneath it.

So the result is not 9 identical APIs.

It is one Python function-calling loop that does not need to be rewritten every time the provider changes.

The implementation and cross-provider tests are available in llm-api-adapter.

Top comments (0)