DEV Community

Vivek Yadav
Vivek Yadav

Posted on

Orchestrating Multiple AI Agents with the OpenAI Decisions API and Responses API

1. Introduction

Modern AI applications increasingly rely on multiple specialized agents rather than a single general-purpose assistant. A travel-planning application, for example, may have separate agents for hotel recommendations, flight search, hotel-and-flight packages, car rentals, and vacation-house rentals.

The challenge is determining which agents should be invoked for each user request.

The OpenAI Decisions API provides a way to classify incoming requests and identify the capabilities required to fulfill them. The selected agents can then use the OpenAI Responses API to perform their respective tasks, reason over available information, and generate results.

This article demonstrates an architecture in which the Decisions API acts as the decision and routing layer, while a Python orchestrator coordinates specialized agents that use the Responses API.

The examples use a travel-planning application, but the same architecture can be applied to enterprise support, financial services, shopping assistants, developer tools, and other multi-agent systems.

Important: Verify API availability, model names, request schemas, and SDK methods against the current OpenAI documentation before deploying. The code below illustrates the architecture; it may require adjustments for the SDK version and API schema enabled for your project.

2. Understanding the two APIs

OpenAI Decisions API

The Decisions API evaluates input against explicitly defined questions. Its documented question types include:

  • Predicate: Estimates the probability that a condition is true.
  • Choice: Selects one option from a predefined list.
  • Score: Evaluates input against a scoring rubric.

For multi-agent orchestration, predicate questions can be useful because multiple predicates can be evaluated against the same request. Each predicate can indicate whether a particular agent capability is needed.

For example, a request to find flights, reserve a hotel, and rent a car can activate three separate agent capabilities.

The Decisions API returns typed answers that the application can use to make routing decisions. It does not execute the agents itself.

Official documentation: https://developers.openai.com/api/docs/guides/decisions

OpenAI Responses API

The Responses API is the execution interface used by each specialized agent in this example. It can generate text, produce structured outputs, and use supported tools to perform more complex tasks.

Each travel agent can have its own instructions, business logic, tools, and integrations while sharing the same underlying API.

For example:

  • A hotel agent searches hotel inventory and compares amenities.
  • A flight agent searches available flights and evaluates itineraries.
  • A car-rental agent compares vehicles, pickup locations, and rental terms.
  • A house-rental agent evaluates vacation homes against the user's requirements.

The actual inventory searches should use authorized travel-provider APIs or other reliable data sources. A language model by itself should not be treated as a source of live prices or availability.

Official documentation: https://developers.openai.com/api/docs/guides/text

3. High-level architecture

The architecture separates classification, orchestration, and execution into distinct responsibilities.

Request flow

  1. The user submits a travel request.
  2. The application sends the request to the Decisions API.
  3. The Decisions API evaluates which travel capabilities are relevant.
  4. The orchestrator applies routing rules, authorization, and execution policies.
  5. The selected agents run using the Responses API and their respective tools.
  6. The orchestrator collects and combines the results.
  7. The application returns a consolidated response to the user.

Conceptually:

User Request
    |
    v
Decisions API
    |
    v
Python Orchestrator
    |
    +---- Hotel Agent -----------+
    +---- Flight Agent ----------+
    +---- Hotel + Flight Agent --+--> Result Aggregator --> User
    +---- Car Rental Agent ------+
    +---- House Rental Agent ----+
Enter fullscreen mode Exit fullscreen mode

The main advantage is that the application does not need to invoke every available agent for every request.

4. Define the travel-agent capabilities

Consider an application with the following five capabilities.

Agent Responsibility Example request
Hotel recommendation agent Find and compare hotels "Find a hotel in San Diego."
Flight agent Find and compare flights "Find flights from Phoenix to Seattle."
Hotel + flight agent Search bundled or coordinated flight-and-hotel options "Find a flight and hotel package for New York."
Car-rental agent Find suitable rental vehicles "Rent a car at the airport."
House-rental agent Find vacation homes and rental properties "Find a beach house for six people."

These capabilities can be combined to support requests involving several travel needs.

For instance, a user might ask:

"Plan a family vacation to Orlando. Find flights, a hotel, a rental car, and a vacation house as an alternative to the hotel."

The orchestrator could invoke the flight, hotel, car-rental, and house-rental agents concurrently.

The application should define an explicit policy for the hotel + flight agent. When the user requests a bundled package, the package agent can handle that search. When the user requests separate comparisons, the individual flight and hotel agents can be invoked instead.

This prevents unnecessary duplicate work.

5. Classify requests with the Decisions API

The following example uses Python and the OpenAI SDK.

Install or upgrade the SDK:

pip install -U openai
Enter fullscreen mode Exit fullscreen mode

Configure the API key through an environment variable.

macOS/Linux:

export OPENAI_API_KEY="your-api-key"
Enter fullscreen mode Exit fullscreen mode

Windows PowerShell:

$env:OPENAI_API_KEY = "your-api-key"
Enter fullscreen mode Exit fullscreen mode

The example evaluates five independent predicates.

from openai import AsyncOpenAI

client = AsyncOpenAI()

AGENT_QUESTIONS = [
    {
        "type": "predicate",
        "name": "hotel",
        "instructions": (
            "Does the user request hotel recommendations, "
            "hotel search, or hotel availability?"
        ),
    },
    {
        "type": "predicate",
        "name": "flight",
        "instructions": (
            "Does the user request flight search, "
            "flight comparisons, or flight availability?"
        ),
    },
    {
        "type": "predicate",
        "name": "hotel_flight_package",
        "instructions": (
            "Does the user explicitly request a combined "
            "flight-and-hotel package or bundled travel deal?"
        ),
    },
    {
        "type": "predicate",
        "name": "car_rental",
        "instructions": (
            "Does the user request a rental car or "
            "car-rental recommendations?"
        ),
    },
    {
        "type": "predicate",
        "name": "house_rental",
        "instructions": (
            "Does the user request a vacation home, "
            "holiday house, or house rental?"
        ),
    },
]


async def classify_request(user_request: str):
    decision = await client.decisions.create(
        model="gpt-6-luna",
        input=user_request,
        questions=AGENT_QUESTIONS,
    )

    return {
        answer.name: answer.probability
        for answer in decision.answers
        if answer.type == "predicate"
        and answer.name is not None
    }
Enter fullscreen mode Exit fullscreen mode

Implementation note: Treat this as an illustrative SDK example. Confirm the model identifier, method signature, answer fields, and supported question schema in the documentation for the Decisions API version available to your account.

Each predicate returns a probability between zero and one. These values represent the model's estimated confidence that a capability is relevant; they are not guaranteed correctness scores.

Because all questions evaluate the same input, the application can classify several capabilities in a single Decisions API request.

For example, the input:

"Find a flight and hotel package in Orlando, and also compare vacation homes for six people."

could produce a result conceptually similar to:

{
  "hotel": 0.96,
  "flight": 0.98,
  "hotel_flight_package": 0.99,
  "car_rental": 0.03,
  "house_rental": 0.97
}
Enter fullscreen mode Exit fullscreen mode

These values are illustrative, not actual API results. The example demonstrates how multiple capabilities can be evaluated independently.

6. Create specialized agents using the Responses API

Each agent should have a focused responsibility and its own instructions.

The following implementation demonstrates a reusable agent wrapper. In a production application, the instructions should be supplemented with appropriate provider integrations, validation, and domain-specific logic.

from openai import AsyncOpenAI

client = AsyncOpenAI()


async def run_travel_agent(
    agent_name: str,
    instructions: str,
    user_request: str,
):
    response = await client.responses.create(
        model="YOUR_SUPPORTED_MODEL",
        instructions=instructions,
        input=user_request,
    )

    return {
        "agent": agent_name,
        "result": response.output_text,
    }
Enter fullscreen mode Exit fullscreen mode

Replace YOUR_SUPPORTED_MODEL with a model available to your project.

Define the agents:

async def hotel_agent(request):
    return await run_travel_agent(
        "hotel",
        """
        You are a hotel recommendation agent.
        Identify the destination, dates, guest count,
        budget, and preferences.
        Compare suitable hotel options.
        Do not invent prices or availability.
        """,
        request,
    )


async def flight_agent(request):
    return await run_travel_agent(
        "flight",
        """
        You are a flight search agent.
        Identify origin, destination, dates,
        passenger count, and cabin preferences.
        Compare itineraries and clearly distinguish
        verified data from estimates.
        """,
        request,
    )


async def hotel_flight_agent(request):
    return await run_travel_agent(
        "hotel_flight_package",
        """
        You are a combined flight-and-hotel agent.
        Evaluate coordinated flight and hotel options.
        Compare the total package cost with separate
        bookings when reliable pricing is available.
        """,
        request,
    )


async def car_rental_agent(request):
    return await run_travel_agent(
        "car_rental",
        """
        You are a car-rental agent.
        Identify pickup and return locations, dates,
        vehicle type, and rental requirements.
        """,
        request,
    )


async def house_rental_agent(request):
    return await run_travel_agent(
        "house_rental",
        """
        You are a vacation-house rental agent.
        Identify dates, guests, budget, location,
        bedrooms, amenities, and property requirements.
        """,
        request,
    )
Enter fullscreen mode Exit fullscreen mode

These wrappers demonstrate the Responses API integration. They do not yet perform live travel searches. Actual search and pricing should be implemented using the appropriate provider APIs, with their results supplied to the agents.

7. Build the orchestrator

The orchestrator connects the classification results to the agent registry.

First, define the registry:

AGENT_REGISTRY = {
    "hotel": hotel_agent,
    "flight": flight_agent,
    "hotel_flight_package": hotel_flight_agent,
    "car_rental": car_rental_agent,
    "house_rental": house_rental_agent,
}
Enter fullscreen mode Exit fullscreen mode

Next, implement the dispatch logic:

import asyncio

THRESHOLD = 0.80


async def orchestrate(user_request: str):
    probabilities = await classify_request(user_request)

    package_requested = (
        probabilities.get("hotel_flight_package", 0.0)
        >= THRESHOLD
    )

    selected_agents = []

    for name, agent in AGENT_REGISTRY.items():
        probability = probabilities.get(name, 0.0)

        if probability < THRESHOLD:
            continue

        # Avoid duplicate hotel and flight work when
        # a combined package search is requested.
        if package_requested and name in {"hotel", "flight"}:
            continue

        selected_agents.append(agent)

    if not selected_agents:
        return {
            "status": "clarification_needed",
            "message": (
                "Please specify whether you need flights, "
                "a hotel, a rental car, or a vacation home."
            ),
        }

    results = await asyncio.gather(
        *(agent(user_request) for agent in selected_agents),
        return_exceptions=True,
    )

    return {
        "status": "completed",
        "results": [
            {"error": str(result)}
            if isinstance(result, Exception)
            else result
            for result in results
        ],
    }
Enter fullscreen mode Exit fullscreen mode

The threshold is illustrative and should be tuned using representative test requests. For important routing decisions, evaluate classification quality, add explicit fallback rules, and handle refusals and ambiguous inputs.

The package-exclusion rule is also simplified. In a real application, the router should distinguish between a package request and a request to compare a package against separate flight and hotel options.

8. Support combinations of agents

A multi-agent system should support both individual capabilities and combinations of capabilities.

User request Agents invoked
"Find a hotel in Miami." Hotel
"Find flights from Phoenix to Miami." Flight
"Find a flight and hotel package." Hotel + flight package
"Find a hotel and rent a car." Hotel + car rental
"Find flights, a hotel, and a car." Flight + hotel + car rental
"Compare a hotel package with separate flights and hotels." Package + flight + hotel
"Find a vacation home and a rental car." House rental + car rental
"Plan a trip with flights, a hotel, a car, and a vacation-home alternative." Flight + hotel + car rental + house rental

The final request illustrates a four-agent combination. Each agent performs its own search, and the orchestrator aggregates the results.

For requests involving comparisons, the package agent should not automatically suppress individual flight and hotel agents. The routing policy should preserve all capabilities needed to satisfy the user's explicit intent.

Parallel versus sequential execution

Parallel execution is appropriate when agents can work independently.

For example, a hotel search and a car-rental search can often run concurrently once the destination and dates are known.

Sequential execution is appropriate when one agent depends on another agent's output. For example, an itinerary planner may first determine the destination and dates, then pass those details to the flight, hotel, and rental-car agents.

The orchestrator should choose the execution strategy based on dependencies rather than treating every multi-agent request as a parallel fan-out.

9. Aggregate results from multiple agents

When multiple agents return results, the application should consolidate them into a coherent travel plan.

For example, an aggregator can:

  • Group results by flight, hotel, car, or house rental.
  • Normalize dates, currencies, and price formats.
  • Identify missing information and conflicting availability.
  • Compare total trip costs where reliable pricing is available.
  • Preserve provider links and booking terms.
  • Ask the user for clarification when essential details are missing.

The aggregator can use another Responses API call to summarize and structure the combined results. For production use, supply a clear output schema and pass the actual agent results as input rather than relying on a model to reconstruct missing facts.

A useful result format might contain:

{
  "destination": "Orlando",
  "flight_options": [],
  "hotel_options": [],
  "package_options": [],
  "car_rental_options": [],
  "house_rental_options": [],
  "comparison": {
    "currency": "USD",
    "total_cost": null,
    "notes": []
  }
}
Enter fullscreen mode Exit fullscreen mode

The empty arrays and null values are placeholders, not search results. The application should populate them from its agent outputs and provider data.

10. Production considerations

A reliable implementation requires more than selecting agents.

Security and authorization: Apply authentication, authorization, and policy checks before executing an agent. A classification result must never grant permission to access a tool, account, or booking system.

Data quality: Use reliable travel-provider integrations for live prices, schedules, inventory, and availability. Clearly identify estimates and stale data.

Failure handling: One agent may fail while other agents succeed. Use timeouts, bounded retries, and partial-result handling rather than failing the entire request unnecessarily.

Clarification: If dates, destination, passenger count, or other required fields are missing, ask the user for them before making a live search or booking.

Booking confirmation: Treat searching and recommending differently from purchasing. Require explicit user confirmation before committing to a booking or payment.

Observability: Record routing decisions, selected agents, latency, token usage, errors, and provider calls. Avoid logging sensitive personal or payment data unnecessarily.

Evaluation: Measure false-positive and false-negative routing decisions, end-to-end latency, agent success rates, and the quality of aggregated results. Do not assume a probability threshold alone guarantees good routing.

Cost optimization: The Decisions API can help avoid invoking irrelevant agents. However, savings depend on the workload, model pricing, the number of agent calls avoided, and the cost of classification. Measure the overall system rather than assuming every request will be cheaper.

11. When should you use the Decisions API?

The Decisions API is particularly useful when your application needs fast, bounded classification decisions, such as identifying the capabilities required by a request.

It is not a replacement for the Responses API or the orchestration layer.

Use the Decisions API for classification and routing, the Responses API for agent execution and synthesis, and application code for permissions, execution policies, retries, and dependency management.

If an application needs to generate arbitrary structured data or ask a model to invoke tools directly, Responses API structured outputs and function calling may be more appropriate for that specific task.

12. Conclusion

The OpenAI Decisions API provides a way to identify which specialized AI capabilities are relevant to a user request. Combined with the Responses API, it supports a modular architecture in which each agent has a focused responsibility and the orchestrator controls execution.

For the travel-planning example, the architecture can route requests to hotel recommendations, flight search, hotel-and-flight packages, car rentals, house rentals, or combinations of these capabilities.

The key design principle is to separate deciding what needs to happen, executing the required agents, and combining their results. This separation improves modularity, enables independent agent development, and provides a foundation for more scalable multi-agent applications.

References

  1. OpenAI Decisions API guide: https://developers.openai.com/api/docs/guides/decisions
  2. OpenAI Decisions API Python reference: https://developers.openai.com/api/reference/python/resources/decisions/methods/create
  3. OpenAI Responses API text guide: https://developers.openai.com/api/docs/guides/text
  4. OpenAI API documentation: https://developers.openai.com/api/docs

Top comments (0)