OpenAI shipped a Decisions API. It isn't a chat model you coax into choosing: it
answers a bounded question directly. You give it a request and a fixed list of
options, it returns one of them with a confidence. No generated text, no option
it invents, nothing to parse. The returned choice is always one you supplied.
An agent is full of bounded questions, so I wired it into a
Strands agent two ways. Both are in the repo and
both run. Here they are side by side.
A quick word on Strands
Strands Agents is an open-source SDK for building AI agents. You create one in a couple of lines: give it a model, a system prompt, and the tools it can call.
from strands import Agent
from strands.models import BedrockModel
agent = Agent(
model=BedrockModel(model_id="global.anthropic.claude-haiku-4-5-20251001-v1:0"),
system_prompt="You are a travel agent that helps people with hotels.",
tools=[get_hotel_room_rates, book_hotel],
)
agent("How much is a room at the Cliffside Resort?")
Tools, hooks, a model router: they attach to the agent like Lego pieces. Add a tool and the agent can call it; add a hook and you run your own code at a point in the agent's lifecycle; swap the model for a router and the agent chooses its model per request. The agent stays the same two lines; you compose behavior by clicking pieces onto it. The two ways below are exactly that: one hook, one router.
The call, once
Both uses share one request shape, through the official async client:
from openai import AsyncOpenAI
oai = AsyncOpenAI() # reads OPENAI_API_KEY
resp = await oai.decisions.create(
model="gpt-6-luna",
input="How much is a room at the Cliffside Resort?",
questions=[{
"type": "choice",
"name": "pick",
"instructions": "Which option best answers this request?",
"choices": [ # one {value, description} per option
{"value": "get_hotel_room_rates", "description": "Nightly room rates for a hotel"},
{"value": "get_weather", "description": "Current weather for a city"},
],
}],
)
answer = next(a for a in resp.answers if a.name == "pick")
answer.choice # -> "get_hotel_room_rates"
answer.confidence # -> 0.97
What changes between the two uses is only what the choices are: tools in one,
models in the other. Both are below. One narrows the toolbox a model sees; the
other picks which model runs. Pick the one your agent needs, or use both.
Way 1: choose the tool, and cut the tokens
Give a small model forty overlapping tools and it picks the wrong one or invents
one. Every tool is also tokens it pays for on every turn: name, description,
schema. Choosing the relevant tool before the model runs fixes both, and in
Strands that choice lives in a hook (BeforeModelCallEvent), because it
changes what the agent sees: the hook swaps the agent's tool registry down to
the pick. The selector is one function; everything else is the hook.
oai_choices = [{"value": t.tool_name, "description": t.tool_spec["description"]} for t in ALL_TOOLS]
async def openai_select(query):
resp = await oai.decisions.create(
model="gpt-6-luna", input=query,
questions=[{"type": "choice", "name": "tool",
"instructions": "Which tool answers this request?", "choices": oai_choices}],
)
answer = next(a for a in resp.answers if a.name == "tool")
return [BY_NAME[answer.choice]]
agent = Agent(model=AGENT_MODEL, tools=ALL_TOOLS, system_prompt=PROMPT,
hooks=[ToolFilterHook(openai_select)], callback_handler=None)
await agent.invoke_async("How much is a room at the Cliffside Resort?")
Over the demo's 40-tool pool, each pick is a bounded choice rather than a
generated tool name:
get_hotel_room_rates (confidence 1.00) <- How much is a room at the Cliffside Resort?
get_weather_forecast (confidence 0.99) <- What's the weather forecast for Tokyo next week?
book_hotel (confidence 0.97) <- Book Cliffside Resort for John Smith for 2 nights
Way 2: choose the model, and match cost to the request
Not every request needs your strongest model. Routing the easy ones to a small,
cheap model and reserving the capable model for hard work keeps quality where it
matters without paying top price on trivial questions.
Strands routes each request to one of several models with ModelRouter, and the
choice is a pluggable RoutingStrategy. This choice does not live in a hook: you
pass the router as the agent's model, which is Strands' seam for picking the
model (the docs say so plainly: pass the router as the model, not through
plugins). Plug the Decisions API in as the classifier: the router's candidates
become the choices, and the decision model picks which model serves the turn. It
judges the request against each candidate's description, and knows nothing
about the models themselves.
class OpenAIDecisionsRoutingStrategy(RoutingStrategy):
def __init__(self):
self._oai = AsyncOpenAI()
async def select(self, context, **kwargs):
if context.attempts: # route the opening attempt only
return None
query = _user_text(context)
choices = [{"value": c.name, "description": c.description or c.name} for c in context.candidates]
resp = await self._oai.decisions.create(
model="gpt-6-luna", input=query,
questions=[{"type": "choice", "name": "model",
"instructions": "Which model should handle this request?", "choices": choices}],
)
answer = next(a for a in resp.answers if a.name == "model")
return next((c for c in context.candidates if c.name == answer.choice), None)
router = ModelRouter(models=[routine, advanced], strategy=OpenAIDecisionsRoutingStrategy())
agent = Agent(model=router, callback_handler=None)
The routine lookups go to the small model, the reasoning tasks to the strong
one:
routine (confidence 1.0) <- What time zone is Tokyo in?
routine (confidence 1.0) <- What is the capital of Australia?
advanced (confidence 1.0) <- Design a rollback-safe idempotency-key migration
advanced (confidence 0.98) <- Compare optimistic vs pessimistic locking, recommend one
So which one?
Both are the same one-call pattern behind different Strands seams: a hook for the
tool, a router for the model. A decision model is a chooser, so you stop asking
a generator to choose and hoping its text parses into a real option. How would
you point it at your agent? Which choice in your loop is the one you don't trust
a chat model to make?
Try it yourself
Both ways run end to end against your own questions. Set OPENAI_API_KEY and go:
git clone https://github.com/elizabethfuentes12/why-agents-fail-sample-for-amazon-agentcore
cd why-agents-fail-sample-for-amazon-agentcore/02-semantic-tools-demo
uv venv && uv pip install -r requirements.txt
If it saves you time, give the repo a ⭐. It helps others find it.
Resources
- OpenAI Decisions API
- Strands hooks
- Strands model routing
- Why AI Agents Fail (the companion repo)
¡Gracias!
Top comments (2)
The always-one-of-your-options part is what got me, chasing a hallucinated tool name through agent logs is pure misery. Way 1 adds a full decisions round trip before the model even runs though, did you measure the latency hit versus just eating the extra tool tokens?
Yes in This blog builder.aws.com/content/3KN6CZDKG0...