DEV Community

Ujjwal B Soni
Ujjwal B Soni

Posted on

Running AWS Strands Decider 2B Locally: A Complete Setup Guide for AI Routing & Multi-RAG Systems

As GenAI applications become more sophisticated, one challenge continues to surface:

How do we make reliable decisions before invoking an LLM?

For example:

  • Which datasource should answer this question?
  • Should I query PLM, Jira, or Confluence?
  • Do I have enough context to answer confidently?
  • Should an AI agent invoke a tool or ask a clarification question?

Traditionally, we let an LLM make these decisions.

Recently, AWS introduced Strands Decider 2B, a lightweight decision model designed specifically for routing, classification, scoring, and orchestrating agent workflows.

Unlike traditional LLMs, Strands Decider doesn't generate arbitrary text. Instead, it selects from predefined options and provides confidence scores, making it ideal for Agentic AI and Multi-RAG systems.

In this article, I'll walk through how I installed and tested Strands Decider 2B locally on Windows using WSL2.


Why Use a Decision Model?

A common architecture today looks like this:

Existing architecture

The problem?

The LLM is responsible for both:

  • Making decisions
  • Generating responses

A better approach is:

Better approach using AWS Strands Decider

Now the LLM focuses on reasoning and generation, while the decision model handles routing and orchestration.


Environment

For this walkthrough I used:

  • Windows 11
  • WSL2
  • Ubuntu
  • Python Virtual Environment
  • Strands Decider 2B

Step 1: Verify WSL2 Installation

Open PowerShell:

wsl -l -v
Enter fullscreen mode Exit fullscreen mode

Example output:

NAME      STATE    VERSION
Ubuntu    Running  2
Enter fullscreen mode Exit fullscreen mode

Launch Ubuntu:

wsl -d Ubuntu
Enter fullscreen mode Exit fullscreen mode

Step 2: Create a Workspace

mkdir -p /mnt/c/GENAI/strands

cd /mnt/c/GENAI/strands
Enter fullscreen mode Exit fullscreen mode

Step 3: Install Required Packages

Update Ubuntu:

sudo apt update
Enter fullscreen mode Exit fullscreen mode

Install required dependencies:

sudo apt install -y \
 python3 \
 python3-pip \
 python3-venv \
 python3-dev \
 build-essential \
 gcc \
 g++
Enter fullscreen mode Exit fullscreen mode

Step 4: Create a Virtual Environment

Create the environment:

python3 -m venv .venv
Enter fullscreen mode Exit fullscreen mode

Activate it:

source .venv/bin/activate
Enter fullscreen mode Exit fullscreen mode

Upgrade pip:

pip install --upgrade pip setuptools wheel
Enter fullscreen mode Exit fullscreen mode

Step 5: Install Strands Decider

pip install strands-decider
Enter fullscreen mode Exit fullscreen mode

Verify installation:

strands-decider --help
Enter fullscreen mode Exit fullscreen mode

Step 6: Discover Available Models

Install Hugging Face Hub:

pip install huggingface_hub
Enter fullscreen mode Exit fullscreen mode

List available Strands models:

python -c "from huggingface_hub import list_models; [print(m.id) for m in list_models(search='strands')]"
Enter fullscreen mode Exit fullscreen mode

The model used in this guide:

StrandsAgents/strands-decider-2B-hobson-v19
Enter fullscreen mode Exit fullscreen mode

Step 7: Start the Model Server

Launch the model:

strands-decider serve StrandsAgents/strands-decider-2B-hobson-v19 --device cpu
Enter fullscreen mode Exit fullscreen mode

Expected output:

Application startup complete.
Uvicorn running on http://127.0.0.1:8000
Enter fullscreen mode Exit fullscreen mode

The first startup downloads and caches the model automatically.


Step 8: Verify the API

Open in your browser:

http://127.0.0.1:8000/docs
Enter fullscreen mode Exit fullscreen mode

Or:

curl http://127.0.0.1:8000/openapi.json
Enter fullscreen mode Exit fullscreen mode

Understanding Question Types

Strands Decider supports three decision formats.

Choice Question

Choose one option from a list.

{
  "type": "choice",
  "instructions": "Select the best datasource.",
  "criteria": {
    "PLM": "Engineering changes and parts",
    "JIRA": "Issue tracking system",
    "CONFLUENCE": "Documentation repository",
    "UNKNOWN": "No suitable source"
  }
}
Enter fullscreen mode Exit fullscreen mode

Noul Question

Yes / No decision.

{
  "type": "noul",
  "instructions": "Determine whether this statement is true."
}
Enter fullscreen mode Exit fullscreen mode

Score Question

Rate against an ordered scale.

{
  "type": "score",
  "instructions": "Rate the sentiment.",
  "criteria": [
    "Very Negative",
    "Negative",
    "Neutral",
    "Positive",
    "Very Positive"
  ]
}
Enter fullscreen mode Exit fullscreen mode

Testing with Python

Create a file called:

decider_demo.py
Enter fullscreen mode Exit fullscreen mode
import requests
import json

payload = {
    "state": "User wants ECO information",
    "questions": {
        "datasource": {
            "type": "choice",
            "instructions": "Select the most appropriate datasource.",
            "criteria": {
                "PLM": "Engineering changes and parts",
                "JIRA": "Issue tracking system",
                "CONFLUENCE": "Documentation repository",
                "UNKNOWN": "No suitable source"
            }
        }
    }
}

response = requests.post(
    "http://127.0.0.1:8000/v1/systemone",
    json=payload
)

print(json.dumps(response.json(), indent=2))
Enter fullscreen mode Exit fullscreen mode

Run:

python decider_demo.py
Enter fullscreen mode Exit fullscreen mode

Sample output:

{
  "answers": {
    "datasource": {
      "choice": "PLM",
      "confidence": 0.94
    }
  }
}
Enter fullscreen mode Exit fullscreen mode

Using It as a Multi-RAG Router

One use case I was particularly interested in was reducing hallucinations across multiple RAG systems.

Routing logic becomes simple:

decision = response["answers"]["datasource"]["choice"]

if decision == "PLM":
    plm_rag.search(query)

elif decision == "JIRA":
    jira_rag.search(query)

elif decision == "CONFLUENCE":
    confluence_rag.search(query)

else:
    print("No reliable datasource identified.")
Enter fullscreen mode Exit fullscreen mode

Instead of asking an LLM to guess which datasource to use, the decision model handles routing first.


Real-World Enterprise Use Cases

I see strong potential in the following scenarios:

✅ Multi-RAG orchestration

✅ Agent tool selection

✅ Engineering Change workflows

✅ PLM assistants

✅ SharePoint routing

✅ Confluence routing

✅ Jira ticket management

✅ Intent classification

✅ Confidence-based validation

✅ Hallucination reduction


Final Thoughts

One of the biggest lessons I've learned building GenAI applications is:

Not every problem requires text generation.

Decision-making and text generation are fundamentally different tasks.

Using a decision model before retrieval and generation creates a much cleaner architecture:

 Decision Model
      ↓
  Retrieval
      ↓
     LLM
Enter fullscreen mode Exit fullscreen mode

For enterprise AI systems, agentic workflows, and multi-RAG architectures, this pattern improves reliability, control, and observability.

If you're building AI agents today, I highly recommend experimenting with decision models as part of your architecture.


GitHub Repository

Repository: https://github.com/ujjwalbsoni/strands-decider-end-to-end


About the Author

Ujjwalkumar Soni

Passionate about AI Agents, RAG Architectures, Knowledge Management, and Enterprise GenAI Solutions.

Let's connect and share ideas around Agentic AI and next-generation enterprise applications.

ai #aws #python #rag #agenticai #llm #genai #machinelearning #softwareengineering #opensource

Top comments (0)