DEV Community

Manish Hatwalne
Manish Hatwalne

Posted on

How to get multiple AI agents working together on one task: a WorkSwarm walkthrough

Most multi-agent demos fall apart at the handoff. One agent returns a paragraph when the next one expects a list. A subtask stalls, and nothing notices. Before long, you are writing dispatch loops and retry logic around every model call.

This tutorial shows a better way. You will build a Cluster-mode Swarm in WorkSwarm (formerly JiuwenSwarm), where a Leader hands a product-analysis task to three specialist Teammates. Every step's output is checked against a JSON schema (an agreed output format), so the next agent always gets the shape it expects. You can even test the whole pipeline offline with pytest before it touches a paid model.

Along the way, you will

  • write the SwarmFlow script that defines the pipeline
  • configure the models your Teammates use
  • run the smoke tests offline
  • watch a live run from the Web UI

The complete code is in the companion repository.

Prerequisites and installation

You will need:

  • Python 3.11, 3.12, or 3.13
  • API credentials for at least one supported LLM provider

WorkSwarm is published on PyPI as workswarm, and installing it also pulls in Agent Core and the SwarmFlow engine.

The companion repository pins version 0.2.6 in requirements.txt. Clone it and install its dependencies:

git clone https://github.com/manishh/openjiuwen-how-to-get-multiple-ai-agents-actually-working-together
cd openjiuwen-how-to-get-multiple-ai-agents-actually-working-together
pip install -r requirements.txt
Enter fullscreen mode Exit fullscreen mode

Next, copy the companion repository's example config into the WorkSwarm config directory:

mkdir -p ~/.jiuwenswarm/config/
cp config/config.yaml ~/.jiuwenswarm/config/config.yaml
Enter fullscreen mode Exit fullscreen mode

Then initialize the workspace and start the services. The command-line tools kept their original names after the rename, so they still start with jiuwenswarm:

jiuwenswarm-init
jiuwenswarm-start
Enter fullscreen mode Exit fullscreen mode

jiuwenswarm-init creates the ~/.jiuwenswarm/ workspace on first run. Because you’ve already copied config.yaml into that directory, init merges it with the default version, so none of your settings are lost.

jiuwenswarm-start launches the Gateway, the AgentServer, and the Web UI at http://localhost:5173:

WorkSwarm

If you'd rather skip pip, the install guide has one-click desktop installers for Windows and macOS, and the README also lists a HarmonyOS desktop version.

How WorkSwarm's coordination engineering model works

WorkSwarm splits a multi-agent job between a Leader and a team of Teammates. The Leader reads the goal, forms the team, and breaks the work into tasks. Each Teammate claims a task that fits its role. If a Teammate gets blocked, it escalates to the Leader (asks the Leader to step in). Once the work is done, the Leader merges the results into the final deliverable.

openJiuwen calls this approach coordination engineering. The Agent Team guide walks through the full Leader and Teammate cycle.

By default, the Leader plans that cycle adaptively, choosing the order as it goes. That flexibility suits exploratory work, but two runs can take different paths. When you need a repeatable pipeline, write a SwarmFlow script. You define the stages as a Python file, and the Leader runs them in the order your code sets. The pipeline keeps the same shape every time, and it lives in version control, where you can review and test it.

Failure handling lives in Agent Core, the engine underneath WorkSwarm. The Agent Core v0.1.18 release added ReliabilityRail, a set of six detectors that watch for repeated tool calls, model errors, tool errors, output-length problems and context-compaction issues, and that catch agents bouncing work back and forth. Each detector has configurable severity actions: correct the problem locally, escalate to the Leader, or notify the user. This is a separate path from Teammate escalation: a Teammate escalates when it is blocked, while ReliabilityRail steps in when it detects a failure pattern. The same release also made third-party agents more resilient: transient timeouts, rate limiting (429), and service-unavailable (5xx) errors are retried, with guidance surfaced to the user and consistent final error codes.

The diagram below shows how this maps to the pipeline you will build. The Leader runs the script through three phases: research, summary, and report. ReliabilityRail watches the run and corrects, escalates, or notifies when something goes wrong.

flowchart TD
    L["Leader<br/>runs swarmflow_pipeline.py"] --> R["Phase 1: Research<br/>market_researcher"]
    R -->|"trends, competitors, opportunities"| S["Phase 2: Summarize<br/>data_summarizer"]
    S -->|"exactly 5 bullets"| P["Phase 3: Report<br/>report_writer"]
    R -->|"full research"| P
    P -->|"report"| L
    E["Agent Core<br/>ReliabilityRail"] -.->|"correct/escalate/notify"| L

Leader agent configuration and task decomposition

A SwarmFlow script has two required top-level pieces: a literal META dict that describes the pipeline, and an async def run(args) entry point that the engine awaits. Everything inside run() is ordinary Python plus three operators the engine supplies: agent() to call a Teammate, phase() to mark a stage, and log() to write a progress line.

import json

from swarmflow import agent, phase, log
Enter fullscreen mode Exit fullscreen mode

META: descriptive, and it must be a literal

The engine reads META from the file's syntax tree without importing the script, so it has to be a plain dict literal. A META built by a function call or computed value is rejected with a MetaError. That's also why the engine can list your pipeline before running any of its code.

META: dict = {
    # Human-readable identifier surfaced in the Web UI (http://localhost:5173)
    "name": "product-analysis-pipeline",
    # Shown on the Swarm task card
    "description": (
        "Three-stage product-analysis Swarm: market research → "
        "data summarization → executive report.  Demonstrates Cluster-mode "
        "Leader/Teammate coordination."
    ),
    "version": "1.0.0",
    # Subtask metadata used for documentation and the Web UI task-progress panel.
    # The engine reads the actual dependency graph from the pipeline
    # structure below, not from this list.
    "subtasks": [
        {
            "role": "market_researcher",
            "description": "Gather and structure market data",
            "success_criteria": "JSON with keys: trends, competitors, opportunities",
        },
       {
            "role": "data_summarizer",
            "description": "Distill the research into five actionable bullets",
            "success_criteria": "JSON with a 'bullets' list of exactly 5 strings",
        },
        {
            "role": "report_writer",
            "description": "Write the executive report from research and summary",
            "success_criteria": "JSON with Executive Summary, Key Findings, Recommendations",
        },

    ],
}
Enter fullscreen mode Exit fullscreen mode

Treat subtasks as documentation. Nothing in it schedules or enforces anything. The real dependency graph is the order of the awaits in run(), and the real contract is the schema attached to each call.

Schemas make each step's contract enforceable

The success_criteria string says what a step should return. The schema= argument is what the engine checks. Each step returns a parsed, validated object that the next step can use directly, with no string parsing in between. This is a new feature introduced in the latest WorkSwarm for obtaining correct and structured output for agents to work with.

RESEARCH_SCHEMA = {
    "type": "object",
    "properties": {
        "trends": {"type": "array", "items": {"type": "string"}},
        "competitors": {"type": "array", "items": {"type": "string"}},
        "opportunities": {"type": "array", "items": {"type": "string"}},
    },
    "required": ["trends", "competitors", "opportunities"],
}
SUMMARY_SCHEMA = {
    "type": "object",
    "properties": {
        "bullets": {
            "type": "array",
            "items": {"type": "string"},
            "minItems": 5,
            "maxItems": 5,
        },
    },
    "required": ["bullets"],
}

REPORT_SCHEMA = {
    "type": "object",
    "properties": {
        "executive_summary": {"type": "string"},
        "key_findings": {"type": "array", "items": {"type": "string"}},
        "recommendations": {"type": "array", "items": {"type": "string"}},
    },
    "required": ["executive_summary", "key_findings", "recommendations"],
}
Enter fullscreen mode Exit fullscreen mode

minItems and maxItems are what turn "exactly 5 bullets" from a request in the prompt into a rule the engine checks.

run(): plain awaits, each feeding the next

The three steps (market research, data summarization, and report writing) depend on each other, so they are sequential awaits. The body of run() is a chain of three agent() calls. Each call sends a prompt to a Teammate, waits for the reply, and returns it as a parsed object that already matches its schema. The next call embeds that object in its own prompt.

async def run(args):
task = args["task"]
# Phase 1: research
phase("Research")
log("[Leader] Phase 1 - dispatching market_researcher Teammate")
research = await agent(
    f"{task}\nReturn the trends, competitors and opportunities you find. "
    "Keep it short and return results quickly.",
    label="market_researcher",
    schema=RESEARCH_SCHEMA,
)
Enter fullscreen mode Exit fullscreen mode

The first step is the only one that sees the raw task. Its schema guarantees research is a dict with trends, competitors and opportunities, each a list of strings. The instruction to keep the answer short and fast keeps the walkthrough quick. Drop it if you want deeper research (it will take longer).

# Phase 2: summarize
phase("Summarize")
log("[Leader] Phase 2 - dispatching data_summarizer Teammate")
summary = await agent(
    "Distill this market research into exactly 5 concise, actionable bullet points:\n"
    + json.dumps(research),
    label="data_summarizer",
    schema=SUMMARY_SCHEMA,
)
Enter fullscreen mode Exit fullscreen mode

The summarizer never sees the original task. It receives only research, serialized with json.dumps so it can be embedded in the prompt as text. Because research was validated in step 1, this step can rely on its shape. SUMMARY_SCHEMA enforces the "exactly 5" requirement, which is why the result is an object with a bullets list.

# Phase 3: report
phase("Report")
log("[Leader] Phase 3 - dispatching report_writer Teammate")
report = await agent(
    "Write an executive report with sections 'Executive Summary', 'Key Findings' "
    "and 'Recommendations' from this research and summary:\n"
    + json.dumps({"research": research, "summary": summary}),
    label="report_writer",
    schema=REPORT_SCHEMA,
)

return {"report": report}
Enter fullscreen mode Exit fullscreen mode

The writer gets both earlier outputs, the full research and the five-bullet summary, so it can use the detail from step 1 and the priorities from step 2. This is the one place where data fans in: step 3 reads from two earlier steps, not just the one before it. REPORT_SCHEMA requires the three named sections, and run() returns the validated report as {"report": report}.

Three details about how this behaves:

  • Order comes from the code. Each await blocks until its Teammate finishes, so step 2 can't start before step 1 returns. Nothing else defines the dependency graph, and META["subtasks"] doesn't affect it.
  • phase(), log() and label= are for readability. They group and name steps in the run's progress view and event stream. Removing them wouldn't change what the pipeline computes.
  • Completed steps are journaled. If a run is resumed from its journal, finished steps replay from the record instead of calling the model again, so a failure in step 3 doesn't force you to pay for steps 1 and 2 twice.

Testing without a model

The swarmflow module, which provides agent, log and phase, only exists while the engine is executing a script. Running python src/swarmflow_pipeline.py directly therefore fails with an import error. Instead, the repository's tests drive the script through Agent Core's run_workflow() with its built-in MockBackend, which returns deterministic replies that match each schema. All nine tests run without an API key.

pytest -v tests/smoke_test.py
Enter fullscreen mode Exit fullscreen mode

The tests check that:

  • the agent, log and phase operators resolve
  • META and run() have the right shape
  • all three agent() calls start and complete
  • a resumed run replays from its journal without calling the model again

They also catch API misuse early. For example, passing an unsupported keyword such as role= to agent() fails here with a TypeError instead of later in the Web UI.

Choosing a model per step

The label="market_researcher" on an agent() call is only a name that shows up in the run tree; it doesn't pick a model. Every agent() call runs on the teammate model by default. To run a step on a different model, name that model on the call:

research = await agent(
    prompt,
    label="market_researcher",
    schema=RESEARCH_SCHEMA,
    model="model_name",  
)
Enter fullscreen mode Exit fullscreen mode

This keeps the pipeline portable: provider details and credentials stay in WorkSwarm's configuration, and the script only says which model a step should use.

Setting up providers

You can add providers in the Web UI (Settings > Models) or in ~/.jiuwenswarm/config/config.yaml. Each model entry has:

  • model_name, for example gpt-4o or deepseek-chat
  • api_base, up to the version level only (for example https://api.openai.com/v1); WorkSwarm appends /chat/completions itself
  • api_key
  • client_provider
  • alias (optional; defaults to model_name), which is the name you refer to the model by

One entry is marked with is_default: true, and that's the fallback when nothing else is specified.
WorkSwarm supports OpenAI, DeepSeek, DashScope, SiliconFlow, InferenceAffinity, OpenRouter and Huawei Cloud MaaS, plus any OpenAI-compatible endpoint. The configuration guide has the full reference.

WorkSwarm: Configure Model

You can also supply keys as environment variables, which take precedence over values in the config file. Export them before you start the services:

export OPENAI_API_KEY=your-openai-key
export DEEPSEEK_API_KEY=your-deepseek-key
Enter fullscreen mode Exit fullscreen mode

How a step picks its model

The SwarmFlow guide lists model among the keys of the options bag that agent() accepts, and recommends leaving it out by default so each worker inherits the teammate model. This tutorial's pipeline does exactly that: it passes no options, so all three Teammates run on the same default model. That's the simplest setup, and it works even with a single LLM.

When to split the work across models

The three steps have quite different workloads:

  • market_researcher starts from a one-line task and has to reason across trends, competitors, and opportunities. That's open-ended work, and a larger model earns its cost here.
  • data_summarizer gets JSON that's already structured and returns five bullets. A smaller, faster model handles that well.
  • report_writer writes three long sections from both earlier outputs, so it belongs on the larger model again.

So the natural refinement is to register a cheaper model under an alias in your configuration, then pass that alias on the summarizer call only:

summary = await agent(
    "Distill this market research into exactly 5 concise, actionable bullet points:\n"
    + json.dumps(research),
    label="data_summarizer",
    schema=SUMMARY_SCHEMA,
    options={"model": "<alias>"},
)
Enter fullscreen mode Exit fullscreen mode

The other two calls stay untouched and keep using the default. You cut the cost of the lightest step, and the pipeline logic doesn't change. Because the schema still checks the output, the smaller model has to produce the same five-bullet shape, or the step doesn't pass.

How to run the Swarm and read the monitoring output

  1. Open http://localhost:5173.
  2. Choose Cluster mode in the chat box.
  3. Load src/swarmflow_pipeline.py with the + button in the chat box.
  4. Ask the Leader to start the Swarm, for example by entering Start the swarm.

The task-progress panel then shows each phase and Teammate as it runs.

Swarm start

You can also open the group chat, which shows the agents' messages to each other.

WorkSwarm: Agents' group chat

A healthy run moves through Research, Summarize, and Report in order, with each [Leader] log line appearing as its phase starts.

WorkSwarm in action

When ReliabilityRail escalates a problem, the Leader's response appears in the group chat, and the run's status changes in the SwarmFlow tree view. From there you can pause, resume, or stop the run. Paused runs don't resume on their own. After you send a message, the Leader decides whether to resume them. If you prefer the terminal, WorkSwarm also has a terminal UI (TUI) for macOS, Windows, and Linux, where /swarmflows opens the same run, phase, and node view.

When all three Teammates finish and the workflow completes, the Leader reports back with a summary.

WorkSwarm completed with summary

Where to go from here

You now have a working three-step Swarm, and how it's built matters as much as what it does:

  • Step order lives in plain Python. One await follows another, so you can read the sequence straight off the code without depending on META.
  • Every handoff is schema-checked. The JSON schema on each agent() call means each step receives a validated object instead of free text it has to parse.
  • The whole flow runs offline. The MockBackend tests exercise all three steps, the event stream, and journal replay without an API key, so structural mistakes show up before a live run spends any tokens.
  • Model choice is per step. All three Teammates can share the default model, and a single options={"model": "<alias>"} moves one step to a cheaper one.

From here, the WorkSwarm docs describe three directions to extend the patterns:

  • Human-in-the-Loop adds approval gates between phases, for workflows where a person has to sign off before the next Teammate starts.
  • Distributed deployment spreads the Leader and Teammates across processes and machines once a single host becomes the bottleneck.
  • The Experience Closed-Loop mechanism lets WorkSwarm reuse task-breakdown templates from earlier successful runs, so decomposing similar tasks gets faster over time.

To get started, clone the repo, read the Agent Team and SwarmFlow guides, and launch your first Cluster-mode Swarm from github.com/openJiuwen-ai/jiuwenswarm.


Frequently asked questions about WorkSwarm multi-agent pipelines

What is WorkSwarm's Cluster mode, and how does it differ from Agent mode?

In Agent mode, a single agent handles the task on its own. Cluster mode turns on multi-agent collaboration: a Leader breaks the task into subtasks and coordinates specialist Teammates, with phase tracking and ReliabilityRail failure detection across the run. Choose Cluster mode when a task splits naturally into stages that need different expertise, or different models.

Is WorkSwarm the same project as JiuwenSwarm?

Yes. JiuwenSwarm was renamed WorkSwarm, and the current release installs from PyPI as workswarm. The Python import package, the command-line tools, the ~/.jiuwenswarm/ config folder, and the GitHub repository still use the old name as of September 2026.

How does a SwarmFlow script get into a Team session?

There are three ways:

  • The Leader writes a script on the fly by calling its swarmflow() tool.
  • You install a Swarm Skill that bundles a workflow script.
  • You prepare a script offline and have the Leader run it with swarmflow(script_path=...).

How do I reuse a SwarmFlow script across projects?

Package it as a Swarm Skill with the built-in swarmskill-creator skill. You can install the resulting skill locally or publish it to the Agentic Hub so other teams can run the same workflow.

Can agent calls run in parallel?

Yes. SwarmFlow has parallel(), pipeline(), and map_parallel() (alias pmap()) operators for fan-out work. This tutorial uses sequential awaits because each step needs the previous step's output.

How do I add a human approval step between phases?

Use the human() operator for a single approval, confirmation, or choice, or human_session() for a multi-turn exchange. The run waits at that node, and you reply from the SwarmFlow tree view in the Web UI.

How do I cap token spend for a Swarm?

Set swarmflow_budget under modes.team.jiuwen_team in ~/.jiuwenswarm/config/config.yaml and restart the backend, or use /swarmflow on --budget in the TUI. The Web UI can switch SwarmFlow on, but it doesn't expose the budget setting. When a run exhausts its budget, it ends as failed and can't be resumed.

How do I pause or stop a run that's already going?

Open the SwarmFlow tree view in the Web UI and pause, resume, or stop the individual run. Paused runs don't resume on their own. After you send a message, the Leader decides whether to resume them.


Top comments (0)