DEV Community

Samuel James Hiotis
Samuel James Hiotis

Posted on

I built a 17-agent AI swarm on my phone — here's how

I built a 17-agent AI swarm on my phone — here's how

Okay, buckle up. This is a weird one. For the past few weeks, I’ve been obsessively tinkering with a project that sounds like science fiction: running a swarm of 17 independent AI agents entirely on my phone. Not connecting to a cloud service, not offloading to a server – everything is local, leveraging the surprisingly capable hardware we carry around in our pockets.

It’s not about building AGI. It’s about pushing the boundaries of what's possible with edge AI and exploring emergent behavior. And honestly, it’s been a lot of fun.

Why a Swarm? And Why on a Phone?

The idea came from a fascination with swarm intelligence. Think ant colonies, bee hives, flocks of birds. Complex behaviors arise from simple individual rules. I wanted to simulate that with AI agents. Each agent would have a limited scope of perception and a simple set of goals, but together, they’d (hopefully) create something interesting.

The phone constraint was deliberate. It forces you to be incredibly efficient. Cloud-based LLMs are powerful, but relying on them removes a huge part of the learning experience – and the potential for truly independent, always-on AI. Plus, the privacy implications of sending everything to a server are… less than ideal. I wanted something that stayed mine.

The Tech Stack

This wasn’t a Python-and-PyTorch kind of project. That would be… challenging to optimize for a mobile device. Instead, I opted for:

  • Godot Engine: This is the heart of the project. Godot is a free and open-source game engine, but it’s surprisingly versatile for non-game applications. It’s lightweight, runs exceptionally well on mobile, and has a robust scripting language (GDScript) that's similar to Python.
  • GDScript: Godot’s scripting language. It's easy to learn and provides direct access to the engine’s APIs.
  • TinyLLM: This is where things get really interesting. TinyLLM is a quantized language model that can run on CPUs, even relatively low-powered ones. I’m using a 1.3B parameter model, quantized down to 4-bit, making it significantly smaller and faster.
  • Local AI Inference: I’m using Godot’s process calling capabilities to execute the TinyLLM inference. Essentially, Godot spawns a Python process that runs the TinyLLM code and returns the results.
  • Vector Database (ChromaDB - Local): Each agent maintains a short-term memory using a local ChromaDB instance. This is crucial for context and preventing the agents from repeating themselves endlessly.

The Agents: Roles and Responsibilities

The swarm consists of 17 agents, each with a specific role and a unique prompt that defines its "personality" and goals. Here's a breakdown of a few:

  • The Observer (x2): Scans the environment (simulated as a series of text descriptions). Reports observations to the Knowledge Base.
  • The Knowledge Base (x1): Stores and organizes information gathered by the Observers. Maintains a summarized “world state.”
  • The Planner (x1): Formulates short-term goals based on the world state.
  • The Executor (x3): Carries out the plans outlined by the Planner. These agents initiate actions.
  • The Analyst (x2): Analyzes the results of Executions, identifying successes and failures.
  • The Critic (x2): Evaluates the actions of other agents, providing constructive feedback.
  • The Storyteller (x1): Attempts to create a narrative based on the current world state.
  • The Randomizer (x3): Introduces controlled chaos to prevent the swarm from getting stuck. They occasionally suggest unexpected actions or questions.
  • The Questioner (x2): Probes the Knowledge Base for information, driving exploration.

A Code Snippet: Agent Interaction (GDScript)

Here’s a simplified example of how one agent (The Questioner) interacts with the Knowledge Base:

# Questioner.gd

extends Node

var knowledge_base_address = "http://localhost:5000/query" # Address of the Knowledge Base (Python process)

func _ready():
    randomize()
    call_deferred("ask_question")

func ask_question():
    var question_topics = ["current location", "recent events", "agent activity", "environmental analysis"]
    var topic = question_topics[randi() % question_topics.size()]
    var question = "What is the current status of " + topic + "?"

    var request = HTTPRequest.new()
    request.url = knowledge_base_address
    request.method = HTTPRequest.METHOD_POST
    request.set_body({"query": question})

    request.request_completed.connect(_on_request_completed)
    request.request_failed.connect(_on_request_failed)

    request.request()

func _on_request_completed(result):
    print("Questioner: Received response: ", result.result)
    # Process the response here (e.g., update internal state)
    call_deferred("ask_question") # Loop to ask another question

func _on_request_failed(error):
    print("Questioner: Error: ", error)
Enter fullscreen mode Exit fullscreen mode

This is a simplification, of course. Error handling, context management, and response parsing are all handled in more detail. The knowledge_base_address points to a Flask API server running in a separate process that handles TinyLLM inference and ChromaDB lookups.

The Python Backend (Knowledge Base – Flask API)

The Flask server receives the questions, queries the ChromaDB vector database using TinyLLM for semantic similarity search, and returns the most relevant answer. Here’s a glimpse of the API endpoint:


python
# app.py (Flask server)

from flask import Flask, request, jsonify
from chromadb.config import Settings
import chromadb
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch

app = Flask(__name__)

# Initialize ChromaDB
chroma_client = chromadb.Client(Settings(
    chroma_db_impl="duckdb+parquet",
    persist_directory="db" # On-device storage
))

collection = chroma_client.get_or_create_collection(name="agent_memory")

# Initialize TinyLLM
tokenizer = AutoTokenizer.from_pretrained("TinyLlama/TinyLlama-1.1B-Chat-v1.0")
model = AutoModelForCausalLM.from_pretrained("TinyL
Enter fullscreen mode Exit fullscreen mode

Top comments (0)