I built a 17-agent AI swarm on my phone — here's how
Okay, buckle up. This is a bit of a weird one. For the last few weeks, I’ve been obsessed with the idea of running a multi-agent system – a swarm of AI agents – entirely on my phone. Not just accessing a cloud-based API, but genuinely running the processing locally. It felt like a fascinating challenge, a bit of a "can it be done?" scenario. And surprisingly… it can.
I call it “IronVision Nexus”, and it currently runs 17 distinct agents, each with a specific role, all coordinating (or sometimes arguing) on my Pixel 7. This article will dive into the “how” – the tech stack, the challenges, the compromises, and a look at the code. It's not a polished product, more a proof of concept, but the results are genuinely exciting.
Why a Swarm on a Phone?
The motivation wasn't practicality (though that’s emerging). It was about pushing boundaries. Large Language Models (LLMs) are incredible, but we're increasingly reliant on cloud infrastructure. I wanted to explore:
- Local Processing: Could we achieve meaningful AI capabilities without constant network connectivity or privacy concerns?
- Edge Computing: Leveraging the growing power of mobile devices for on-device intelligence.
- Multi-Agent Coordination: Understanding how multiple smaller LLMs could work together to accomplish tasks.
- Just...because! It seemed cool.
The Tech Stack: A Surprisingly Effective Combo
Given the limitations of mobile processing power, I needed to be clever with my choices.
- Language: Python. It's the easiest for rapid prototyping and has a huge library ecosystem.
-
LLM Backend: TinyLlama 1.1B. This is key. While GPT-3.5/4 would be better quality, they're simply too large for practical on-device inference. TinyLlama, while small, is surprisingly capable when properly prompted and directed. I quantized it to 4-bit using
llama.cppto further reduce memory usage. -
Inference Engine:
llama.cpp(through its Python bindings). This library is designed for efficient LLM inference on CPUs and, crucially, supports quantization and offloading layers to the GPU on Android. - Framework: My own lightweight agent framework (details below). I didn't want the overhead of something like LangChain for this specific project.
- UI: Termux & a simple Python script for command-line interaction. No fancy GUI… yet.
Building the Agent Framework
The core of IronVision Nexus is a very basic agent framework built around a shared "blackboard" – a dictionary representing the system's knowledge and task list. Each agent operates in a loop:
- Observe: Read the blackboard for new tasks and relevant information.
- Think: Generate a response using TinyLlama, based on its role, the current blackboard state, and any relevant context.
- Act: Update the blackboard with its findings, propose new tasks, or modify existing ones.
Here’s a simplified snippet illustrating the agent loop:
class Agent:
def __init__(self, name, role, model):
self.name = name
self.role = role
self.model = model
def run(self, blackboard):
while True:
# 1. Observe
tasks = blackboard.get("tasks", [])
relevant_info = blackboard.get("knowledge", {})
# 2. Think
prompt = self.generate_prompt(tasks, relevant_info)
response = self.model.generate(prompt)
# 3. Act
self.process_response(response, blackboard)
time.sleep(0.5) # Add a small delay to avoid hogging the CPU
def generate_prompt(self, tasks, knowledge):
# This is where the magic happens - crafting the prompt based on the agent's role
# and the current context
prompt = f"You are {self.role}. Your name is {self.name}.\n"
prompt += "Current Tasks: " + str(tasks) + "\n"
prompt += "Current Knowledge: " + str(knowledge) + "\n"
prompt += "What should you do next? Be concise.\n"
return prompt
def process_response(self, response, blackboard):
# Parse the response and update the blackboard accordingly
print(f"{self.name}: {response}")
# ... (logic to parse response and update the blackboard)
pass
This is a highly simplified example. The generate_prompt function is the crucial piece, dynamically creating a prompt tailored to the agent's role.
The 17 Agents: A Microscopic Society
So, what kind of agents are we talking about? Here's a breakdown of the current cast:
- The Planner: Responsible for breaking down goals into actionable tasks.
-
The Researcher: Uses a simplified search API (implemented with Python's
requestslibrary, making limited calls to search engines when connectivity is available) to gather information. - The Summarizer: Condenses lengthy information into concise summaries.
- The Critic: Reviews and critiques the work of other agents, identifying errors or inconsistencies.
- The Memory Keeper: Maintains a long-term knowledge base (stored as a simple JSON file).
- The Prioritizer: Determines the order in which tasks should be executed.
- The Fact Checker: Attempts to verify information found by the Researcher.
- The Translator: Translates between English and (currently) Spanish.
- The Storyteller: Creates narrative summaries of the system's activity.
- And 7 others: Covering roles like Sentiment Analyzer, Code Explainer, Pattern Recognizer, etc.
Challenges and Compromises
It wasn't all smooth sailing. Here are some major hurdles:
- Performance: TinyLlama 1.1B is slow on a phone CPU. Quantization helps, but inference times are still significant. I’ve optimized prompts to be as concise as possible and limited the maximum output length.
- Memory Constraints: Even with quantization, the model takes up a considerable amount of RAM. I’ve had to aggressively manage memory usage and sometimes deal with out-of-memory errors.
- Prompt Engineering: Getting the agents to cooperate effectively requires extremely careful prompt engineering. It's a delicate balancing act between providing enough context and keeping the prompts short enough to fit within the LLM's context window.
- Debugging: Debugging a multi-agent system running on a phone is… challenging. Relying heavily on logging and careful observation of
Top comments (0)