I built a 17-agent AI swarm on my phone — here's how
Okay, buckle up. This is going to be a deep dive. For the past few weeks, I’ve been obsessed with a seemingly ludicrous idea: could I run a multi-agent AI system, a swarm even, directly on my phone? Not just simple chatbots, but a network of interacting agents with distinct roles, communicating and collaborating to achieve a (relatively) complex goal. The answer, surprisingly, is… yes. And it's been a wild ride.
This wasn't about raw processing power (my phone is a perfectly respectable Pixel 7, but it's not a data center). It was about clever architecture, resource management, and exploiting the advancements in on-device machine learning. Here's how I built my 17-agent AI swarm, affectionately dubbed “Project Nightingale,” running entirely on Android.
The Goal: Decentralized Information Gathering & Summarization
Before getting into the tech, let's define the goal. I wanted a system that could autonomously gather information about a specific topic (currently, emerging trends in web3 security), analyze it from multiple angles, and present a concise, summarized report. Instead of one massive LLM trying to do everything, I envisioned a swarm of specialized agents:
- Scrapers (4 agents): Responsible for web scraping from specific sources (news sites, blogs, Twitter, Reddit).
- Analyzers (6 agents): These agents perform sentiment analysis, topic extraction, and key phrase identification on the scraped content. They’re specialized – some focused on technical details, others on market impact, etc.
- Validators (3 agents): These agents assess the credibility of the sources and the information. They cross-reference data and flag potential misinformation.
- Summarizers (2 agents): The final layer, taking the analyzed data and producing a cohesive summary report.
- Coordinator (2 agents): Two redundant coordinator agents manage the workflow, distribute tasks, and collect results. This redundancy is key for robustness.
The Tech Stack: Python, llama.cpp, and a sprinkle of Magic
My entire system is built using Python. It’s the most flexible option for prototyping and accessing the necessary libraries. The core of each agent leverages llama.cpp – a fantastic library that allows running Large Language Models (LLMs) locally, even on devices with limited resources. I chose a quantized version of Mistral 7B (specifically Q4_K_M) for its balance of performance and size.
Here’s a simplified example of how an analyzer agent initializes the LLM using llama.cpp (using the python bindings):
from llama_cpp import Llama
llm = Llama(model_path="./models/mistral-7b-instruct-v0.1.Q4_K_M.gguf", n_ctx=2048) #Adjust n_ctx based on phone RAM
Key components:
- Agent Class: A base class defining common functionalities like message handling, task execution, and reporting.
- Task Queue (Redis): A Redis server running locally on the phone (using a lightweight Redis implementation) acts as a task queue. The coordinator agents push tasks (e.g., "scrape this URL") onto the queue, and worker agents pull tasks as they become available. This decouples the agents and allows for asynchronous operation.
- Data Storage (SQLite): SQLite is used for persistent storage of scraped data, analysis results, and agent logs. It’s lightweight and ideal for on-device use.
- Communication (Simple Text-Based Protocol): Agents communicate via a simple text-based protocol over TCP sockets. This keeps overhead low and avoids complex serialization/deserialization.
The Architecture: A Decentralized Mesh
This isn't a hierarchical system. While the coordinators initiate tasks, agents aren't strictly ordered. An analyzer agent, for example, can directly request additional information from a scraper if it needs clarification.
Here's a snippet of the agent communication loop:
import socket
class Agent:
def __init__(self, agent_id, role):
self.agent_id = agent_id
self.role = role
self.socket = socket.socket(socket.AF_INET, socket.SOCK_STREAM)
self.socket.bind(('localhost', 5000 + agent_id)) #Assign unique ports
self.socket.listen(5)
def run(self):
print(f"Agent {self.agent_id} ({self.role}) listening...")
while True:
conn, addr = self.socket.accept()
with conn:
data = conn.recv(1024).decode()
response = self.process_message(data)
conn.sendall(response.encode())
Each agent listens on a unique port. The coordinator agents act as central hubs, but agents can also directly connect to each other when needed. This mesh network provides resilience – if one agent fails, others can often compensate.
Challenges and Optimizations
Running 17 LLMs, even quantized ones, on a phone is challenging. Here's what I tackled:
- Memory Management: Crucial. I used context size (n_ctx) strategically. Analyzers dealing with longer texts used larger contexts, while scrapers used smaller ones. I also implemented aggressive garbage collection.
- CPU Throttling: Mobile CPUs throttle under sustained load. I introduced sleep intervals between tasks and limited the number of concurrent LLM inferences.
- Battery Life: This thing is a battery hog. I added a battery monitoring module that automatically reduces the number of active agents when battery levels are low.
- Context Switching: With so many agents running concurrently, context switching overhead became significant. Using Redis for task queuing and keeping agent code streamlined helped mitigate this.
Code Snippets & Key Functionality
Let’s look at a simplified example of a scraper agent:
import requests
from bs4 import BeautifulSoup
def scrape_website(url):
try:
response = requests.get(url, timeout=10)
response.raise_for_status() # Raises HTTPError for bad responses (4xx or 5xx)
soup = BeautifulSoup(response.content, 'html.parser')
text = soup.get_text()
return text[:4000] #Limit text length - LLM context window
except requests.exceptions.RequestException as e:
print(f"Error scraping {url}: {e}")
return ""
And a simplified example of how an Analyzer might use the LLM to perform sentiment analysis:
python
def analyze_sentiment(text):
prompt = f
Top comments (0)