I built a 17-agent AI swarm on my phone — here's how
Okay, buckle up. This is going to be a bit of a deep dive. For the past few weeks, I’ve been obsessed with the idea of running a surprisingly complex AI system entirely on my smartphone. Not just a single model, but a swarm of interacting agents. Sounds crazy? Maybe. But I did it, and I'm going to walk you through how.
The goal wasn't to build the next ChatGPT. It was to explore the limits of on-device AI, specifically focusing on emergent behavior from a distributed system. The end result? A 17-agent system I've dubbed "Nexus", focused on information gathering, analysis, and a rudimentary form of "decision-making" based on collective assessment.
Why a phone? And why a swarm?
Why not a powerful server? Well, for one, the challenge. The constraints of a mobile device – limited processing power, RAM, and battery life – force you to be incredibly efficient. Plus, the potential applications are huge: truly private AI, off-grid operation, and novel interfaces are just a few.
As for the swarm, it's all about leveraging the power of distributed intelligence. Each agent doesn't need to be individually brilliant. The intelligence comes from how they interact, share information, and collectively arrive at a result. Think of it like ant colonies – individual ants aren't planning world domination, but the colony is incredibly effective.
The Tech Stack: Python, PyTorch Mobile, and a lot of Optimization
My stack is built around:
- Python: Because, well, it's Python. I'm comfortable with it, and it has excellent libraries.
- PyTorch Mobile: This is the game changer. It allows you to run PyTorch models directly on mobile devices. It was key to making this viable.
- Kivy: A Python framework for building cross-platform mobile apps (Android and iOS). It's not the fastest option, but it's quick to prototype with.
- SQLite: For lightweight data storage of agent memories and communication logs.
- A LOT of Optimization: This is where the real work happened.
The Agents: Roles and Responsibilities
My Nexus swarm consists of 17 agents, each with a specific, relatively simple role. Here's a breakdown:
-
3 x Scrapers: These agents pull data from pre-defined URLs (news sites, APIs, etc.). They're the "eyes" of the system. I used
requestslibrary for this. - 4 x Summarizers: Take the raw text from the Scrapers and generate concise summaries using a pre-trained DistilBERT model (converted to PyTorch Mobile format).
- 4 x Sentiment Analyzers: Analyze the sentiment (positive, negative, neutral) of the summarized text. Another DistilBERT variant was used here.
- 2 x Trend Detectors: Look for recurring themes and keywords across the summaries.
- 2 x "Contextualizers": These agents add relevant background information to the analyzed data, pulling from a locally stored knowledge base (SQLite). Think of them as providing context.
- 2 x "Decision-Makers": The final stage. These agents weigh the data from all other agents and generate a "report" – a concise assessment of the situation.
Code Snippet: Simplified Summarizer Agent
This is a drastically simplified example, but it shows the core structure of an agent:
import torch
from transformers import DistilBertTokenizer, DistilBertForSequenceClassification
class SummarizerAgent:
def __init__(self, model_path):
self.tokenizer = DistilBertTokenizer.from_pretrained(model_path)
self.model = DistilBertForSequenceClassification.from_pretrained(model_path)
self.model.eval() # crucial for inference
def summarize(self, text):
inputs = self.tokenizer(text, return_tensors="pt", truncation=True, max_length=512)
with torch.no_grad(): # important to disable gradient calculation
outputs = self.model(**inputs)
# Simplified summary extraction - in reality, you'd use more sophisticated methods
predicted_class = torch.argmax(outputs.logits, dim=-1).item()
if predicted_class == 1:
summary = f"Positive sentiment: {text[:100]}..."
elif predicted_class == 0:
summary = f"Negative sentiment: {text[:100]}..."
else:
summary = f"Neutral: {text[:100]}..."
return summary
Key Optimization Techniques (This is where it gets gritty)
Running 17 AI models on a phone requires serious optimization. Here's what I did:
- Quantization: PyTorch Mobile allows you to quantize your models to 8-bit integers, significantly reducing their size and improving inference speed. This was huge.
- Model Pruning: Removing unnecessary weights from the models. This reduces complexity without significant accuracy loss.
- Batching (limited): I attempted batch processing where feasible, but the phone's limited memory made this tricky.
- Efficient Data Structures: Using lists and dictionaries instead of more complex objects to minimize memory overhead.
- Background Processing: Agents run in separate threads, allowing them to work concurrently without blocking the UI.
- Aggressive Caching: Caching results of expensive operations (like tokenization) to avoid redundant calculations.
- Limited Data Scope: The Scrapers only pull a small amount of data per cycle. We’re not indexing the entire internet!
Communication: The Heart of the Swarm
The agents don't directly call each other’s methods. Instead, they communicate through a central message queue (implemented with Python lists and locks).
Each agent listens for messages tagged with specific keywords. For example, a Summarizer agent would listen for messages tagged “raw_text”. When an agent processes a message, it creates new messages containing its output, tagged appropriately. This loosely-coupled architecture allows for flexibility and scalability.
Challenges & Limitations
This wasn’t a smooth ride. The biggest hurdles were:
- Memory Management: The phone's RAM is a constant constraint. Careful memory allocation and garbage collection are crucial.
- Battery Life: Running multiple AI models drains the battery quickly. Optimization is paramount, but you're still limited.
- Model Conversion: Converting PyTorch models to PyTorch Mobile format isn’t always straightforward.
- Debugging: Debugging a distributed system on a phone is… challenging.
Current State & Future Plans
Right now, Nexus is a proof of concept.
Top comments (0)