I built a 17-agent AI swarm on my phone — here's how
Okay, buckle up. This is a bit of a journey. For the last few weeks, I’ve been obsessed with the idea of running a fully functional, multi-agent AI system… on my phone. Not just an AI, but a swarm of them, each with a specific role and able to interact, learn, and (hopefully) achieve a common goal. Sounds crazy? Maybe. But it's done. And here's how.
Why a phone? Why a swarm?
Let’s address the elephant in the room. Why limit myself to the processing power of a phone? The answer is two-fold: portability and accessibility. I wanted something I could tinker with anywhere, without relying on cloud connections or expensive hardware. It's a fascinating constraint that forced me to get really creative with optimization.
As for the swarm, I’m fascinated by emergent behavior. The idea that simple agents, following simple rules, can collectively achieve complex results is incredibly powerful. Think ants building a colony or birds flocking. I wanted to create something analogous in the digital space. Plus, it’s just…cool.
The Architecture: Agent Framework & Language Choice
The core of the project relies on a framework I developed (and continue to refine) called FractalMesh. It’s essentially a lightweight agent management system built in Python and leveraging the llama.cpp library. Why Python? It's mature, well-documented, and has fantastic libraries. And llama.cpp? This is key. It allows running large language models (LLMs) – in my case, quantized versions of Mistral-7B – locally on devices with limited resources.
Here's a simplified illustration:
[Phone CPU/GPU] <-----> [FractalMesh Framework] <-----> [17 AI Agents]
|
v
[LLM (Mistral-7B Quantized)]
Each agent within FractalMesh is defined by:
- Role: A description of its function (e.g., “Research Assistant,” “Data Analyst,” “Creative Writer,” “Code Reviewer”).
- Memory: A small, persistent store (using SQLite) to retain information between interactions.
- Tools: Functions the agent can access (e.g., web search, file I/O, basic calculations).
- Prompt Template: A carefully crafted prompt that guides the LLM's response based on the agent's role, current context, and memory.
The Agents: A Diverse Team
I settled on 17 agents, categorized into three main groups:
-
Information Gatherers (5 Agents): These agents are responsible for acquiring data. Roles include "Web Researcher," "News Aggregator," "Academic Paper Summarizer," "Social Media Monitor," and "Fact Checker." They use Python’s
requestsandBeautifulSoup4libraries to scrape data and parse information. - Processors & Analysts (7 Agents): These agents process the information gathered by the first group. Roles include "Data Analyst," "Trend Identifier," "Sentiment Analyzer," "Risk Assessor," “Summary Generator,” “Report Writer," and "Pattern Recognizer." They use basic statistical functions and LLM-powered analysis.
- Output & Interaction (5 Agents): These agents handle the output and interaction with the user. Roles include "Creative Writer," "Code Generator," "Email Composer," "Presentation Creator," and “Executive Summarizer”. They focus on generating human-readable results.
Code Snippets: Agent Definition & Interaction
Here’s a simplified example of how an agent is defined within the FractalMesh framework:
class WebResearcherAgent:
def __init__(self, framework):
self.framework = framework
self.role = "Web Researcher"
self.memory = [] # Initialize agent memory
self.prompt_template = """
You are a meticulous web researcher. Your task is to find relevant information about the given query.
Remember previous queries and results.
Query: {query}
Previous Interactions: {memory}
Response:
"""
def execute(self, query):
# Simulate web search (replace with actual scraping logic)
search_results = f"Results for '{query}': [link1, link2, link3]"
response = self.framework.llm_predict(self.prompt_template.format(query=query, memory=self.memory), model="mistral-7b-instruct")
self.memory.append(f"Query: {query}, Results: {search_results}") # Update agent memory
return response
Interaction between agents is managed by the FractalMesh framework. Agents pass messages and data to each other based on a defined workflow. For example, the "Web Researcher" might pass search results to the "Data Analyst" for processing. This messaging system relies heavily on JSON for structured data exchange.
Challenges & Optimization: Fitting it all on a Phone
This wasn’t without its hurdles. My phone (a Pixel 7 Pro) has respectable specs, but it’s still not a server-grade machine. Here’s what I had to tackle:
- LLM Quantization: Running a full Mistral-7B model would be impossible. I used a 4-bit quantized version of the model which significantly reduces memory usage but slightly impacts quality.
- Memory Management: Each agent has limited memory. I implemented a "forgetting" mechanism where older interactions are gradually removed from memory to prevent overload.
-
Concurrency: I used Python’s
asynciolibrary to enable concurrent execution of agents. This allowed multiple agents to operate “simultaneously,” maximizing CPU utilization. However, overdoing it leads to thermal throttling, so it required careful balancing. - Prompt Engineering: Designing effective prompts is crucial. Clear, concise prompts help the LLM stay on task and generate relevant responses.
- Battery Life: This is a big one. Running LLMs drains battery fast. I implemented a "sleep" mode for agents when they're not actively processing, and I'm constantly optimizing the code for efficiency.
The Results: Is it Useful?
So, after all this effort, does it actually work? Surprisingly, yes! It's slow, definitely slower than a cloud-based solution, but it's functional. I’ve been using it to research topics, generate summaries, and even write basic code snippets. The emergent behavior is fascinating. For example, when asked to “investigate the potential impact of AI on the automotive industry”, the swarm independently identified key trends like self-driving technology, supply chain disruptions, and the changing role of automotive
Top comments (0)