I built a 17-agent AI swarm on my phone — here's how
Okay, buckle up. This is going to be a bit of a ride. For the last few weeks, I've been obsessively working on a project that sounds straight out of a sci-fi novel: a swarm of 17 independent AI agents running entirely on my phone. Not leveraging a cloud service, not offloading computation, genuinely all happening locally. Sounds impossible? It was challenging, but here's how I did it.
Why a Swarm? And Why on a Phone?
The core idea stemmed from wanting to explore emergent behavior. I'm fascinated by the idea that complex behavior can arise from relatively simple agents interacting with each other. A swarm seemed like the perfect vehicle for this. Each agent could have a specific task, and the collective would hopefully accomplish something more significant.
As for the phone… well, constraints breed creativity. I wanted to prove it could be done. Mobile devices are incredibly powerful now, but we often default to cloud-based solutions. This was a challenge to push the limits of on-device AI and see what's achievable without constant network connectivity. Plus, the portability aspect is pretty cool.
The Tech Stack: Minimalism is Key
Given the limited resources of a mobile device (relative to a server, anyway), I needed to be incredibly selective about my tools. Here’s what I ended up with:
- Language: Python. It’s versatile, has a wealth of AI libraries, and importantly, can be run on Android via tools like Pydroid 3.
-
AI Library:
llama-cpp-python. This is a Python binding for llama.cpp, which allows running LLMs locally (including quantized models) with excellent performance. The key here was quantization – reducing the model size without significant loss of accuracy. -
Model: TinyLlama-1.1B-Chat-v1.0. This is a relatively small LLM (1.1 billion parameters) specifically designed for resource-constrained environments. I quantized it down to 4-bit using
llama.cppto further reduce its footprint. - Communication: Simple in-memory queues. Each agent publishes messages to a queue, and other agents subscribe to relevant queues. Think pub/sub, but extremely lightweight.
- Framework: No overarching framework. This was deliberately built from the ground up using basic Python classes and data structures. This allowed for maximum control and minimal overhead.
The Agents: Roles and Responsibilities
Each agent within the swarm has a defined role. This is where things get interesting. Here’s a breakdown of the 17 agents and their functions:
- The Brain (1 Agent): Acts as the central coordinator. Receives high-level goals, breaks them down into tasks, and distributes them to other agents.
- Data Collectors (3 Agents): These agents simulate gathering data from various sources (e.g., a simplified "market feed," "news headlines," "sensor readings"). They generate random, but contextually relevant, data.
- Analysis Agents (3 Agents): Analyze the data received from the Data Collectors, identifying trends, anomalies, and potential opportunities.
- Action Proposal Agents (4 Agents): Based on the analysis, these agents propose specific actions that could be taken. They also estimate the risk and reward associated with each action.
- Evaluation Agents (3 Agents): Evaluate the proposed actions, providing feedback on their feasibility, potential impact, and alignment with the overall goal.
- Execution Agent (1 Agent): Takes the most promising action (determined by consensus amongst the Evaluation Agents) and simulates its execution.
- Reporting Agent (2 Agents): Summarize the results of the execution, providing a report back to The Brain.
Code Snippets: A Glimpse Under the Hood
Let’s look at a simplified example of an agent class. This is a highly streamlined version; in practice, error handling and more robust queue management were crucial.
import time
import random
import threading
from queue import Queue
class AnalysisAgent:
def __init__(self, agent_id, data_queue, proposal_queue):
self.agent_id = agent_id
self.data_queue = data_queue
self.proposal_queue = proposal_queue
def run(self):
while True:
try:
data = self.data_queue.get(timeout=1) #Wait up to 1 sec
#Simulate Analysis
analysis_result = f"Agent {self.agent_id}: Analyzed data - Trend: {random.choice(['Up', 'Down', 'Stable'])}"
self.proposal_queue.put(analysis_result)
except Exception as e:
print(f"Analysis Agent {self.agent_id} error: {e}")
time.sleep(1)
#Example usage (simplified)
data_queue = Queue()
proposal_queue = Queue()
agent1 = AnalysisAgent(1, data_queue, proposal_queue)
thread = threading.Thread(target=agent1.run)
thread.daemon = True #Allow program to exit
thread.start()
This is a barebones example. The real code involved significantly more sophisticated prompting for the LLM, leveraging llama-cpp-python to interact with the quantized TinyLlama model. Each agent has its own thread to allow for concurrent operation.
The Challenges: Heat, Memory, and Latency
This wasn’t a smooth sail. I encountered several significant hurdles:
- Thermal Throttling: My phone heated up. Running 17 concurrent processes, even with a small LLM, is intensive. I had to implement throttling mechanisms to prevent the phone from overheating and shutting down. This involved pausing agents intermittently and strategically.
- Memory Constraints: Even with quantization, the model and the data needed to be managed carefully. I used techniques like garbage collection and optimized data structures to minimize memory usage.
- Latency: The LLM inference takes time, even on a powerful phone. The communication between agents via queues added further latency. I experimented with batching requests and optimizing the queue management to mitigate this.
- Prompt Engineering: Getting the agents to cooperate effectively required extensive prompt engineering. Each agent's prompt had to be carefully crafted to ensure it understood its role and communicated effectively with other agents.
What can it actually do?
Currently, the swarm simulates a basic “market prediction” scenario. The Data Collectors generate random price movements. The Analysis Agents try to identify trends. The Action Proposal Agents suggest buying or selling. The Evaluation Agents assess risk. The Execution Agent "simulates" the trade, and the Reporting Agents provide feedback.
It’s not
Top comments (0)