I built a 17-agent AI swarm on my phone — here's how
Okay, buckle up. This is a bit of a rabbit hole, but a really fun one. I’ve spent the last few weeks building and running a swarm of 17 independent AI agents, all contained within and operating on my phone. Yes, you read that right. No cloud servers, no expensive GPUs. Just a Pixel 7 Pro and a lot of Python.
This wasn’t about solving world hunger (though, hey, maybe someday!), it was about pushing the boundaries of what’s possible with on-device AI and exploring emergent behaviour within a limited computational environment. It’s a proof-of-concept for what I’m calling “FractalMesh,” a system designed for decentralized, localized AI processing.
Why on-device?
The current AI landscape is dominated by massive models running in the cloud. That's fantastic for powerful applications, but comes with downsides: latency, privacy concerns, and reliance on network connectivity. On-device AI, however, offers a lot of promise. It’s faster, more private, and can operate offline. My phone is essentially a mini-supercomputer just waiting to be tapped. The challenge? Squeezing 17 brains into it.
The Architecture: Agents and the Coordinator
The system is built around a central “Coordinator” agent and 17 “Worker” agents. Think of it like a hive mind.
- Coordinator: This is the orchestrator. It defines the task, divides it into sub-tasks, distributes those sub-tasks to the Workers, collects their responses, and synthesizes them into a final output. It doesn't do the core processing, it manages it.
-
Workers: These are the specialists. Each agent is designed with a specific, focused role. In this iteration, I've got agents for:
- Sentiment Analysis (x3): Evaluate text for emotional tone.
- Topic Extraction (x3): Identify key themes within text.
- Keyword Generation (x3): Brainstorm related keywords.
- Text Summarization (x3): Condense text into concise summaries.
- Question Answering (x3): Answer questions based on provided text.
- Translation (x2): Translate text between English and Spanish.
Tech Stack & Optimizations: Python, llama.cpp, and a LOT of patience
The core of the project is Python, running via a custom script I’ve packaged with Termux on Android. Termux provides a Linux-like environment on Android allowing me to run Python and install packages. However, running full-blown transformer models (even smaller ones) on a phone isn’t feasible without serious optimization.
This is where llama.cpp came in. llama.cpp is a fantastic project that allows you to run large language models (LLMs) quantized for CPU usage. This is crucial for phone performance. I used a quantized version of the TinyLlama model (1.1B parameters, Q4_K_M quantization) for each agent. It's not GPT-4, but it's surprisingly capable for its size.
Here’s a simplified snippet demonstrating how an agent (let's say a Sentiment Analyzer) is instantiated and used:
from llama_cpp import Llama
import time
class SentimentAgent:
def __init__(self, model_path):
self.llm = Llama(model_path=model_path, n_ctx=512, n_threads=4) # Threads adjusted for phone cores
def analyze_sentiment(self, text):
prompt = f"Analyze the sentiment of the following text: '{text}'. Respond with 'Positive', 'Negative', or 'Neutral'."
output = self.llm(prompt, max_tokens=20, stop=["\n"], echo=False)
return output["choices"][0]["text"].strip()
# Example usage
sentiment_agent = SentimentAgent("/sdcard/models/tinyllama-q4_k_m.gguf") # Path on phone
text_to_analyze = "This is an absolutely amazing project! I love it!"
sentiment = sentiment_agent.analyze_sentiment(text_to_analyze)
print(f"Sentiment: {sentiment}")
Key Optimizations:
- Quantization: Using Q4_K_M quantization reduced the model size significantly, allowing me to fit more agents.
-
Context Length: Limiting the context window (
n_ctxinllama.cpp) to 512 tokens helped keep memory usage down. -
Thread Control:
n_threadswas set to 4, matching my phone's core count. Experimenting with this value is crucial for performance. - Batching: While not fully implemented yet, I'm planning to batch requests to the same agent type to improve throughput.
- Prioritization: The Coordinator prioritizes tasks and agents based on estimated complexity. More complex tasks get more resources (longer processing time).
The Coordinator's Role: Orchestrating the Swarm
The Coordinator is where the real magic happens. It receives an initial text input, breaks it down, and assigns portions to the appropriate Workers. Here's a simplified view of the Coordinator's logic:
import threading
class Coordinator:
def __init__(self, agents):
self.agents = agents # Dictionary of agent types (sentiment, topic, etc.)
self.results = {}
def process_text(self, text):
self.results = {}
threads = []
# Divide the task & assign to workers
for agent_type, agent_list in self.agents.items():
for agent in agent_list:
thread = threading.Thread(target=self.run_agent, args=(agent, text, agent_type))
threads.append(thread)
thread.start()
for thread in threads:
thread.join()
# Synthesize results (currently just prints them)
print("--- Results ---")
for agent_type, results in self.results.items():
print(f"{agent_type}: {results}")
def run_agent(self, agent, text, agent_type):
# Run the agent's function
if agent_type == "sentiment":
result = agent.analyze_sentiment(text)
elif agent_type == "topic":
result = agent.extract_topics(text)
# ... other agents ...
self.results[agent_type] = result
This example uses threading to run agents concurrently. The actual implementation is more complex, handling error
Top comments (0)