I built a 17-agent AI swarm on my phone — here's how
Okay, buckle up. This is a bit of a wild ride. For the past few weeks, I've been obsessively working on something I’m calling “IronVision Nexus” – a multi-agent AI system running entirely on my Android phone. Not a server in sight, no cloud dependency. Just 17 little AI brains buzzing around, interacting and attempting to solve a basic task: collaborative object recognition and scene description. It's been a learning curve, to say the least, and I want to share the journey, the challenges, and how I made it happen.
Why a Phone? And Why a Swarm?
Good questions! The “why phone” part stems from a fascination with edge computing. I wanted to push the boundaries of what’s possible locally. We're so reliant on the cloud, and the idea of having self-contained, privacy-focused AI felt compelling. Plus, it’s a cool technical challenge.
The “swarm” aspect comes from the desire to explore emergent behavior. Individual AI agents are relatively simple. But when you throw a bunch of them together, each with their own slightly different goals and perspectives, interesting things happen. It's about leveraging collective intelligence, even in a limited scope.
The Architecture: Tasky & TinyLLM
The core of IronVision Nexus is built on two main components: Tasky (my lightweight task management system) and a quantized version of TinyLLM.
-
TinyLLM: This is where the “thinking” happens. I'm using a 1.1B parameter version of TinyLLM, quantized down to 4-bit precision using
llama.cpp. This is crucial for running on a phone. The full model would be impossible. I chose TinyLLM because it’s designed to be small and relatively fast, while still exhibiting basic language understanding. - Tasky: This is my custom system for orchestrating the agents. Think of it as a task queue and communication bus. Each agent pulls tasks from the queue, processes them, and then posts their results back, potentially creating new tasks for other agents.
The Agents: Specialization is Key
I didn't want 17 copies of the same agent. That would be… boring. Instead, each agent is specialized for a particular role. Here's a breakdown:
-
Image Captioners (x4): These agents receive raw image data (from the phone's camera) and generate initial textual descriptions. Prompt:
“Describe the image in as much detail as possible.” - Object Detectors (x4): Leveraging a pre-trained MobileNetV2 model (using TensorFlow Lite), these agents identify objects within the image and report their bounding boxes and confidence levels.
-
Scene Understanding Agents (x3): These take the outputs from the Image Captioners and Object Detectors, and attempt to build a higher-level understanding of the scene. Prompt:
“Based on the image description and detected objects, what is happening in this scene?” -
Detail Enhancement Agents (x3): These agents focus on refining descriptions. They take initial descriptions and ask clarifying questions, or look for additional details. Prompt:
“Expand on the previous description. What specific features stand out?” - Consensus Agent (x3): These agents are the arbiters. They receive outputs from all other agents and attempt to create a final, consensus-driven description of the scene. This is where the “swarm intelligence” really comes into play.
Code Snippets: A Glimpse Under the Hood
Let’s look at a simplified Python snippet illustrating how Tasky manages an agent's workload. This is running via Pydroid 3 on the phone:
import queue
class Agent:
def __init__(self, agent_id, role):
self.agent_id = agent_id
self.role = role
def process_task(self, task_data):
# Simulated task processing – in reality, this would interact with TinyLLM
if self.role == "Image Captioner":
description = f"Image described by agent {self.agent_id}: {task_data['image']}"
elif self.role == "Object Detector":
description = f"Object detected by agent {self.agent_id}: {task_data['image']}"
else:
description = f"Agent {self.agent_id} processed: {task_data['image']}"
return {"agent_id": self.agent_id, "description": description}
class Tasky:
def __init__(self):
self.task_queue = queue.Queue()
self.agents = []
def add_agent(self, agent):
self.agents.append(agent)
def submit_task(self, task_data):
self.task_queue.put(task_data)
def process_tasks(self):
while not self.task_queue.empty():
task = self.task_queue.get()
results = []
for agent in self.agents:
results.append(agent.process_task(task))
return results
# Example Usage
tasky = Tasky()
#Create agents
agents = [Agent(i, "Image Captioner") for i in range(4)]
tasky.agents.extend(agents)
tasky.submit_task({"image": "A cat sitting on a mat"})
results = tasky.process_tasks()
print(results)
This is a very simplified example. The actual implementation handles communication between agents, manages resource allocation (memory is a real constraint!), and integrates with the TinyLLM inference engine.
Challenges & Optimizations
Running this on a phone is not without its hurdles:
- Memory Management: The biggest challenge. 4-bit quantization helps, but 17 TinyLLM instances still eat memory. I'm aggressively caching results and unloading agents when they're not actively processing tasks.
-
Computational Power: The phone’s CPU is the bottleneck. I’m using TensorFlow Lite for the MobileNetV2 model and heavily optimizing TinyLLM inference through
llama.cppparameters like-ngl(layer offloading to the GPU – surprisingly effective!). - Inter-Process Communication (IPC): Python’s multiprocessing can be slow. I'm experimenting with shared memory to reduce overhead for passing data between agents.
- Battery Life: This thing drains battery. I'm implementing a “sleep mode” where inactive agents are paused to conserve power.
Results & Future Directions
The results are… promising, considering the constraints. The swarm does converge on a reasonable description of
Top comments (0)