DEV Community

Samuel James Hiotis
Samuel James Hiotis

Posted on

I built a 17-agent AI swarm on my phone — here's how

I built a 17-agent AI swarm on my phone — here's how

Okay, buckle up. This is a bit of a wild ride. For the past few weeks, I've been obsessively working on something I’m calling “IronVision Nexus” – a multi-agent AI system running entirely on my Android phone. Not a server in sight, no cloud dependency. Just 17 little AI brains buzzing around, interacting and attempting to solve a basic task: collaborative object recognition and scene description. It's been a learning curve, to say the least, and I want to share the journey, the challenges, and how I made it happen.

Why a Phone? And Why a Swarm?

Good questions! The “why phone” part stems from a fascination with edge computing. I wanted to push the boundaries of what’s possible locally. We're so reliant on the cloud, and the idea of having self-contained, privacy-focused AI felt compelling. Plus, it’s a cool technical challenge.

The “swarm” aspect comes from the desire to explore emergent behavior. Individual AI agents are relatively simple. But when you throw a bunch of them together, each with their own slightly different goals and perspectives, interesting things happen. It's about leveraging collective intelligence, even in a limited scope.

The Architecture: Tasky & TinyLLM

The core of IronVision Nexus is built on two main components: Tasky (my lightweight task management system) and a quantized version of TinyLLM.

  • TinyLLM: This is where the “thinking” happens. I'm using a 1.1B parameter version of TinyLLM, quantized down to 4-bit precision using llama.cpp. This is crucial for running on a phone. The full model would be impossible. I chose TinyLLM because it’s designed to be small and relatively fast, while still exhibiting basic language understanding.
  • Tasky: This is my custom system for orchestrating the agents. Think of it as a task queue and communication bus. Each agent pulls tasks from the queue, processes them, and then posts their results back, potentially creating new tasks for other agents.

The Agents: Specialization is Key

I didn't want 17 copies of the same agent. That would be… boring. Instead, each agent is specialized for a particular role. Here's a breakdown:

  • Image Captioners (x4): These agents receive raw image data (from the phone's camera) and generate initial textual descriptions. Prompt: “Describe the image in as much detail as possible.”
  • Object Detectors (x4): Leveraging a pre-trained MobileNetV2 model (using TensorFlow Lite), these agents identify objects within the image and report their bounding boxes and confidence levels.
  • Scene Understanding Agents (x3): These take the outputs from the Image Captioners and Object Detectors, and attempt to build a higher-level understanding of the scene. Prompt: “Based on the image description and detected objects, what is happening in this scene?”
  • Detail Enhancement Agents (x3): These agents focus on refining descriptions. They take initial descriptions and ask clarifying questions, or look for additional details. Prompt: “Expand on the previous description. What specific features stand out?”
  • Consensus Agent (x3): These agents are the arbiters. They receive outputs from all other agents and attempt to create a final, consensus-driven description of the scene. This is where the “swarm intelligence” really comes into play.

Code Snippets: A Glimpse Under the Hood

Let’s look at a simplified Python snippet illustrating how Tasky manages an agent's workload. This is running via Pydroid 3 on the phone:

import queue

class Agent:
  def __init__(self, agent_id, role):
    self.agent_id = agent_id
    self.role = role

  def process_task(self, task_data):
    # Simulated task processing – in reality, this would interact with TinyLLM
    if self.role == "Image Captioner":
      description = f"Image described by agent {self.agent_id}: {task_data['image']}"
    elif self.role == "Object Detector":
      description = f"Object detected by agent {self.agent_id}: {task_data['image']}"
    else:
      description = f"Agent {self.agent_id} processed: {task_data['image']}"

    return {"agent_id": self.agent_id, "description": description}

class Tasky:
  def __init__(self):
    self.task_queue = queue.Queue()
    self.agents = []

  def add_agent(self, agent):
    self.agents.append(agent)

  def submit_task(self, task_data):
    self.task_queue.put(task_data)

  def process_tasks(self):
    while not self.task_queue.empty():
      task = self.task_queue.get()
      results = []
      for agent in self.agents:
        results.append(agent.process_task(task))
      return results

# Example Usage
tasky = Tasky()
#Create agents
agents = [Agent(i, "Image Captioner") for i in range(4)]
tasky.agents.extend(agents)

tasky.submit_task({"image": "A cat sitting on a mat"})
results = tasky.process_tasks()
print(results)
Enter fullscreen mode Exit fullscreen mode

This is a very simplified example. The actual implementation handles communication between agents, manages resource allocation (memory is a real constraint!), and integrates with the TinyLLM inference engine.

Challenges & Optimizations

Running this on a phone is not without its hurdles:

  • Memory Management: The biggest challenge. 4-bit quantization helps, but 17 TinyLLM instances still eat memory. I'm aggressively caching results and unloading agents when they're not actively processing tasks.
  • Computational Power: The phone’s CPU is the bottleneck. I’m using TensorFlow Lite for the MobileNetV2 model and heavily optimizing TinyLLM inference through llama.cpp parameters like -ngl (layer offloading to the GPU – surprisingly effective!).
  • Inter-Process Communication (IPC): Python’s multiprocessing can be slow. I'm experimenting with shared memory to reduce overhead for passing data between agents.
  • Battery Life: This thing drains battery. I'm implementing a “sleep mode” where inactive agents are paused to conserve power.

Results & Future Directions

The results are… promising, considering the constraints. The swarm does converge on a reasonable description of

Top comments (0)