DEV Community

Samuel James Hiotis
Samuel James Hiotis

Posted on

I built a 17-agent AI swarm on my phone — here's how

I built a 17-agent AI swarm on my phone — here's how

Okay, buckle up. This is a bit of a weird one. For the last few weeks, I’ve been obsessed with the idea of running a multi-agent system – a swarm of AI agents – entirely on my phone. Not just accessing a cloud-based API, but genuinely running the processing locally. It felt like a fascinating challenge, a bit of a "can it be done?" scenario. And surprisingly… it can.

I call it “IronVision Nexus”, and it currently runs 17 distinct agents, each with a specific role, all coordinating (or sometimes arguing) on my Pixel 7. This article will dive into the “how” – the tech stack, the challenges, the compromises, and a look at the code. It's not a polished product, more a proof of concept, but the results are genuinely exciting.

Why a Swarm on a Phone?

The motivation wasn't practicality (though that’s emerging). It was about pushing boundaries. Large Language Models (LLMs) are incredible, but we're increasingly reliant on cloud infrastructure. I wanted to explore:

  • Local Processing: Could we achieve meaningful AI capabilities without constant network connectivity or privacy concerns?
  • Edge Computing: Leveraging the growing power of mobile devices for on-device intelligence.
  • Multi-Agent Coordination: Understanding how multiple smaller LLMs could work together to accomplish tasks.
  • Just...because! It seemed cool.

The Tech Stack: A Surprisingly Effective Combo

Given the limitations of mobile processing power, I needed to be clever with my choices.

  • Language: Python. It's the easiest for rapid prototyping and has a huge library ecosystem.
  • LLM Backend: TinyLlama 1.1B. This is key. While GPT-3.5/4 would be better quality, they're simply too large for practical on-device inference. TinyLlama, while small, is surprisingly capable when properly prompted and directed. I quantized it to 4-bit using llama.cpp to further reduce memory usage.
  • Inference Engine: llama.cpp (through its Python bindings). This library is designed for efficient LLM inference on CPUs and, crucially, supports quantization and offloading layers to the GPU on Android.
  • Framework: My own lightweight agent framework (details below). I didn't want the overhead of something like LangChain for this specific project.
  • UI: Termux & a simple Python script for command-line interaction. No fancy GUI… yet.

Building the Agent Framework

The core of IronVision Nexus is a very basic agent framework built around a shared "blackboard" – a dictionary representing the system's knowledge and task list. Each agent operates in a loop:

  1. Observe: Read the blackboard for new tasks and relevant information.
  2. Think: Generate a response using TinyLlama, based on its role, the current blackboard state, and any relevant context.
  3. Act: Update the blackboard with its findings, propose new tasks, or modify existing ones.

Here’s a simplified snippet illustrating the agent loop:

class Agent:
  def __init__(self, name, role, model):
    self.name = name
    self.role = role
    self.model = model

  def run(self, blackboard):
    while True:
      # 1. Observe
      tasks = blackboard.get("tasks", [])
      relevant_info = blackboard.get("knowledge", {})

      # 2. Think
      prompt = self.generate_prompt(tasks, relevant_info)
      response = self.model.generate(prompt)

      # 3. Act
      self.process_response(response, blackboard)
      time.sleep(0.5)  # Add a small delay to avoid hogging the CPU

  def generate_prompt(self, tasks, knowledge):
    # This is where the magic happens - crafting the prompt based on the agent's role
    # and the current context
    prompt = f"You are {self.role}. Your name is {self.name}.\n"
    prompt += "Current Tasks: " + str(tasks) + "\n"
    prompt += "Current Knowledge: " + str(knowledge) + "\n"
    prompt += "What should you do next? Be concise.\n"
    return prompt

  def process_response(self, response, blackboard):
    # Parse the response and update the blackboard accordingly
    print(f"{self.name}: {response}")
    # ... (logic to parse response and update the blackboard)
    pass
Enter fullscreen mode Exit fullscreen mode

This is a highly simplified example. The generate_prompt function is the crucial piece, dynamically creating a prompt tailored to the agent's role.

The 17 Agents: A Microscopic Society

So, what kind of agents are we talking about? Here's a breakdown of the current cast:

  • The Planner: Responsible for breaking down goals into actionable tasks.
  • The Researcher: Uses a simplified search API (implemented with Python's requests library, making limited calls to search engines when connectivity is available) to gather information.
  • The Summarizer: Condenses lengthy information into concise summaries.
  • The Critic: Reviews and critiques the work of other agents, identifying errors or inconsistencies.
  • The Memory Keeper: Maintains a long-term knowledge base (stored as a simple JSON file).
  • The Prioritizer: Determines the order in which tasks should be executed.
  • The Fact Checker: Attempts to verify information found by the Researcher.
  • The Translator: Translates between English and (currently) Spanish.
  • The Storyteller: Creates narrative summaries of the system's activity.
  • And 7 others: Covering roles like Sentiment Analyzer, Code Explainer, Pattern Recognizer, etc.

Challenges and Compromises

It wasn't all smooth sailing. Here are some major hurdles:

  • Performance: TinyLlama 1.1B is slow on a phone CPU. Quantization helps, but inference times are still significant. I’ve optimized prompts to be as concise as possible and limited the maximum output length.
  • Memory Constraints: Even with quantization, the model takes up a considerable amount of RAM. I’ve had to aggressively manage memory usage and sometimes deal with out-of-memory errors.
  • Prompt Engineering: Getting the agents to cooperate effectively requires extremely careful prompt engineering. It's a delicate balancing act between providing enough context and keeping the prompts short enough to fit within the LLM's context window.
  • Debugging: Debugging a multi-agent system running on a phone is… challenging. Relying heavily on logging and careful observation of

Top comments (0)