DEV Community

Samuel James Hiotis
Samuel James Hiotis

Posted on

I built a 17-agent AI swarm on my phone — here's how

I built a 17-agent AI swarm on my phone — here's how

Okay, buckle up. This is going to be a deep dive. For the past few weeks, I’ve been obsessed with a seemingly ludicrous idea: could I run a multi-agent AI system, a swarm even, directly on my phone? Not just simple chatbots, but a network of interacting agents with distinct roles, communicating and collaborating to achieve a (relatively) complex goal. The answer, surprisingly, is… yes. And it's been a wild ride.

This wasn't about raw processing power (my phone is a perfectly respectable Pixel 7, but it's not a data center). It was about clever architecture, resource management, and exploiting the advancements in on-device machine learning. Here's how I built my 17-agent AI swarm, affectionately dubbed “Project Nightingale,” running entirely on Android.

The Goal: Decentralized Information Gathering & Summarization

Before getting into the tech, let's define the goal. I wanted a system that could autonomously gather information about a specific topic (currently, emerging trends in web3 security), analyze it from multiple angles, and present a concise, summarized report. Instead of one massive LLM trying to do everything, I envisioned a swarm of specialized agents:

  • Scrapers (4 agents): Responsible for web scraping from specific sources (news sites, blogs, Twitter, Reddit).
  • Analyzers (6 agents): These agents perform sentiment analysis, topic extraction, and key phrase identification on the scraped content. They’re specialized – some focused on technical details, others on market impact, etc.
  • Validators (3 agents): These agents assess the credibility of the sources and the information. They cross-reference data and flag potential misinformation.
  • Summarizers (2 agents): The final layer, taking the analyzed data and producing a cohesive summary report.
  • Coordinator (2 agents): Two redundant coordinator agents manage the workflow, distribute tasks, and collect results. This redundancy is key for robustness.

The Tech Stack: Python, llama.cpp, and a sprinkle of Magic

My entire system is built using Python. It’s the most flexible option for prototyping and accessing the necessary libraries. The core of each agent leverages llama.cpp – a fantastic library that allows running Large Language Models (LLMs) locally, even on devices with limited resources. I chose a quantized version of Mistral 7B (specifically Q4_K_M) for its balance of performance and size.

Here’s a simplified example of how an analyzer agent initializes the LLM using llama.cpp (using the python bindings):

from llama_cpp import Llama

llm = Llama(model_path="./models/mistral-7b-instruct-v0.1.Q4_K_M.gguf", n_ctx=2048) #Adjust n_ctx based on phone RAM
Enter fullscreen mode Exit fullscreen mode

Key components:

  • Agent Class: A base class defining common functionalities like message handling, task execution, and reporting.
  • Task Queue (Redis): A Redis server running locally on the phone (using a lightweight Redis implementation) acts as a task queue. The coordinator agents push tasks (e.g., "scrape this URL") onto the queue, and worker agents pull tasks as they become available. This decouples the agents and allows for asynchronous operation.
  • Data Storage (SQLite): SQLite is used for persistent storage of scraped data, analysis results, and agent logs. It’s lightweight and ideal for on-device use.
  • Communication (Simple Text-Based Protocol): Agents communicate via a simple text-based protocol over TCP sockets. This keeps overhead low and avoids complex serialization/deserialization.

The Architecture: A Decentralized Mesh

This isn't a hierarchical system. While the coordinators initiate tasks, agents aren't strictly ordered. An analyzer agent, for example, can directly request additional information from a scraper if it needs clarification.

Here's a snippet of the agent communication loop:

import socket

class Agent:
    def __init__(self, agent_id, role):
        self.agent_id = agent_id
        self.role = role
        self.socket = socket.socket(socket.AF_INET, socket.SOCK_STREAM)
        self.socket.bind(('localhost', 5000 + agent_id)) #Assign unique ports
        self.socket.listen(5)

    def run(self):
        print(f"Agent {self.agent_id} ({self.role}) listening...")
        while True:
            conn, addr = self.socket.accept()
            with conn:
                data = conn.recv(1024).decode()
                response = self.process_message(data)
                conn.sendall(response.encode())
Enter fullscreen mode Exit fullscreen mode

Each agent listens on a unique port. The coordinator agents act as central hubs, but agents can also directly connect to each other when needed. This mesh network provides resilience – if one agent fails, others can often compensate.

Challenges and Optimizations

Running 17 LLMs, even quantized ones, on a phone is challenging. Here's what I tackled:

  • Memory Management: Crucial. I used context size (n_ctx) strategically. Analyzers dealing with longer texts used larger contexts, while scrapers used smaller ones. I also implemented aggressive garbage collection.
  • CPU Throttling: Mobile CPUs throttle under sustained load. I introduced sleep intervals between tasks and limited the number of concurrent LLM inferences.
  • Battery Life: This thing is a battery hog. I added a battery monitoring module that automatically reduces the number of active agents when battery levels are low.
  • Context Switching: With so many agents running concurrently, context switching overhead became significant. Using Redis for task queuing and keeping agent code streamlined helped mitigate this.

Code Snippets & Key Functionality

Let’s look at a simplified example of a scraper agent:

import requests
from bs4 import BeautifulSoup

def scrape_website(url):
  try:
    response = requests.get(url, timeout=10)
    response.raise_for_status() # Raises HTTPError for bad responses (4xx or 5xx)
    soup = BeautifulSoup(response.content, 'html.parser')
    text = soup.get_text()
    return text[:4000] #Limit text length - LLM context window
  except requests.exceptions.RequestException as e:
    print(f"Error scraping {url}: {e}")
    return ""
Enter fullscreen mode Exit fullscreen mode

And a simplified example of how an Analyzer might use the LLM to perform sentiment analysis:


python
def analyze_sentiment(text):
  prompt = f
Enter fullscreen mode Exit fullscreen mode

Top comments (0)