DEV Community

Samuel James Hiotis
Samuel James Hiotis

Posted on

I built a 17-agent AI swarm on my phone — here's how

I built a 17-agent AI swarm on my phone — here's how

Okay, buckle up. This is a bit of a rabbit hole, but a really fun one. I’ve spent the last few weeks building and running a swarm of 17 independent AI agents, all contained within and operating on my phone. Yes, you read that right. No cloud servers, no expensive GPUs. Just a Pixel 7 Pro and a lot of Python.

This wasn’t about solving world hunger (though, hey, maybe someday!), it was about pushing the boundaries of what’s possible with on-device AI and exploring emergent behaviour within a limited computational environment. It’s a proof-of-concept for what I’m calling “FractalMesh,” a system designed for decentralized, localized AI processing.

Why on-device?

The current AI landscape is dominated by massive models running in the cloud. That's fantastic for powerful applications, but comes with downsides: latency, privacy concerns, and reliance on network connectivity. On-device AI, however, offers a lot of promise. It’s faster, more private, and can operate offline. My phone is essentially a mini-supercomputer just waiting to be tapped. The challenge? Squeezing 17 brains into it.

The Architecture: Agents and the Coordinator

The system is built around a central “Coordinator” agent and 17 “Worker” agents. Think of it like a hive mind.

  • Coordinator: This is the orchestrator. It defines the task, divides it into sub-tasks, distributes those sub-tasks to the Workers, collects their responses, and synthesizes them into a final output. It doesn't do the core processing, it manages it.
  • Workers: These are the specialists. Each agent is designed with a specific, focused role. In this iteration, I've got agents for:
    • Sentiment Analysis (x3): Evaluate text for emotional tone.
    • Topic Extraction (x3): Identify key themes within text.
    • Keyword Generation (x3): Brainstorm related keywords.
    • Text Summarization (x3): Condense text into concise summaries.
    • Question Answering (x3): Answer questions based on provided text.
    • Translation (x2): Translate text between English and Spanish.

Tech Stack & Optimizations: Python, llama.cpp, and a LOT of patience

The core of the project is Python, running via a custom script I’ve packaged with Termux on Android. Termux provides a Linux-like environment on Android allowing me to run Python and install packages. However, running full-blown transformer models (even smaller ones) on a phone isn’t feasible without serious optimization.

This is where llama.cpp came in. llama.cpp is a fantastic project that allows you to run large language models (LLMs) quantized for CPU usage. This is crucial for phone performance. I used a quantized version of the TinyLlama model (1.1B parameters, Q4_K_M quantization) for each agent. It's not GPT-4, but it's surprisingly capable for its size.

Here’s a simplified snippet demonstrating how an agent (let's say a Sentiment Analyzer) is instantiated and used:

from llama_cpp import Llama
import time

class SentimentAgent:
    def __init__(self, model_path):
        self.llm = Llama(model_path=model_path, n_ctx=512, n_threads=4) # Threads adjusted for phone cores

    def analyze_sentiment(self, text):
        prompt = f"Analyze the sentiment of the following text: '{text}'.  Respond with 'Positive', 'Negative', or 'Neutral'."
        output = self.llm(prompt, max_tokens=20, stop=["\n"], echo=False)
        return output["choices"][0]["text"].strip()

# Example usage
sentiment_agent = SentimentAgent("/sdcard/models/tinyllama-q4_k_m.gguf") # Path on phone
text_to_analyze = "This is an absolutely amazing project! I love it!"
sentiment = sentiment_agent.analyze_sentiment(text_to_analyze)
print(f"Sentiment: {sentiment}")
Enter fullscreen mode Exit fullscreen mode

Key Optimizations:

  • Quantization: Using Q4_K_M quantization reduced the model size significantly, allowing me to fit more agents.
  • Context Length: Limiting the context window ( n_ctx in llama.cpp) to 512 tokens helped keep memory usage down.
  • Thread Control: n_threads was set to 4, matching my phone's core count. Experimenting with this value is crucial for performance.
  • Batching: While not fully implemented yet, I'm planning to batch requests to the same agent type to improve throughput.
  • Prioritization: The Coordinator prioritizes tasks and agents based on estimated complexity. More complex tasks get more resources (longer processing time).

The Coordinator's Role: Orchestrating the Swarm

The Coordinator is where the real magic happens. It receives an initial text input, breaks it down, and assigns portions to the appropriate Workers. Here's a simplified view of the Coordinator's logic:

import threading

class Coordinator:
    def __init__(self, agents):
        self.agents = agents  # Dictionary of agent types (sentiment, topic, etc.)
        self.results = {}

    def process_text(self, text):
        self.results = {}
        threads = []

        # Divide the task & assign to workers
        for agent_type, agent_list in self.agents.items():
            for agent in agent_list:
                thread = threading.Thread(target=self.run_agent, args=(agent, text, agent_type))
                threads.append(thread)
                thread.start()

        for thread in threads:
            thread.join()

        # Synthesize results (currently just prints them)
        print("--- Results ---")
        for agent_type, results in self.results.items():
            print(f"{agent_type}: {results}")

    def run_agent(self, agent, text, agent_type):
        # Run the agent's function
        if agent_type == "sentiment":
            result = agent.analyze_sentiment(text)
        elif agent_type == "topic":
            result = agent.extract_topics(text)
        # ... other agents ...

        self.results[agent_type] = result
Enter fullscreen mode Exit fullscreen mode

This example uses threading to run agents concurrently. The actual implementation is more complex, handling error

Top comments (0)