DEV Community

Samuel James Hiotis
Samuel James Hiotis

Posted on

I built a 17-agent AI swarm on my phone — here's how

I built a 17-agent AI swarm on my phone — here's how

Okay, buckle up. This is a bit of a rabbit hole. For the past few weeks, I’ve been obsessively tinkering with running a multi-agent system – a swarm of 17 individual AI “agents” – entirely on my Android phone. Not cloud-connected, not relying on a server. Entirely on-device. And it works. It's slow, resource intensive, and occasionally crashes my phone, but it's a proof-of-concept I'm incredibly proud of. Here’s how I did it, and why.

Why bother? The Motivation.

The current AI hype is largely focused on massive models requiring significant compute. That’s fantastic, but leaves a lot of potential unexplored. I'm interested in the idea of distributed intelligence – can we create useful systems with less power, more privacy, and greater resilience by relying on many small, specialized agents operating locally? My phone felt like a good, contained environment to experiment with that. Plus, the challenge was just... compelling.

The Tech Stack: A Surprisingly Feasible Combo

The core technologies powering this aren't as intimidating as you might think. Here’s what I used:

  • Python: The language of choice for rapid prototyping and AI/ML.
  • Pyodide: This is the game-changer. Pyodide allows you to run Python inside a web browser, using WebAssembly (WASM). This means I could run my Python code directly in a web view on Android.
  • Termux (Android Terminal Emulator): Required for installing some dependencies and managing the Pyodide environment.
  • LLaMA.cpp & quantized models: Running full-size LLMs on a phone isn’t feasible. LLaMA.cpp is a C++ port of the LLaMA architecture optimized for running on CPUs (and even better, on Apple Silicon… but that’s for another article!). Crucially, it supports quantized models – drastically reduced precision versions of the LLM, making them much smaller and faster. I used a 4-bit quantized version of Mistral 7B.
  • HTML/JavaScript: To build the basic web interface and handle communication between the Python/Pyodide environment and the Android app.

The Architecture: A Swarm of Specialists

The idea isn’t to have 17 general-purpose AI assistants. It’s to have 17 specialized agents. Each agent has a specific persona, role, and limited knowledge base. Think of it like a miniature, AI-powered department.

Here are some examples of my agents:

  • "FactChecker": Dedicated to verifying information.
  • "CreativeWriter": Generates stories, poems, or marketing copy.
  • "CodeReviewer": Attempts to find bugs in provided code snippets.
  • "Summarizer": Condenses long texts into shorter summaries.
  • "TaskPlanner": Breaks down complex tasks into smaller, manageable steps.
  • "SentimentAnalyzer": Analyzes the emotional tone of text.
  • …and 11 more, each with a unique purpose.

The Code: A Simplified Example (Agent Communication)

This is heavily simplified, but it shows the core logic of how agents communicate. The entire system is message-passing based.

# Python (running in Pyodide)
import llama_cpp

llm = llama_cpp.Llama(model_path="./models/mistral-7b-instruct-v0.2.Q4_K_M.gguf") # Path to quantized model

def generate_response(prompt, agent_persona):
  """Generates a response from the LLM with a specific persona."""
  full_prompt = f"You are {agent_persona}. Respond to the following:\n{prompt}"
  output = llm(full_prompt, max_tokens=150, stop=["Q:", "\n\n"]) # Limit token length
  return output["choices"][0]["text"].strip()

def agent_interaction(query, agent_name):
  """Simulates interaction with a specific agent."""
  agent_persona = AGENT_PERSONAS.get(agent_name, "a helpful AI assistant")
  response = generate_response(query, agent_persona)
  return response

AGENT_PERSONAS = {
    "FactChecker": "a meticulous fact-checker. You only state verified information.",
    "CreativeWriter": "a creative and imaginative writer.",
    "CodeReviewer": "an experienced software developer focused on finding bugs."
}

# Example usage:
query = "What is the capital of France?"
fact_checker_response = agent_interaction(query, "FactChecker")
print(f"FactChecker says: {fact_checker_response}")
Enter fullscreen mode Exit fullscreen mode

This is running inside Pyodide. The Javascript then grabs the output from Pyodide and displays it in the web view. The important part is generate_response. It takes a prompt and a persona, formats them, and sends them to the LLaMA model.

The Communication Layer: JavaScript & Pyodide Bridging

Pyodide runs in a sandboxed environment. To interact with it from JavaScript, you use Pyodide's runPython() function. This lets you execute Python code from JavaScript, and vice-versa. This is how the user interface interacts with the agents.

// JavaScript
async function sendMessageToAgent(agentName, message) {
  const result = await pyodide.runPython(`
    import sys
    sys.stdout.flush()  # Important for capturing output
    from main import agent_interaction  # Assuming your Python script is named 'main.py'
    response = agent_interaction("${message}", "${agentName}")
    response
  `);
  return result;
}
Enter fullscreen mode Exit fullscreen mode

This JavaScript function takes the agent name and message, sends it to the Python script, runs the agent_interaction function, and returns the response. The sys.stdout.flush() is crucial – Pyodide often buffers output, and you need to flush it to get the results back to JavaScript immediately.

Challenges & Optimizations (and why it's slow!)

This project was not easy. Here were the major hurdles:

  • Performance: Running quantized LLMs on a phone is still slow. Mistral 7B, even quantized, takes several seconds to generate a response. This is why the UI can feel sluggish.
  • Memory Management: LLMs are memory hogs. I had to carefully manage the model loading and unloading, and use a very low quantization level (4-bit) to fit everything into my phone’s memory.
  • Pyodide Limitations:

Top comments (0)