DEV Community

Samuel James Hiotis
Samuel James Hiotis

Posted on

I built a 17-agent AI swarm on my phone — here's how

I built a 17-agent AI swarm on my phone — here's how

Okay, buckle up. This is going to be a bit of a deep dive. For the past few weeks I've been obsessing over the idea of running a genuinely useful, multi-agent system entirely on a smartphone. Not just a demo, not just a simplified example, but something that could actually perform a relatively complex task. And I did it. I built a 17-agent AI swarm, running on my Pixel 7, that analyzes text and generates targeted social media content.

Why a phone? Honestly, it's a challenge. We're so used to thinking of AI as cloud-based, reliant on massive servers. But the capabilities of modern smartphone processors are frankly astonishing. It forces you to be incredibly efficient, to think about model size, quantization, and creative code architecture. Plus, it’s portable. Where else can you carry a swarm intelligence around in your pocket?

This isn't about replacing large language models (LLMs) hosted in the cloud. It's about exploring the boundaries of what’s possible on-device.

The Core Idea: Social Media Content Alchemy

The goal was to take a relatively long-form piece of text (think a blog post, article, or even a transcript) and automatically generate a set of tailored social media posts for different platforms – Twitter (now X), LinkedIn, Facebook, and Instagram.

The 'swarm' architecture was crucial. Instead of relying on one big model, I broke the problem down into specialized agents, each handling a specific stage of the process. Think of it like an assembly line, but powered by AI.

The Agents: A Breakdown of the Swarm

Here’s a look at the 17 agents and their roles:

  1. Input Manager (1 agent): Handles loading the initial text input.
  2. Summarizer (2 agents): Each independently summarizes the text. Redundancy is built in for robustness.
  3. Keyword Extractor (2 agents): Identifies key themes and phrases. Again, dual for reliability.
  4. Tone Analyzer (2 agents): Detects the overall sentiment (positive, negative, neutral) and style of the text.
  5. Platform Profiler (4 agents): Each is trained (through prompt engineering) to understand the nuances of a specific platform: Twitter, LinkedIn, Facebook, and Instagram. They "know" what types of content perform best on each.
  6. Post Generator (4 agents): The workhorses. These generate the actual social media posts, using the summaries, keywords, tone analysis, and platform profiles as input. Each specializes in a different post length (short, medium, long).
  7. Hashtag Generator (2 agents): Generates relevant hashtags based on keywords and platform.

Tech Stack & Challenges

  • Language Model: This was the biggest hurdle. I needed something small enough to run efficiently on a phone, yet powerful enough to be useful. I settled on Phi-2, a 2.7 billion parameter model from Microsoft. It’s surprisingly capable for its size.
  • Framework: I used llama.cpp for running Phi-2 on Android. It provides excellent performance and quantization options.
  • Quantization: Crucially, I quantized the model down to Q4_K_M. This reduces the model size from ~5.5GB to around 2.7GB, making it manageable on my phone’s storage. It does come with a slight accuracy cost, but it’s a necessary trade-off.
  • Language: Python with a heavy reliance on the transformers and llama-cpp-python libraries.
  • Platform: Android - specifically, a Pixel 7 running Android 14.
  • Challenges:
    • Memory Management: Running a 2.7GB model on 8GB of RAM is tight. I implemented aggressive garbage collection and carefully managed memory allocation.
    • Processing Power: Inference is slower on a phone than on a dedicated GPU. I optimized prompts and batch sizes to maximize throughput.
    • Battery Life: Running this constantly drains the battery quickly! Optimization is key.

Code Snippets (Illustrative)

Let's look at some simplified snippets to give you a flavour.

1. Loading the Model (Python - llama-cpp-python)

from llama_cpp import Llama

llm = Llama(model_path="./phi-2.Q4_K_M.gguf", n_ctx=2048, n_gpu_layers=-1)  # Use all available GPU layers
Enter fullscreen mode Exit fullscreen mode

This initializes the Llama model, loading the quantized GGUF file. The -1 argument tries to offload as many layers as possible to the GPU (which my Pixel 7 has).

2. A Simple Agent Function (Post Generator)

def generate_post(summary, keywords, platform, length):
    prompt = f"Generate a {length} social media post for {platform} about the following:\n\nSummary: {summary}\n\nKeywords: {keywords}\n\nPost:"
    output = llm(prompt, max_tokens=150, stop=["\n\n"], echo=False)
    return output['choices'][0]['text'].strip()
Enter fullscreen mode Exit fullscreen mode

This function takes the summarized text, keywords, platform, and desired post length as input and constructs a prompt for Phi-2. The stop parameter prevents the model from generating endlessly.

3. Orchestrating the Swarm (Simplified)


python
# Load input text
input_text = load_text_from_file("my_article.txt")

# Summarize (using the Summarizer agents)
summaries = [summarize_text(input_text) for _ in range(2)]

# Extract Keywords (using the Keyword Extractor agents)
keywords_sets = [extract_keywords(input_text) for _ in range(2)]

# Average the keyword sets
keywords = list(set(keywords_sets[0] + keywords_sets[1])) #Remove Duplicates


# Generate posts for each platform
platform_posts = {}
for platform in ["Twitter", "LinkedIn", "Facebook", "Instagram"]:
    platform_posts[platform] = [
        generate_post(summaries[0], keywords, platform, "short"),
        generate_post(summaries[0], keywords, platform, "medium"),
        generate_post(summaries[0], keywords, platform, "long")
    ]

# Print Results
for platform, posts in platform_posts.items():
    print(f"--- {platform} ---")
    for i, post in enumerate(posts):
        print(
Enter fullscreen mode Exit fullscreen mode

Top comments (0)