DEV Community

Samuel James Hiotis
Samuel James Hiotis

Posted on

I built a 17-agent AI swarm on my phone — here's how

I built a 17-agent AI swarm on my phone — here's how

Okay, buckle up. This is a bit of a wild ride. For the past few weeks, I've been obsessively working on a project that sounded utterly insane when I first conceived it: building a multi-agent AI swarm that runs entirely on my smartphone. Not leveraging cloud APIs, not offloading processing… genuinely running 17 independent AI “agents” locally.

Why? Because I could. And because the potential for truly personal, always-on, privacy-focused AI is massively exciting. I’ve always been fascinated by swarm intelligence, and the idea of democratizing access to complex AI capabilities – not gatekeeping it behind expensive compute and corporate servers – felt important.

This isn’t a polished, production-ready product (yet!). It’s a proof-of-concept, a testament to the surprising capabilities of modern smartphone hardware, and a deep dive into some fascinating (and frustrating) optimization techniques. Here's how I did it.

The Goal: A Decentralized Information Processing Unit

The core idea wasn’t just to run 17 AI models; it was to create a swarm. Each agent would have a specific task, operate independently, and communicate with others to achieve a higher-level goal. I envisioned it as a decentralized information processing unit – constantly learning, adapting, and providing contextually relevant insights.

My initial use case? A "Personal Situation Awareness Engine". Think: subtly monitoring environmental factors, news feeds, calendar appointments, and even my phone’s usage patterns to provide preemptive information. Example: “Traffic is heavy on your usual route home, consider leaving now” or “You have a meeting in 15 minutes, do you want to silence notifications?”.

The Tech Stack: Python, TensorFlow Lite, and a LOT of Optimisation

The foundation of this project is Python. It’s versatile, relatively easy to debug, and boasts excellent machine learning libraries. However, directly running standard Python/TensorFlow models on a phone (a Samsung Galaxy S22 in my case) would be… slow. We’re talking minutes per inference, which defeats the purpose of a swarm.

This is where TensorFlow Lite (TFLite) came in. TFLite is TensorFlow’s lightweight solution for mobile and embedded devices. It allows you to convert trained TensorFlow models into a smaller, optimized format. It still wasn’t enough.

Agent Design & Model Choices

Each agent is a relatively simple TFLite model, designed for a specific task. Here’s a breakdown of a few key agents:

  • Sentiment Analyzer (x2): Processes news headlines or text messages to detect sentiment (positive, negative, neutral). Uses a pre-trained BERT model, heavily quantized.
  • Traffic Predictor (x3): Analyzes historical traffic data (a small, locally stored dataset) and current news reports to predict traffic congestion. A simplified LSTM network.
  • Calendar Event Processor (x2): Parses calendar event details, extracts key information (time, location, attendees). Uses a rule-based system combined with a small NER (Named Entity Recognition) model.
  • Usage Pattern Monitor (x4): Tracks app usage patterns to predict likely activities. A simple feedforward neural network.
  • News Summarizer (x3): Uses a lightweight transformer model to summarize news articles.
  • Contextual Router: This central agent analyzes data from all other agents and decides what information is relevant to the user. A relatively complex model, but crucial for coordinating the swarm.

The Code: A Snippet of Agent Initialisation

This shows how I load and initialise a TFLite model within an agent class. It’s Python using the TFLite interpreter:

import tensorflow as tf

class SentimentAgent:
    def __init__(self, model_path):
        self.interpreter = tf.lite.Interpreter(model_path=model_path)
        self.interpreter.allocate_tensors()
        self.input_details = self.interpreter.get_input_details()
        self.output_details = self.interpreter.get_output_details()

    def analyze_sentiment(self, text):
        # Preprocess text (tokenization, padding, etc.) - omitted for brevity
        input_data = np.array(processed_text, dtype=np.float32)
        self.interpreter.set_tensor(self.input_details[0]['index'], input_data)
        self.interpreter.invoke()
        output_data = self.interpreter.get_tensor(self.output_details[0]['index'])
        return output_data[0][0] # Sentiment score
Enter fullscreen mode Exit fullscreen mode

Quantization, Pruning, and the Art of Squeezing Performance

This is where things got really interesting (and painful). Just converting to TFLite wasn't sufficient. The initial inference times were still unacceptable. I had to dive deep into model optimization:

  • Quantization: Reducing the precision of model weights from 32-bit floating-point to 8-bit integer. This halved model size and significantly improved speed, but introduced some accuracy loss. I experimented with different quantization techniques (dynamic range, full integer) to find the best balance.
  • Pruning: Removing unnecessary connections (weights) in the neural networks. This made the models smaller and faster, but required careful retraining to maintain performance.
  • Operator Fusion: TFLite automatically fuses certain operations for efficiency, but I also explored manual operator fusion where possible.
  • Threading: Leveraging multi-threading to parallelize computations. This provided a modest performance boost, but was limited by the phone's CPU architecture.
  • XNNPACK Backend: TFLite supports different backends for optimized execution. XNNPACK, designed for ARM CPUs, offered substantial improvements on my phone.

Communication & Orchestration: The Agent Network

The agents don't operate in isolation. They communicate via a lightweight messaging system built on Python’s multiprocessing.Queue. The Contextual Router agent acts as the central coordinator, receiving messages from all other agents, analyzing the information, and generating appropriate notifications.

Here’s a simplified example of how an agent sends data:

def send_data(queue, data):
    queue.put(data)

# Example usage:
traffic_data = {"location": "Highway 101", "congestion": 0.8}
agent_queue.send_data(context_router_queue, traffic_data)
Enter fullscreen mode Exit fullscreen mode

Challenges & Lessons Learned

This project wasn’t without its hurdles:

  • Memory Management: 17 agents, each with its own model and data, can quickly eat up memory. I had to implement careful memory management techniques to avoid crashes.
  • Battery Drain: Constantly running AI models

Top comments (0)