I built a 17-agent AI swarm on my phone — here's how
Okay, buckle up. This is going to be a bit of a deep dive. For the last few weeks, I've been obsessed with the idea of running a genuinely useful, multi-agent AI system entirely on a mobile device. Not offloading to the cloud, not relying on APIs, just raw, on-device processing. It sounds crazy, right? Performance is… challenging. But I’ve managed to get a 17-agent swarm up and running on my Pixel 7, and it’s surprisingly capable. This article details how I did it, the hurdles I faced, and why I think this approach could be significant.
The Goal: Local, Autonomous Problem Solving
The idea isn’t just about technical bragging rights. Think about scenarios where connectivity is unreliable or unavailable – disaster relief, remote fieldwork, or even just wanting privacy. An AI capable of processing information and making decisions locally, without a data connection, is incredibly valuable.
My initial test case? A simplified logistics problem. I wanted the swarm to figure out the most efficient route for “resource delivery” (imagining packages, medical supplies, etc.) based on a dynamically changing map with obstacles and varying urgency levels for different delivery points.
Why a Swarm?
I chose a swarm architecture for several reasons:
- Fault Tolerance: If one agent fails, the others keep working. Critical in a resource-constrained environment.
- Parallelism: Distributing the workload across multiple agents should improve performance (even on a phone!).
- Emergent Behavior: I wanted to see if complex problem-solving could emerge from the interaction of relatively simple agents.
The Tech Stack: Python, TensorFlow Lite, and a whole lot of Optimization
This wasn't happening in JavaScript or Flutter. I needed the power of numerical computation and existing AI tooling. Python was the clear winner, but running full-fat Python on a phone is… unwise. That's where TensorFlow Lite (TFLite) came in. TFLite allows you to convert trained TensorFlow models into a compact format optimized for mobile and embedded devices.
Here's a breakdown:
- Python: For initial model training and swarm logic.
- TensorFlow/Keras: To build the core AI agents.
- TensorFlow Lite: For on-device execution.
- NumPy (carefully): Used primarily during training, minimized in on-device code.
- SQLite: For lightweight, persistent data storage on the phone for map data and agent "memory".
- A custom “Agent Communication” protocol: Using Python dictionaries serialized to JSON for efficient data exchange between agents.
Agent Architecture: Simple Neural Networks & Reward Systems
Each of my 17 agents is designed for a specific sub-task. They aren’t monolithic “route planners.” Instead, they specialize:
- 4 x Map Sensor Agents: These agents scan the SQLite map data (represented as a grid) and identify obstacles and delivery points within their "vision" radius. They output a simplified representation of their surroundings.
- 4 x Urgency Assessment Agents: These agents take delivery point data (coordinates, urgency level) and output a “priority score”. The urgency level is a dynamically changing parameter.
- 4 x Route Evaluation Agents: These agents take a potential route (a list of coordinates) and evaluate its distance, estimated travel time (based on terrain – another SQLite data point), and potential risks (nearby obstacles).
- 3 x Route Coordinator Agents: These are the "leaders." They receive evaluations from the Route Evaluation Agents and orchestrate the final route selection based on a combination of factors – urgency, distance, and risk.
- 2 x Environmental Update Agents: These agents monitor simulated “environmental factors” (like new obstacles appearing on the map) and broadcast updates to the swarm.
Each agent’s “brain” is a small, fully connected neural network built in Keras. I deliberately kept them shallow (1-2 hidden layers with 16-32 neurons each) to minimize size and computational cost.
Here's a snippet of the Keras code for a Route Evaluation Agent:
import tensorflow as tf
from tensorflow import keras
from tensorflow.keras.layers import Dense
def create_route_evaluator():
model = keras.Sequential([
Dense(32, activation='relu', input_shape=(6,)), # Input: distance, travel_time, obstacle_count, urgency, map_quality, route_complexity
Dense(16, activation='relu'),
Dense(1, activation='sigmoid') # Output: Route score (0-1)
])
model.compile(optimizer='adam',
loss='binary_crossentropy',
metrics=['accuracy'])
return model
Training & Quantization: Making it Fit
Training these agents was done on a more powerful machine, naturally. I used a synthetic dataset generated based on randomly created maps and delivery scenarios. The training loop involved assigning rewards to agents based on their performance – for example, Route Evaluation Agents receive a high reward if they accurately predict the best route.
The crucial step was quantization. TFLite supports quantization, reducing the precision of model weights from 32-bit floating point to 8-bit integer. This significantly reduces model size and improves inference speed on mobile devices. I used post-training quantization in TensorFlow Lite:
converter = tf.lite.TFLiteConverter.from_keras_model(model)
converter.optimizations = [tf.lite.Optimize.DEFAULT]
tflite_model = converter.convert()
# Save the TFLite model
with open('route_evaluator.tflite', 'wb') as f:
f.write(tflite_model)
The On-Device Swarm: Threading & Communication
Getting everything to run on the phone required careful threading and communication management.
-
Multi-threading: I used Python's
threadingmodule to run each agent in its own thread. This allows for concurrent processing, maximizing the use of the phone's CPU cores. - Agent Communication: Agents communicate via a central “message broker” – essentially a Python queue. Agents publish their outputs to the queue, and other agents subscribe to the specific data they need. This decoupled architecture reduces dependencies and allows for easier scaling.
- SQLite Integration: The phone’s SQLite database stores the map data and acts as a shared “world state” accessible to all agents.
Here’s a simplified example of how an Agent might publish its output:
python
import queue
import json
message_queue = queue.Queue()
def publish_message(message_type, data):
message = {"type": message_type, "data": data}
message_queue.put(json.dumps(message))
Top comments (0)