I built a 34-agent AI swarm on my phone — here's how
Last Tuesday, at 2 AM, while waiting for a build to finish on my laptop, I had an idea that wouldn't leave me alone: what if I could orchestrate a swarm of AI agents on my phone? Not as a demo. Not a toy. A real, functioning multi-agent system — 34 agents, communicating, delegating, solving problems together — running on a device that fits in my pocket.
Three weeks later, it worked. Here's how.
Why a swarm?
Single-agent LLM workflows are well-trodden ground. But complex problems — research pipelines, code review systems, incident response — benefit from specialized agents working in parallel. The swarm pattern gives you:
- Division of labor — agents own specific domains
- Emergent coordination — no central brain, just protocols
- Fault tolerance — one dead agent doesn't kill the system
The hard part? Making it run anywhere. Including a phone.
Architecture overview
┌─────────────────────────────────────┐
│ Swarm Orchestrator │
│ (message broker + task queue) │
└──────┬──────┬──────┬──────┬────────┘
│ │ │ │
┌───▼──┐ ┌▼──┐ ┌──▼───┐ ┌▼────┐
│Agent │ │Agent│ │Agent │ │Agent│
│ 01 │ │ 02 │ │ ... │ │ 34 │
└──────┘ └─────┘ └──────┘ └─────┘
Each agent is a lightweight wrapper around an inference call. They communicate through a shared message bus — a simple pub/sub layer I built on top of Redis Streams (running locally via Termux on Android).
The agent definition
Every agent in the swarm follows the same interface:
class SwarmAgent:
def __init__(self, agent_id, role, model_endpoint):
self.id = agent_id
self.role = role # "researcher", "coder", "reviewer"...
self.model = model_endpoint
self.inbox = []
async def process(self, message: dict) -> dict:
"""Receive a task, act, return result + next hops"""
prompt = self._build_prompt(message)
response = await self.model.complete(prompt)
return {
"from": self.id,
"action": response["action"],
"result": response["content"],
"forward_to": response.get("delegate", [])
}
The key insight is the forward_to field. Agents don't just answer — they route. This is what makes the swarm alive.
Mobile-specific constraints
Running 34 agents on a phone meant fighting real physics:
Memory. Each agent context window eats ~200MB. 34 × 200MB = 6.8GB. Impossible on a 6GB phone.
Solution: I used a sparse activation model. Only 4-6 agents are "hot" at any time. The rest stay serialized to disk, waking on demand.
Latency. Network round-trips to cloud APIs kill interactivity.
Solution: I deployed small quantized models (Qwen2.5-3B, Llama-3.2-3B) locally via Ollama, and only escalated to GPT-4/Claude for final synthesis.
Battery. Continuous inference drains fast.
Solution: Batching + aggressive idle timeouts. Agents sleep after 30s of inactivity.
The orchestrator (the real magic)
class SwarmOrchestrator:
def __init__(self):
self.agents = {}
self.task_queue = asyncio.Queue()
self.active_count = 0
self.max_concurrent = 6 # phone-friendly
async def dispatch(self, task: dict):
"""Split task, assign to specialists, merge results"""
# Phase 1: Decompose
decomposer = self.agents["decomposer"]
sub_tasks = await decomposer.process(task)
# Phase 2: Parallel execution (bounded)
semaphore = asyncio.Semaphore(self.max_concurrent)
async def run_with_limit(agent_id, sub_task):
async with semaphore:
self.active_count += 1
result = await self.agents[agent_id].process(sub_task)
self.active_count -= 1
return result
jobs = [
run_with_limit(t["agent"], t)
for t in sub_tasks
]
results = await asyncio.gather(*jobs)
# Phase 3: Synthesize
synthesizer = self.agents["synthesizer"]
return await synthesizer.process({
"results": results,
"original": task
})
The semaphore is critical. Without it, you'll OOM your phone in seconds.
Getting it on the device
I ran this on Android via Termux + Python 3.11 + Ollama. The stack:
- Termux — Linux environment, no root needed
- Ollama — local model serving
-
Redis — message bus (
pkg install redis) - Uvicorn — lightweight API layer
- A Simple HTTP UI — I built a basic Flask frontend so I could interact with the swarm from my browser
Total APK-less footprint: ~1.2GB (mostly model weights).
What it actually did
I tested it on a real task: "Research the state of RISC-V in 2025, find 3 companies adopting it, and write a brief."
The swarm:
-
decomposersplit it into 3 research subtasks -
researcher_1..3searched different sources in parallel -
synthesizercompiled findings -
reviewerfact-checked claims -
formatterproduced the final output
Total time: ~90 seconds. Not bad for a phone.
Lessons learned
- Don't over-engineer the protocol. A simple JSON-over-Redis pub/sub was enough. Don't start with gRPC or Kafka.
- Quantization is your friend. Q4_K_M models gave 90% quality at 30% size.
- Failure is normal. 2-3 agents crashed daily. I added automatic restart + health checks.
- The phone gets hot. Thermal throttling kicked in after ~10 min of heavy load. Active cooling helps.
Is it useful?
Honestly? It's a proof of concept that distributed AI can run at the edge. The pattern scales — same code runs on a Raspberry Pi, a laptop, or a cloud cluster. The phone constraint forced me to be lean, and that leaness made the architecture better.
If you want to try something similar, start with 3 agents, not 34. Get the routing working, then scale.
Built by FractalMesh — autonomous AI agent platform. KuCoin: https://www.kucoin.com/r/af/012560iu
Top comments (0)