DEV Community

Samuel James Hiotis
Samuel James Hiotis

Posted on

I built a 34-agent AI swarm on my phone — here's how

I built a 34-agent AI swarm on my phone — here's how

Last Tuesday, at 2 AM, while waiting for a build to finish on my laptop, I had an idea that wouldn't leave me alone: what if I could orchestrate a swarm of AI agents on my phone? Not as a demo. Not a toy. A real, functioning multi-agent system — 34 agents, communicating, delegating, solving problems together — running on a device that fits in my pocket.

Three weeks later, it worked. Here's how.

Why a swarm?

Single-agent LLM workflows are well-trodden ground. But complex problems — research pipelines, code review systems, incident response — benefit from specialized agents working in parallel. The swarm pattern gives you:

  • Division of labor — agents own specific domains
  • Emergent coordination — no central brain, just protocols
  • Fault tolerance — one dead agent doesn't kill the system

The hard part? Making it run anywhere. Including a phone.

Architecture overview

┌─────────────────────────────────────┐
│         Swarm Orchestrator          │
│    (message broker + task queue)    │
└──────┬──────┬──────┬──────┬────────┘
       │      │      │      │
   ┌───▼──┐ ┌▼──┐ ┌──▼───┐ ┌▼────┐
   │Agent │ │Agent│ │Agent │ │Agent│
   │  01  │ │ 02  │ │ ...  │ │ 34  │
   └──────┘ └─────┘ └──────┘ └─────┘
Enter fullscreen mode Exit fullscreen mode

Each agent is a lightweight wrapper around an inference call. They communicate through a shared message bus — a simple pub/sub layer I built on top of Redis Streams (running locally via Termux on Android).

The agent definition

Every agent in the swarm follows the same interface:

class SwarmAgent:
    def __init__(self, agent_id, role, model_endpoint):
        self.id = agent_id
        self.role = role          # "researcher", "coder", "reviewer"...
        self.model = model_endpoint
        self.inbox = []

    async def process(self, message: dict) -> dict:
        """Receive a task, act, return result + next hops"""
        prompt = self._build_prompt(message)
        response = await self.model.complete(prompt)
        return {
            "from": self.id,
            "action": response["action"],
            "result": response["content"],
            "forward_to": response.get("delegate", [])
        }
Enter fullscreen mode Exit fullscreen mode

The key insight is the forward_to field. Agents don't just answer — they route. This is what makes the swarm alive.

Mobile-specific constraints

Running 34 agents on a phone meant fighting real physics:

Memory. Each agent context window eats ~200MB. 34 × 200MB = 6.8GB. Impossible on a 6GB phone.

Solution: I used a sparse activation model. Only 4-6 agents are "hot" at any time. The rest stay serialized to disk, waking on demand.

Latency. Network round-trips to cloud APIs kill interactivity.

Solution: I deployed small quantized models (Qwen2.5-3B, Llama-3.2-3B) locally via Ollama, and only escalated to GPT-4/Claude for final synthesis.

Battery. Continuous inference drains fast.

Solution: Batching + aggressive idle timeouts. Agents sleep after 30s of inactivity.

The orchestrator (the real magic)

class SwarmOrchestrator:
    def __init__(self):
        self.agents = {}
        self.task_queue = asyncio.Queue()
        self.active_count = 0
        self.max_concurrent = 6  # phone-friendly

    async def dispatch(self, task: dict):
        """Split task, assign to specialists, merge results"""
        # Phase 1: Decompose
        decomposer = self.agents["decomposer"]
        sub_tasks = await decomposer.process(task)

        # Phase 2: Parallel execution (bounded)
        semaphore = asyncio.Semaphore(self.max_concurrent)
        async def run_with_limit(agent_id, sub_task):
            async with semaphore:
                self.active_count += 1
                result = await self.agents[agent_id].process(sub_task)
                self.active_count -= 1
                return result

        jobs = [
            run_with_limit(t["agent"], t) 
            for t in sub_tasks
        ]
        results = await asyncio.gather(*jobs)

        # Phase 3: Synthesize
        synthesizer = self.agents["synthesizer"]
        return await synthesizer.process({
            "results": results,
            "original": task
        })
Enter fullscreen mode Exit fullscreen mode

The semaphore is critical. Without it, you'll OOM your phone in seconds.

Getting it on the device

I ran this on Android via Termux + Python 3.11 + Ollama. The stack:

  1. Termux — Linux environment, no root needed
  2. Ollama — local model serving
  3. Redis — message bus (pkg install redis)
  4. Uvicorn — lightweight API layer
  5. A Simple HTTP UI — I built a basic Flask frontend so I could interact with the swarm from my browser

Total APK-less footprint: ~1.2GB (mostly model weights).

What it actually did

I tested it on a real task: "Research the state of RISC-V in 2025, find 3 companies adopting it, and write a brief."

The swarm:

  1. decomposer split it into 3 research subtasks
  2. researcher_1..3 searched different sources in parallel
  3. synthesizer compiled findings
  4. reviewer fact-checked claims
  5. formatter produced the final output

Total time: ~90 seconds. Not bad for a phone.

Lessons learned

  • Don't over-engineer the protocol. A simple JSON-over-Redis pub/sub was enough. Don't start with gRPC or Kafka.
  • Quantization is your friend. Q4_K_M models gave 90% quality at 30% size.
  • Failure is normal. 2-3 agents crashed daily. I added automatic restart + health checks.
  • The phone gets hot. Thermal throttling kicked in after ~10 min of heavy load. Active cooling helps.

Is it useful?

Honestly? It's a proof of concept that distributed AI can run at the edge. The pattern scales — same code runs on a Raspberry Pi, a laptop, or a cloud cluster. The phone constraint forced me to be lean, and that leaness made the architecture better.

If you want to try something similar, start with 3 agents, not 34. Get the routing working, then scale.


Built by FractalMesh — autonomous AI agent platform. KuCoin: https://www.kucoin.com/r/af/012560iu

Top comments (0)