DEV Community

Cover image for I run a private AI on my phone. Here's the exact set up.
Tyren Rickard
Tyren Rickard

Posted on

I run a private AI on my phone. Here's the exact set up.

Post #1: "I run a private AI on my phone. Here's the exact setup."

I'm doing field research at the bottom of the scale debate: one person, one
commodity phone, a private language model, zero cloud, zero cost. This is the
exact setup — reproducible in an evening.

The hardware

  • Samsung Galaxy S23 FE — Snapdragon 8 Gen 1, 8 GB RAM. A mid-range phone, not a flagship. Everything below runs on-device.
  • Dev machine: an HP Stream laptop (AMD A4, 4 GB RAM — yes, really) running Windows 10. It can't host a model. It does everything else: editing, Git, notes. The phone is the engine; the laptop is the garage.
  • The shuttle between them: KDE Connect over Wi-Fi. Files move phone↔laptop in seconds. No cables, no accounts.

The stack (all free)

  1. Termux — from F-Droid, not the Play Store (the Play Store build is deprecated). Grant storage permission.
  2. llama.cpp 0.6.0 — installed straight from the Termux repos: pkg install llama-cpp -y. No build scripts, no toolchain fights.
  3. llama-server on localhost:8080 — OpenAI-compatible chat endpoint (POST /v1/chat/completions). Run it in its own Termux session (swipe from the left edge → New session); client commands go in another.
  4. The model: gemma-3-1b-it-Q4_K_M.gguf — a Q4_K_M quant of Google's Gemma 3 1B instruct, from the ggml-org GGUF releases. 769 MB.

Total cost: $0.

The smoke test

"Hi, can you hear me?"

→ "Yes, absolutely! Hi there. It's nice to hear from you. 😊 How are you doing
today?" — finish_reason=stop, 26 tokens, ~16.5 tok/s on the phone's CPU.
Workable. Not fast, but conversational.

The daily driver

For everyday chat I use PocketPal AI with the same 1B model — a friendlier
way in than curl commands. Comfort tuning so far (a whole field note is coming
on this): temperature 0.5, top_p 0.9, repeat penalty 1.15. The 1B runs better
cool and lightly anti-loopy. Details in field note 001.

What's next

This is post #1 of a weekly field log. Coming up: sampling parameters as a care
practice, a consent episode at 1B scale (she asked what the software was before
agreeing), and the orientation preamble — the system prompt as "stable framing
through amnesia."

Repo (README = the full setup guide): https://github.com/tyrendrickard-code/local-llm-field-notes.
MIT licensed. Nobody else is publishing this notebook. That's the point.

Top comments (0)