DEV Community

CR4CODE
CR4CODE

Posted on

I gave a language model hands inside Android — and it built me an app factory

Most LLM assistants are "brains in a jar." They can reason, write code, explain concepts — but they can't do anything in the real world. Everything they produce is text that a human copies and pastes somewhere.

I decided to fix that. Termux Assistant AI is an Android app that bridges a WebView with DeepSeek Chat and simultaneously gives the assistant direct shell access inside Termux via local HTTP servers. In effect — the language model gets hands.

The project is fully open source (MIT): github.com/CR4CODE/Termux-Assistant-AI.

What Termux is, and why it's the key

Termux is a terminal emulator for Android with a full Linux environment — no root required. Bash, Python, git, ssh, package manager — all working. Essentially a regular Linux distro in your pocket.

The key insight: if I have Termux, I have a shell. And if a language model has a shell, it can run arbitrary commands, work with files, compile code, publish releases.

Two questions remained:

  1. How to bridge the WebView app with Termux?
  2. How to safely pass commands from the model to the shell?

Architecture: three layers

The system consists of three isolated layers.

Layer 1: WebView + DeepSeek

The main app is an Android WebView that loads chat.deepseek.com. The user chats with the model as usual. But the app intercepts the "Send" button and can insert prompts into the input field and read responses from the DOM.

Implemented via a JavaScript bridge VibeWebBridge: methods sendMessage, readLastAnswer, lastAnswerHash, isThinking, isStopped execute JS in the WebView and return results through callbacks.

Response reading works by finding [class*=markdown] elements and hashing the last block to distinguish a new response from an old one.

Layer 2: HTTP server in Termux (port 8767)

Inside Termux runs a Python server ai-tasker-server. It listens on 127.0.0.1:8767 and accepts tasks from the app:
the app accepts tasks:

POST /task
Body: {"task": "auto:uptime"}
Response: {"status":"success","output":"...","exit_code":0}

Routes are grouped by prefix:

  • auto: — direct bash via bash -c
  • project_* — read/write/build Android projects
  • vk_, tg_ — posting to VK and Telegram
  • release_* — release pipeline: CHANGELOG, tag, GitHub Release
  • bot_* — PiarBot management
  • apply_patches: — apply patches to source files
  • log: — read logs

A key design decision: auto: commands go straight to bash, without LLM-based classification. The first version tried to guess the command type, which caused hangs on non-existent routes. Now it's just subprocess.run(["bash","-c", code], timeout=120).

Layer 3: VibeService (port 8768)

A separate Android service that manages the WebView and exposes an HTTP API:

  • POST /send — insert text into input field and press "Send"
  • GET /read — read last DeepSeek response
  • GET /hash — hash of last response (for change detection)
  • GET /thinking — is the model generating right now?
  • GET /stopped — is the model stuck in "Stopped" state?
  • POST /install-apk — install APK from path

This service is how the assistant talks to DeepSeek and runs autonomous loops.

Full development cycle: from prompt to APK

This is the most interesting part — Vibe mode. You write one task, and the system does the rest.

Workflow:

  1. User runs vibe-run <project> "<task>"
  2. Script collects project tree and content of all .java/.xml files
  3. Builds a prompt requiring files back in format === FILE: path === ... === END FILE ===
  4. Sends to DeepSeek via VibeService (/send)
  5. Waits for response stabilization: hash unchanged for N seconds in a row
  6. Parses response, applies files to project
  7. Builds APK via aapt2 → javac → d8 → apksigner
  8. If build fails — sends error log back to the model and retries (up to 5 iterations)
  9. On success — APK is handed over for installation via /install-apk

The key trick is self-healing. If the model stops in "Stopped" state (e.g., token limit), the background watchdog vibe-auto-resume checks /stopped every 5 seconds and automatically sends "continue". Without this, the autonomous run just hangs.

Real use cases

  • Fast prototypes — idea → working APK in minutes
  • Phone automation — scheduled tasks, parsing, data dispatch
  • Content publishing — scheduled VK/Telegram posts with LLM-generated text
  • Release management — build, CHANGELOG, tag, GitHub Release, social announcement

Honest limitations

  • Phone resources. Heavy projects (Gradle, Kotlin, Compose) won't fit — the pipeline is tuned for lightweight Java+XML
  • Fragility of autonomous runs. One typo in a patch can crash the server, so all changes are backed up
  • Dependency on an external LLM. DeepSeek Chat is not local. A local option via llama.cpp in Termux is the next step
  • No vision. The assistant "sees" only the WebView DOM, not the whole screen

What's next

  • Voice control via Termux API (TTS/STT)
  • Local LLM via llama.cpp for full autonomy
  • Vision — OCR + uiautomator to control any app
  • APK marketplace — build → sign → publish to own channel

Conclusion

Termux Assistant AI is an experiment: what happens if you give a language model hands? It turns out — a lot. The model builds apps, ships releases, fixes its own bugs, runs social media. All inside a single Android phone, without a PC or cloud infrastructure.

The project is fully open source:

I'm not a programmer. I don't know a single line of code. Everything you see was built by one person and one AI that became my hands.

Questions and criticism are welcome in the comments.

Top comments (0)