This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
localchat is a chat app for an AI model that runs on your own machine. A Go server renders the UI, and Google's Gemma 4 E2B writes the replies, streamed into your browser token by token. Your messages stay in memory on your laptop for as long as you run the app.
I talked to a friend before about what a bare-basic and simple Go web app with AI integration would look like, without too many difficult parts. So I spent some time setting this up and kept the project small enough to read in one sitting.
I picked Gemma while developing because I like running small models on my own hardware. Gemma 4 E2B fits in a few gigabytes of RAM and starts answering in about a second on my laptop. It's fast and easy to develop against while still having realistic output.
My friend and I get two things:
- Code you can read in an evening. A HTML form posts to a Go handler, the handler calls the model, and the reply streams back into the page.
- A base to build on. You can change the system prompt, swap the model with one environment variable, or install the app as a background service on Linux, macOS or Windows. You can also run it as a container if you like.
Code
Repo: git.b0b.be/bdeb/localchat
cmd/localchat/ entry point: serve + install/start/stop as an OS service
internal/llm/ ~150-line OpenAI-compatible streaming client (stdlib only)
internal/chat/ in-memory conversations, one per browser session
internal/web/ routes, SSE streaming, embedded static assets
internal/web/views/ Templ components
scripts/e2e.py Playwright browser test against the real model
compose.yaml app + llama.cpp + Gemma 4 E2B
Demo
On my M3 MacBook Air, Gemma 4 E2B (4-bit, through oMLX) sends its first token after about a second and finishes a short answer with a code block in 3 to 5 seconds.
How I Built It
Open-source AI: Google's Gemma 4 E2B (open weights, 4-bit). I ran it with oMLX (Apple MLX) on my Mac while developing, and the container uses llama.cpp (llama-server with a GGUF build).
App stack: Go 1.27, Templ, HTMX 2 with its SSE extension, goldmark and kardianos/service.
Browser ── HTMX + SSE extension
│ POST /chat → user bubble + empty reply bubble (sse-connect)
│ GET /chat/stream/{id} ← "token" events (append) … "done" (swap in Markdown)
Go (net/http + Templ)
│ POST /v1/chat/completions {stream: true}
Local model server ── Gemma 4 E2B
Streaming works as plain HTML over server-sent events:
- Your browser posts the message, and the server answers with two Templ fragments: your message and an empty reply bubble with
sse-connect="/chat/stream/{id}". - The Go handler streams the model's reply and wraps each chunk in an HTML-escaped
<span>inside atokenevent. HTMX appends each span withhx-swap="beforeend", so you see the answer appear word by word. I wrote no JavaScript for the streaming. - Once the model finishes, the handler sends a
doneevent carrying the reply as server-rendered Markdown (goldmark, raw HTML stripped). HTMX swaps the whole bubble for it, and removing thesse-connectelement closes the stream. - Browsers reconnect a dropped EventSource on their own. If one reconnects for a finished reply, the handler sends back the final HTML and skips the model. A test checks that the model runs once per reply.
A few more details:
- I used no AI SDK. The client reads OpenAI-compatible streaming with a
bufio.Scanneroverdata:lines, in about 150 lines of standard-library Go. -
localchat install --userregisters the app as a launchd agent, a systemd unit or a Windows service, pointed at your.envfile. - I embedded htmx, the SSE extension, the CSS and the icon in the binary, so the app works offline.
- Handler tests run against a fake OpenAI-style server and cover streaming, history, session isolation, XSS-safe Markdown and an offline model. A Playwright test drives a real browser against Gemma, both on the Mac and in Docker.
Prize Categories
- Best Use of Gemma: localchat runs Google's Gemma 4 E2B on your own machine, through MLX on a Mac or a GGUF build in llama.cpp under Docker.

Top comments (0)