<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Bob D</title>
    <description>The latest articles on DEV Community by Bob D (@bdeb1337).</description>
    <link>https://dev.to/bdeb1337</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4153208%2F0a8d2ca4-d225-4b1a-9f9b-800b140ddd94.jpg</url>
      <title>DEV Community: Bob D</title>
      <link>https://dev.to/bdeb1337</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/bdeb1337"/>
    <language>en</language>
    <item>
      <title>localchat: yet another local chat app POC w/GoLang + HTMX</title>
      <dc:creator>Bob D</dc:creator>
      <pubDate>Mon, 05 Oct 2026 06:01:14 +0000</pubDate>
      <link>https://dev.to/bdeb1337/localchat-yet-another-local-chat-app-poc-wgolang-htmx-265d</link>
      <guid>https://dev.to/bdeb1337/localchat-yet-another-local-chat-app-poc-wgolang-htmx-265d</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/hacktoberfest-weekend-2026-10-01"&gt;Hacktoberfest Weekend Challenge: Build for a Friend&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Built
&lt;/h2&gt;

&lt;p&gt;localchat is a chat app for an AI model that runs on your own machine. A Go server renders the UI, and Google's &lt;strong&gt;Gemma 4 E2B&lt;/strong&gt; writes the replies, streamed into your browser token by token. Your messages stay in memory on your laptop for as long as you run the app.&lt;/p&gt;

&lt;p&gt;I talked to a friend before about what a bare-basic and simple Go web app with AI integration would look like, without too many difficult parts. So I spent some time setting this up and kept the project small enough to read in one sitting.&lt;/p&gt;

&lt;p&gt;I picked Gemma while developing because I like running small models on my own hardware. Gemma 4 E2B fits in a few gigabytes of RAM and starts answering in about a second on my laptop. It's fast and easy to develop against while still having realistic output.&lt;/p&gt;

&lt;p&gt;My friend and I get two things:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Code you can read in an evening.&lt;/strong&gt; A HTML form posts to a Go handler, the handler calls the model, and the reply streams back into the page.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A base to build on.&lt;/strong&gt; You can change the system prompt, swap the model with one environment variable, or install the app as a background service on Linux, macOS or Windows. You can also run it as a container if you like.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Repo:&lt;/strong&gt; &lt;a href="https://git.b0b.be/bdeb/localchat" rel="noopener noreferrer"&gt;git.b0b.be/bdeb/localchat&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;cmd/localchat/          entry point: serve + install/start/stop as an OS service
internal/llm/           ~150-line OpenAI-compatible streaming client (stdlib only)
internal/chat/          in-memory conversations, one per browser session
internal/web/           routes, SSE streaming, embedded static assets
internal/web/views/     Templ components
scripts/e2e.py          Playwright browser test against the real model
compose.yaml            app + llama.cpp + Gemma 4 E2B
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Demo
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F78cot3ygsv3378ybyp9p.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F78cot3ygsv3378ybyp9p.png" alt="localchat answering with streamed Markdown" width="800" height="676"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://streamable.com/e/4frgmk" rel="noopener noreferrer"&gt;Video Demonstration&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;On my M3 MacBook Air, Gemma 4 E2B (4-bit, through oMLX) sends its first token after about a second and finishes a short answer with a code block in 3 to 5 seconds.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I Built It
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Open-source AI:&lt;/strong&gt; Google's Gemma 4 E2B (open weights, 4-bit). I ran it with &lt;strong&gt;oMLX&lt;/strong&gt; (Apple MLX) on my Mac while developing, and the container uses &lt;strong&gt;llama.cpp&lt;/strong&gt; (&lt;code&gt;llama-server&lt;/code&gt; with a GGUF build).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;App stack:&lt;/strong&gt; Go 1.27, Templ, HTMX 2 with its SSE extension, goldmark and kardianos/service.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Browser ── HTMX + SSE extension
  │  POST /chat              → user bubble + empty reply bubble (sse-connect)
  │  GET  /chat/stream/{id}  ← "token" events (append) … "done" (swap in Markdown)
Go (net/http + Templ)
  │  POST /v1/chat/completions {stream: true}
Local model server ── Gemma 4 E2B
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Streaming works as plain HTML over server-sent events:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Your browser posts the message, and the server answers with two Templ fragments: your message and an empty reply bubble with &lt;code&gt;sse-connect="/chat/stream/{id}"&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;The Go handler streams the model's reply and wraps each chunk in an HTML-escaped &lt;code&gt;&amp;lt;span&amp;gt;&lt;/code&gt; inside a &lt;code&gt;token&lt;/code&gt; event. HTMX appends each span with &lt;code&gt;hx-swap="beforeend"&lt;/code&gt;, so you see the answer appear word by word. I wrote no JavaScript for the streaming.&lt;/li&gt;
&lt;li&gt;Once the model finishes, the handler sends a &lt;code&gt;done&lt;/code&gt; event carrying the reply as server-rendered Markdown (goldmark, raw HTML stripped). HTMX swaps the whole bubble for it, and removing the &lt;code&gt;sse-connect&lt;/code&gt; element closes the stream.&lt;/li&gt;
&lt;li&gt;Browsers reconnect a dropped EventSource on their own. If one reconnects for a finished reply, the handler sends back the final HTML and skips the model. A test checks that the model runs once per reply.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;A few more details:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;I used no AI SDK. The client reads OpenAI-compatible streaming with a &lt;code&gt;bufio.Scanner&lt;/code&gt; over &lt;code&gt;data:&lt;/code&gt; lines, in about 150 lines of standard-library Go.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;localchat install --user&lt;/code&gt; registers the app as a launchd agent, a systemd unit or a Windows service, pointed at your &lt;code&gt;.env&lt;/code&gt; file.&lt;/li&gt;
&lt;li&gt;I embedded htmx, the SSE extension, the CSS and the icon in the binary, so the app works offline.&lt;/li&gt;
&lt;li&gt;Handler tests run against a fake OpenAI-style server and cover streaming, history, session isolation, XSS-safe Markdown and an offline model. A Playwright test drives a real browser against Gemma, both on the Mac and in Docker.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Prize Categories
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Best Use of Gemma:&lt;/strong&gt; localchat runs Google's Gemma 4 E2B on your own machine, through MLX on a Mac or a GGUF build in llama.cpp under Docker.&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>devchallenge</category>
      <category>weekendchallenge</category>
      <category>hf26challenge</category>
    </item>
  </channel>
</rss>
