<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Vinks Goyal</title>
    <description>The latest articles on DEV Community by Vinks Goyal (@vinksgoyal).</description>
    <link>https://dev.to/vinksgoyal</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4152525%2F686420db-c567-4ba8-bd07-bb8f35b38fbe.jpg</url>
      <title>DEV Community: Vinks Goyal</title>
      <link>https://dev.to/vinksgoyal</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/vinksgoyal"/>
    <language>en</language>
    <item>
      <title>Unstick: One Tiny Action Instead of a Plan</title>
      <dc:creator>Vinks Goyal</dc:creator>
      <pubDate>Sun, 04 Oct 2026 06:11:52 +0000</pubDate>
      <link>https://dev.to/vinksgoyal/unstick-one-tiny-action-instead-of-a-plan-582d</link>
      <guid>https://dev.to/vinksgoyal/unstick-one-tiny-action-instead-of-a-plan-582d</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/hacktoberfest-weekend-2026-10-01"&gt;Hacktoberfest Weekend Challenge: Build for a Friend&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Built
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Unstick&lt;/strong&gt; gives you one tiny action instead of a plan.&lt;/p&gt;

&lt;p&gt;Every productivity tool, every AI chatbot, every roadmap does the same thing: it shows you the whole mountain. Ask ChatGPT how to start a creative portfolio and it gives you five sections, twenty sub-bullets, and an implied skill tree you didn't ask to see. That's technically helpful and completely useless if you're already overwhelmed.&lt;/p&gt;

&lt;p&gt;I built Unstick because I've been that person. I've started my portfolio fourteen times and never finished. Not because I'm lazy — because every time I opened a tool, it showed me how much I didn't know yet, and I closed the tab.&lt;/p&gt;

&lt;p&gt;Unstick does the opposite. You give it a goal. It gives you exactly three cards:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Start here&lt;/strong&gt; — one physical action you can do in five minutes&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;If you have 15 minutes&lt;/strong&gt; — a slightly bigger one&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Not today&lt;/strong&gt; — a greyed-out card that says: &lt;em&gt;"The rest of the plan exists. You don't need to see it."&lt;/em&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That third card is the whole product. The absence of a plan is the feature. No progress bars. No "step 3 of 47." No skill tree.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Who it's for:&lt;/strong&gt; me. I'm the friend I built this for. I don't have someone else to hand it to this weekend, and I'm not going to invent one. I've started my portfolio fourteen times. This is the tool that would have made me finish.&lt;/p&gt;

&lt;h2&gt;
  
  
  Demo
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Live app:&lt;/strong&gt; &lt;a href="https://unstick-api-c7va.onrender.com" rel="noopener noreferrer"&gt;https://unstick-api-c7va.onrender.com&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Type a goal. Click the button. Three cards.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;API endpoint&lt;/strong&gt; (for the terminal-inclined):&lt;/p&gt;

&lt;p&gt;bash&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-X&lt;/span&gt; POST https://unstick-api-c7va.onrender.com/api/unstick &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{"goal":"start a creative portfolio","minutes":15,"energy":2,"location":"home","time":"21:30"}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Video:&lt;/strong&gt; &lt;a href="https://www.youtube.com/video/JtzAP9eyUZc" rel="noopener noreferrer"&gt;https://www.youtube.com/video/JtzAP9eyUZc&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Note on the demo:&lt;/strong&gt; the fine-tuned classifier is trained specifically for &lt;strong&gt;creative portfolio goals&lt;/strong&gt;. For other goals it still returns scope-safe actions, but the domain fit is weak — a "finish my thesis" prompt will give you portfolio-shaped actions. That's a real limitation, not a hidden one. A next step would be broadening the training set.&lt;/p&gt;

&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;GitHub:&lt;/strong&gt; &lt;a href="https://github.com/vinksgoyal/unstick" rel="noopener noreferrer"&gt;github.com/vinksgoyal/unstick&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Structure:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;app/&lt;/code&gt; — FastAPI backend + Tinker + Backboard clients&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;frontend/&lt;/code&gt; — the three-card UI&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;training/&lt;/code&gt; — dataset, label helper, Tinker fine-tune script&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;eval/&lt;/code&gt; — base vs fine-tuned evaluation harness&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;unstick-evidence/&lt;/code&gt; — every screenshot, log, and number cited in this post&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  How I Built It
&lt;/h2&gt;

&lt;p&gt;The prompt says to use open-source AI at the core. Here's what that actually means in this project.&lt;/p&gt;

&lt;h3&gt;
  
  
  The engine: local open-weight models
&lt;/h3&gt;

&lt;p&gt;The candidate generator runs &lt;strong&gt;Gemma 2B&lt;/strong&gt; locally through &lt;strong&gt;Ollama&lt;/strong&gt;. When you type a goal, the app asks Gemma to decompose it into candidate actions. That call happens &lt;strong&gt;on your machine&lt;/strong&gt;. No cloud API sees the raw brain-dump.&lt;/p&gt;

&lt;p&gt;Gemma is not good at this task out of the box. Here's what it produced when I asked it to generate actions for "Start a creative portfolio":&lt;/p&gt;

&lt;p&gt;text&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;"Stand up and stretch for 30 seconds."
"Go to the kitchen and pour a glass of water."
"Pick up a pen and paper and start doodling for 5 minutes."
"Shuffle a deck of cards."
"Take a five-minute walk in your neighborhood."
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;None of those have anything to do with starting a portfolio.&lt;/strong&gt; The base model drifts into generic wellness advice because it can't hold the goal in mind. That's the "before" state.&lt;/p&gt;

&lt;h3&gt;
  
  
  The classifier: Tinker fine-tune
&lt;/h3&gt;

&lt;p&gt;The scope-safety classifier is the actual product. Its job is to look at a (goal, action) pair and decide: &lt;strong&gt;is this action safe to show an overwhelmed person?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I defined six ways an action can fail:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;strong&gt;Failure&lt;/strong&gt;&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;Example&lt;/strong&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;reveals_scope&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;"Sketch your homepage layout"&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;needs_new_skill&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;"Learn Figma"&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;needs_decision&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;"Pick your best five pieces"&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;not_physical&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;"Think about your design style"&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;needs_other_person&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;"Ask a friend for feedback"&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;off_topic&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;"Pour a glass of water" (for a portfolio goal)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;I hand-labeled 128 (goal, state, action) triples across those six categories, plus safe examples. 100 went to training, 44 to a held-out set.&lt;/p&gt;

&lt;p&gt;Then I fine-tuned &lt;strong&gt;Qwen3-8B&lt;/strong&gt; with &lt;strong&gt;LoRA&lt;/strong&gt; on &lt;strong&gt;Tinker&lt;/strong&gt;, using the SDK directly with masked cross-entropy loss on the answer tokens only.&lt;/p&gt;

&lt;h3&gt;
  
  
  Results
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;strong&gt;Model&lt;/strong&gt;&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;Accuracy&lt;/strong&gt;&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;Safe-F1&lt;/strong&gt;&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;Fail-F1&lt;/strong&gt;&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;Latency&lt;/strong&gt;&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;Rows&lt;/strong&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Base Qwen3-8B&lt;/td&gt;
&lt;td&gt;0.2727&lt;/td&gt;
&lt;td&gt;0.4000&lt;/td&gt;
&lt;td&gt;0.3721&lt;/td&gt;
&lt;td&gt;96s&lt;/td&gt;
&lt;td&gt;44&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Fine-tuned (Tinker LoRA)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.7045&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.7778&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.9429&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;118s&lt;/td&gt;
&lt;td&gt;44&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Base model: 27% accuracy. Fine-tune: 70%.&lt;/strong&gt; The fail-class F1 — how well the model catches bad actions — went from 0.37 to 0.94. That's the difference between a demo and something you'd actually use.&lt;/p&gt;

&lt;h3&gt;
  
  
  The memory layer: Backboard
&lt;/h3&gt;

&lt;p&gt;Unstick remembers &lt;strong&gt;how&lt;/strong&gt; you act, never &lt;strong&gt;what&lt;/strong&gt; you said.&lt;/p&gt;

&lt;p&gt;After you complete an action, the app stores an abstract signal in &lt;strong&gt;Backboard&lt;/strong&gt;:&lt;/p&gt;

&lt;p&gt;json&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"goal_id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"abc123"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"energy"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"minutes"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;15&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"time_bucket"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"evening"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"action_type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"file_rename"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"completed"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;No content. No goal text. No draft titles. Just six fields about &lt;em&gt;what kind of moment this was&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;On the next request, Backboard's semantic search returns the signals that match the current state: &lt;em&gt;"This user completes file-rename actions at 9pm with energy 2 about 80% of the time."&lt;/em&gt; That context shapes which candidate the model ranks highest.&lt;/p&gt;

&lt;p&gt;The &lt;code&gt;record()&lt;/code&gt; method enforces the boundary at the code level — any field outside the allowed set raises &lt;code&gt;ValueError&lt;/code&gt;. There is no path by which your draft titles or half-formed thoughts reach Backboard.&lt;/p&gt;

&lt;h3&gt;
  
  
  Deployment: Render
&lt;/h3&gt;

&lt;p&gt;The whole thing runs on &lt;strong&gt;Render&lt;/strong&gt; — FastAPI backend, static frontend, one service. The deployed version uses Tinker for both candidate generation and classification (Ollama doesn't run on Render's free tier).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Live URL:&lt;/strong&gt; &lt;a href="https://unstick-api-c7va.onrender.com/" rel="noopener noreferrer"&gt;unstick-api-c7va.onrender.com&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Cold start is 60–90 seconds on the free tier. Subsequent requests are fast.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Does Open Innovation Matter?
&lt;/h2&gt;

&lt;p&gt;Three specific reasons, not abstract philosophy.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. The content never leaves your machine.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The raw material for a portfolio is unfinished work. Half-drawn things. Abandoned drafts. Songs you're embarrassed by. That's the most personal stuff a person owns, and it's exactly the stuff that has to go into the tool for it to be useful.&lt;/p&gt;

&lt;p&gt;If this ran on a closed API, every abandoned draft and every "I don't know what I'm doing" brain-dump would sit on someone else's server. With Gemma running locally through Ollama, it doesn't. That's not a talking point — it's the reason the tool can be honest with you.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. The classifier is trained on one person's patterns.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I fine-tuned the model on 100 examples I labeled by hand. Not 100,000 scraped examples. Not a commercial dataset. &lt;strong&gt;Mine.&lt;/strong&gt; The fail taxonomy — &lt;code&gt;reveals_scope&lt;/code&gt;, &lt;code&gt;needs_new_skill&lt;/code&gt;, &lt;code&gt;needs_decision&lt;/code&gt; — those categories exist because I noticed those were the exact things that make &lt;em&gt;me&lt;/em&gt; freeze.&lt;/p&gt;

&lt;p&gt;You cannot fine-tune a closed API on one person's patterns. The model belongs to a vendor, and you get the vendor's judgment about what "good" means. With Tinker and an open-weight base model, the model belongs to me.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. It costs nothing to run.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Gemma 2B runs on a laptop with no GPU. Ollama serves it with one command. The local version of Unstick costs $0 per message. I can use it 30 times a night without thinking about the bill. For a tool designed to be used 30 times a night, that matters.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. The model is swappable.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I ran Gemma, then Llama, then Qwen on the same prompt during development. I swapped the classifier architecture between fine-tuning runs. With a closed API, you get one model and one behavior, forever. With open weights, the whole design space stays available.&lt;/p&gt;

&lt;p&gt;That's what open innovation made possible here: a tool that was honest because it was private, and personal because it could be trained on one person.&lt;/p&gt;

&lt;h2&gt;
  
  
  My Agent Session
&lt;/h2&gt;

&lt;p&gt;The whole project was built with &lt;strong&gt;GitHub Copilot's coding agent&lt;/strong&gt;. You can see it in the git history — Copilot commits its own session state:&lt;/p&gt;

&lt;p&gt;text&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;261a8ba Wire Backboard behavioural memory (signals only, never content)
0d7f165 Agent host session 177818fb-a6ca-41c7-be2d-f5476e619709 - turn 1
db22d40 Drop separate static service; FastAPI serves the frontend
242fbbb Serve frontend from FastAPI at root; use relative API base
33be99c Agent host session d5a1d459-9a19-48f0-93b5-d32eae080cc7 - turn 7
5aafbfb Wire Tinker fine-tuned classifier into the app
26d73d6 Fine-tuned Qwen3-8B via Tinker: 27% -&amp;gt; 70% accuracy on held-out set
03d3beb Agent host session 4de8b0e3-5d57-4107-917a-5535793c011b - turn 3
7cb3fd0 Build Unstick scaffold
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Three Copilot agent sessions are visible in the log:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;4de8b0e3-5d57-4107-917a-5535793c011b&lt;/code&gt; — initial FastAPI + frontend scaffold&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;d5a1d459-9a19-48f0-93b5-d32eae080cc7&lt;/code&gt; — Tinker integration&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;177818fb-a6ca-41c7-be2d-f5476e619709&lt;/code&gt; — Backboard wire-up&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;What Copilot did well:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Scaffolded the entire repo from a single mega-prompt&lt;/li&gt;
&lt;li&gt;Wrote the FastAPI backend, frontend, and test suite&lt;/li&gt;
&lt;li&gt;Created the GitHub Actions workflow and initial &lt;code&gt;render.yaml&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Correctly refused to invent Tinker API endpoints (left &lt;code&gt;TODO: verify docs&lt;/code&gt; comments)&lt;/li&gt;
&lt;li&gt;Debugged the &lt;code&gt;Datum&lt;/code&gt; length mismatch after I described the exact error&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Two bugs Copilot missed that I fixed by hand:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Constraint checker&lt;/strong&gt; was treating the user's available minutes as the action's duration, so every action failed the "5 minutes" rule. Fixed by removing the state read and only checking the action text.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tinker&lt;/strong&gt; &lt;code&gt;Datum&lt;/code&gt; &lt;strong&gt;construction&lt;/strong&gt; used only the prompt as &lt;code&gt;model_input&lt;/code&gt; while &lt;code&gt;target_tokens&lt;/code&gt; used the full sequence, causing a 3-token mismatch. Fixed by using &lt;code&gt;full_ids[:-1]&lt;/code&gt; as input and &lt;code&gt;full_ids[1:]&lt;/code&gt; as targets, then adding an assertion to catch future mismatches.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Prize Categories
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Tinker (Thinking Machines)&lt;/strong&gt; — Fine-tuned Qwen3-8B with LoRA on 100 hand-labeled examples. Base model: 27% accuracy. Fine-tune: 70%. Fail-F1: 0.37 → 0.94. Full recipe in &lt;code&gt;training/tinker_finetune.py&lt;/code&gt;, evaluation in &lt;code&gt;eval/run_eval.py&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Render&lt;/strong&gt; — Live app at &lt;a href="https://unstick-api-c7va.onrender.com/" rel="noopener noreferrer"&gt;unstick-api-c7va.onrender.com&lt;/a&gt;. FastAPI backend + static frontend, one service. Deploy config in &lt;code&gt;render.yaml&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Backboard&lt;/strong&gt; — Behavioural memory layer storing only abstract signals (never content). Verified end-to-end: stored a signal, searched for it, got it back with a 0.61 relevance score.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;GitHub Copilot&lt;/strong&gt; — Whole project scaffolded in agent mode. Three session IDs visible in the git log. Two Copilot bugs caught and fixed by hand.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Ollama + Gemma 2B&lt;/strong&gt; — Local open-weight model for candidate generation. Runs offline, costs $0 per message, keeps the raw brain-dump on the user's machine.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;FastAPI + Pydantic&lt;/strong&gt; — Backend. Tests pass, CI green.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Learned
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The hard part wasn't the fine-tune. It was deciding what "bad" means.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Defining the six failure categories was the real product work. Once I could name &lt;em&gt;why&lt;/em&gt; an action was bad — reveals scope, needs a decision, requires a new skill — the classifier became tractable. Without those labels, the fine-tune was just vibes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Open-source AI is a specific advantage, not a slogan.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;For this project, "open" meant four concrete things: the drafts stay on my laptop, the model is trainable on my patterns, running it costs nothing, and I can swap the base model when I want. All four of those are real features, not philosophy.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The base model is 27% wrong in ways that look helpful.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The most important finding: the base model didn't produce garbage. It produced &lt;em&gt;generic wellness advice that looked reasonable.&lt;/em&gt; "Stand up and stretch." "Take a walk." Those are fine sentences. They're just not portfolio actions. A closed API would have shipped that behavior forever because it looks fine in a demo. The fine-tune is what catches it.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Built for the &lt;a href="https://dev.to/challenges/hacktoberfest-weekend-2026-10-01"&gt;Hacktoberfest Weekend Challenge: Build for a Friend&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>weekendchallenge</category>
      <category>hf26challenge</category>
    </item>
  </channel>
</rss>
