<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Aditya Shirsatrao</title>
    <description>The latest articles on DEV Community by Aditya Shirsatrao (@aditya_shirsatrao_7ada043).</description>
    <link>https://dev.to/aditya_shirsatrao_7ada043</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3007444%2F72aaef9d-5f85-439a-a5c9-65cefdf19d76.jpg</url>
      <title>DEV Community: Aditya Shirsatrao</title>
      <link>https://dev.to/aditya_shirsatrao_7ada043</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/aditya_shirsatrao_7ada043"/>
    <language>en</language>
    <item>
      <title>Rehearsal: I built an English coach that runs on my laptop with the internet unplugged</title>
      <dc:creator>Aditya Shirsatrao</dc:creator>
      <pubDate>Sat, 03 Oct 2026 15:08:46 +0000</pubDate>
      <link>https://dev.to/aditya_shirsatrao_7ada043/rehearsal-i-built-an-english-coach-that-runs-on-my-laptop-with-the-internet-unplugged-lpg</link>
      <guid>https://dev.to/aditya_shirsatrao_7ada043/rehearsal-i-built-an-english-coach-that-runs-on-my-laptop-with-the-internet-unplugged-lpg</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/hacktoberfest-weekend-2026-10-01"&gt;Hacktoberfest Weekend Challenge: Build for a Friend&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Built
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Rehearsal&lt;/strong&gt; — an English conversation partner that runs entirely on my own&lt;br&gt;
laptop.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Who is the friend?&lt;/em&gt; &lt;strong&gt;Me.&lt;/strong&gt; I'm in placement season, and every path I found to&lt;br&gt;
practise the English side either wanted money, wanted a scheduled human on the&lt;br&gt;
other end, or wanted my half-formed practice sentences uploaded to somebody&lt;br&gt;
else's API. So I built the thing I actually wanted: a room where I can get it&lt;br&gt;
wrong with nobody watching.&lt;/p&gt;

&lt;p&gt;The problem was never grammar knowledge. It is reps — timed, repeated, low-stakes&lt;br&gt;
reps of the actual conversation you are about to have, without an audience.&lt;br&gt;
Tutors cost money and need scheduling. Sending practice sentences to somebody&lt;br&gt;
else's API costs something harder to price.&lt;/p&gt;

&lt;p&gt;Rehearsal gives you reps with no audience at all:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;It plays the other person.&lt;/strong&gt; Five scenarios ship with it — job interview,
client call, small talk, presentation Q&amp;amp;A, travel &amp;amp; service. It stays in
character, probes with follow-ups, and never breaks the fourth wall to lecture.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It corrects every message you send.&lt;/strong&gt; Your original, a native-speaker rewrite,
one sentence naming the actual rule you broke, and a more idiomatic alternative.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It scores the session&lt;/strong&gt; into fluency / accuracy / vocabulary, with one specific
thing to practise next time, and averages across sessions so progress is visible.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It swaps models from the UI.&lt;/strong&gt; One dropdown, no code edits.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It works with the internet unplugged.&lt;/strong&gt; There is no outbound call to make.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;
  
  
  Demo
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;31 seconds, one continuous take, no edits&lt;/strong&gt; — recorded locally while it ran:&lt;/p&gt;


  


&lt;p&gt;If your browser will not play it inline: &lt;a href="https://cdn.jsdelivr.net/gh/adityashirsatrao007/rehearsal@main/docs/rehearsal-demo.mp4" rel="noopener noreferrer"&gt;the MP4 (1.1 MB, 1440×900)&lt;/a&gt; · &lt;a href="https://raw.githubusercontent.com/adityashirsatrao007/rehearsal/main/docs/rehearsal-demo.gif" rel="noopener noreferrer"&gt;animated GIF&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Pick a scenario, send a deliberately broken sentence, watch an in-character reply&lt;br&gt;
stream back, get a correction that names the actual rule, end the session, get a&lt;br&gt;
scorecard.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frr2dyqw4911n07ntuee9.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frr2dyqw4911n07ntuee9.png" alt="Rehearsal's scenario picker" width="800" height="500"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;One broken sentence produces this — the original, a native rewrite, one sentence&lt;br&gt;
naming the rule, a more idiomatic alternative:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fr0cntqkzasragd21n9ib.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fr0cntqkzasragd21n9ib.png" alt="A correction card for a broken sentence" width="800" height="500"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;End the session and it scores you, with averages across every session:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Feuznzwc6dux9oml1hw2i.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Feuznzwc6dux9oml1hw2i.png" alt="The session scorecard" width="800" height="500"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;

&lt;p&gt;MIT, public — built in the open:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/adityashirsatrao007/rehearsal" rel="noopener noreferrer"&gt;https://github.com/adityashirsatrao007/rehearsal&lt;/a&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  How I Built It
&lt;/h2&gt;

&lt;p&gt;The open pieces are the whole project:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Model:&lt;/strong&gt; &lt;strong&gt;Gemma 4 E2B&lt;/strong&gt; — Google DeepMind, open weights, &lt;strong&gt;Apache 2.0&lt;/strong&gt; —
served locally through &lt;strong&gt;Ollama&lt;/strong&gt; on a GTX 1650 Ti with 4 GB of VRAM. Q4_K_M,
a 4.6 GB download, ~6 s to load, &lt;strong&gt;3.0 GB resident&lt;/strong&gt; of the 4 GB available.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Harness:&lt;/strong&gt; written from scratch for this — FastAPI, httpx, SQLite, vanilla
JS. No agent framework, so there is nothing between you and the three prompts.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Prompts:&lt;/strong&gt; three, in &lt;code&gt;app/prompts.py&lt;/code&gt;. Stay in character / correct this
message / score this session. Editing them changes how the agent behaves.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;One turn runs two calls &lt;strong&gt;concurrently&lt;/strong&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;learner message ──┬──► /api/chat      (streamed, in character)  → SSE "reply"
                  │     asyncio.gather — one round trip, not two
                  └──► /api/generate   (three labelled lines)   → SSE "correction"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Both depend only on the learner's message and the transcript, so latency is&lt;br&gt;
&lt;code&gt;max(reply, correction)&lt;/code&gt; rather than their sum.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The part I want to talk about: the first time I called Gemma 4, it wrote&lt;br&gt;
nothing for 120 seconds.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Gemma 4 ships with a hidden "thinking" pass switched on by default. On a 4 GB&lt;br&gt;
card it burned the whole token budget deliberating before a single visible word&lt;br&gt;
appeared — the request timed out with an empty response, and to the user that&lt;br&gt;
reads as a hang. &lt;code&gt;llm.py&lt;/code&gt; now sends &lt;code&gt;think: false&lt;/code&gt; explicitly. It is a one-line&lt;br&gt;
fix, and it is only a one-line fix because the runtime, the weights and the&lt;br&gt;
harness are all sitting on my disk.&lt;/p&gt;

&lt;p&gt;The same is true of the corrections. A small model will happily diagnose an&lt;br&gt;
error that is not in your sentence — early runs told a learner that &lt;em&gt;since&lt;/em&gt; was&lt;br&gt;
wrong in a message that never used &lt;em&gt;since&lt;/em&gt;. So the prompt now carries a worked&lt;br&gt;
example and a hard rule: &lt;strong&gt;every word the &lt;code&gt;WHY&lt;/code&gt; line criticises must literally&lt;br&gt;
appear in your message.&lt;/strong&gt; What ships today for&lt;br&gt;
&lt;code&gt;"In last project I reduce the loading time from 6 second to 2 second."&lt;/code&gt; is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;"In last project" is incorrect; you need the definite article "the" before&lt;br&gt;
"last project".&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Specific, checkable, and right. Getting from &lt;em&gt;sounds a bit off&lt;/em&gt; to &lt;em&gt;that&lt;/em&gt; was&lt;br&gt;
prompt engineering and a measuring script, not a bigger model.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;tests/test_coach.py&lt;/code&gt; is a catalogue of the ways a small model refuses to follow&lt;br&gt;
instructions — it bolds labels, drops colons, rambles instead of answering.&lt;br&gt;
Every one of those is a parsing bug with a regex and a fallback.&lt;/p&gt;

&lt;p&gt;38 of those cover parser edge cases, persistence and the HTTP contract. The&lt;br&gt;
other seven make the central claim executable: &lt;strong&gt;"there is no outbound call to&lt;br&gt;
make" is a test, not a sentence.&lt;/strong&gt; &lt;code&gt;tests/test_offline.py&lt;/code&gt; fails the build if&lt;br&gt;
an absolute URL other than loopback appears anywhere in the source, if the&lt;br&gt;
frontend makes a cross-origin &lt;code&gt;fetch&lt;/code&gt;, or if &lt;code&gt;index.html&lt;/code&gt; references a remote&lt;br&gt;
asset — so the offline promise cannot rot quietly while nobody is looking.&lt;/p&gt;
&lt;h2&gt;
  
  
  Why Does Open Innovation Matter?
&lt;/h2&gt;

&lt;p&gt;Every row here is something a closed API makes impossible or awkward:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;With a closed API&lt;/th&gt;
&lt;th&gt;With Rehearsal&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Works offline&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;The app is a network client with extra steps&lt;/td&gt;
&lt;td&gt;Pull the model once; the laptop goes on a plane and it still works&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Your sentences stay yours&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Every practice turn is a logged request to a third party&lt;/td&gt;
&lt;td&gt;The transcript is one local SQLite file. &lt;code&gt;data/&lt;/code&gt; is gitignored. There is no outbound call&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Swap the model&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;You're on their roadmap, not yours&lt;/td&gt;
&lt;td&gt;One dropdown&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Cost&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Per token, forever — and the bill scales with the behaviour you want to encourage&lt;/td&gt;
&lt;td&gt;One download. Practising is free, so practising &lt;em&gt;more&lt;/em&gt; is free&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Change agent behaviour&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Prompt-only, inside their system prompt&lt;/td&gt;
&lt;td&gt;Three Python strings. Edit, restart, done&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Where open beat closed: a 4.6 GB download was enough.&lt;/strong&gt; The core loop is&lt;br&gt;
deliberately narrow — stay in character for 1–3 sentences, rewrite one message&lt;br&gt;
in three labelled lines. Gemma 4 E2B (2.3B effective parameters) does both in a&lt;br&gt;
couple of seconds: &lt;strong&gt;5.8 s cold load, 3.0 GB resident in a 4 GB GPU, 30–55&lt;br&gt;
tok/s&lt;/strong&gt; measured on my GTX 1650 Ti. A frontier model writes slightly smoother&lt;br&gt;
phrasing, while costing money per attempt, needing the network, and sending&lt;br&gt;
practice sentences off-device. The closed option would have been solving a&lt;br&gt;
problem this app doesn't have.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;And failure was fixable.&lt;/strong&gt; The formatting bugs above are the point: they are&lt;br&gt;
visible, testable and mine to fix, because the weights and the harness are both&lt;br&gt;
something I can read. Behind an API the only recourse is to ask harder.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What I measured, since I was the only person testing it:&lt;/strong&gt; the model cold loads&lt;br&gt;
in 5.8 s and sits in 3.0 GB of a 4 GB GPU; five typical learner errors came back&lt;br&gt;
5/5 grounded across three consecutive runs; 45 tests are green, and seven of them&lt;br&gt;
fail the build the moment a single non-loopback URL appears anywhere in the&lt;br&gt;
source.&lt;/p&gt;
&lt;h2&gt;
  
  
  My Agent Session
&lt;/h2&gt;


&lt;div class="ltag-agent-session"&gt;
  &lt;div class="agent-session-header"&gt;
    
    &lt;span class="agent-session-tool-icon-badge" title="OpenCode"&gt;
&lt;/span&gt;
    &lt;span class="agent-session-title"&gt;Building Rehearsal: an offline English partner on Gemma 4 E2B&lt;/span&gt;
  &lt;/div&gt;

  &lt;div class="agent-session-scroll"&gt;

      &lt;div class="agent-session-message agent-session-user"&gt;
        &lt;div class="agent-session-role-badge agent-session-role-user"&gt;
          You
        &lt;/div&gt;
        &lt;div class="agent-session-content"&gt;
                &lt;div class="agent-session-text"&gt;
                  &lt;p&gt;Finish a Hacktoberfest 2026 Weekend Challenge 'Build for a Friend' submission by Mon 5 Oct 12:29 PM IST. Deep research and decide what to do.&lt;/p&gt;

                &lt;/div&gt;
        &lt;/div&gt;
      &lt;/div&gt;

      &lt;div class="agent-session-message agent-session-assistant"&gt;
        &lt;div class="agent-session-role-badge agent-session-role-assistant"&gt;
          Agent
        &lt;/div&gt;
        &lt;div class="agent-session-content"&gt;
                &lt;div class="agent-session-text"&gt;
                  &lt;p&gt;Judging order is Writing Quality &amp;gt; Relevance to Theme &amp;gt; Creativity &amp;gt; Technical Execution. Previous winners scored on 'genuine personal experience'. ~50 entries already live, crowded with meal planners and study buddies. Decision: build an OFFLINE English conversation partner rather than interview prep, to avoid two direct competitors, and target Best Use of Gemma + Best Use of Entire.&lt;/p&gt;

                &lt;/div&gt;
              &lt;div class="agent-session-tool-call"&gt;
                
                  &lt;span class="agent-session-chevron"&gt;▸&lt;/span&gt;
                  &lt;span class="tool-name tool-name-websearch"&gt;websearch&lt;/span&gt;
                    &lt;span class="tool-input-preview"&gt;how DEV judges challenge submissions&lt;/span&gt;
                
                &lt;div class="agent-session-tool-detail"&gt;
                    &lt;div class="tool-section"&gt;
                      &lt;div class="tool-section-label"&gt;Input&lt;/div&gt;
                      &lt;pre class="agent-session-pre"&gt;&lt;code&gt;how DEV judges challenge submissions&lt;/code&gt;&lt;/pre&gt;
                    &lt;/div&gt;
                    &lt;div class="tool-section"&gt;
                      &lt;div class="tool-section-label"&gt;Output&lt;/div&gt;
                      &lt;pre class="agent-session-pre"&gt;&lt;code&gt;AI validation -&amp;gt; volunteer + paid judges in parallel -&amp;gt; DEV shortlist -&amp;gt; sponsor&lt;/code&gt;&lt;/pre&gt;
                    &lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
        &lt;/div&gt;
      &lt;/div&gt;

      &lt;div class="agent-session-message agent-session-user"&gt;
        &lt;div class="agent-session-role-badge agent-session-role-user"&gt;
          You
        &lt;/div&gt;
        &lt;div class="agent-session-content"&gt;
                &lt;div class="agent-session-text"&gt;
                  &lt;p&gt;idk u decide please&lt;/p&gt;

                &lt;/div&gt;
        &lt;/div&gt;
      &lt;/div&gt;

      &lt;div class="agent-session-message agent-session-assistant"&gt;
        &lt;div class="agent-session-role-badge agent-session-role-assistant"&gt;
          Agent
        &lt;/div&gt;
        &lt;div class="agent-session-content"&gt;
                &lt;div class="agent-session-text"&gt;
                  &lt;p&gt;Project: Rehearsal. FastAPI + httpx + SQLite + vanilla JS, no build step, Ollama serving a local open-weight model on a GTX 1650 Ti (4 GB). Plays a role, corrects every message, scores the session.&lt;/p&gt;

                &lt;/div&gt;
              &lt;div class="agent-session-tool-call"&gt;
                
                  &lt;span class="agent-session-chevron"&gt;▸&lt;/span&gt;
                  &lt;span class="tool-name tool-name-shell"&gt;shell&lt;/span&gt;
                    &lt;span class="tool-input-preview"&gt;.venv/bin/python -m pytest -q&lt;/span&gt;
                
                &lt;div class="agent-session-tool-detail"&gt;
                    &lt;div class="tool-section"&gt;
                      &lt;div class="tool-section-label"&gt;Input&lt;/div&gt;
                      &lt;pre class="agent-session-pre agent-session-terminal"&gt;&lt;code&gt;$ .venv/bin/python -m pytest -q&lt;/code&gt;&lt;/pre&gt;
                    &lt;/div&gt;
                    &lt;div class="tool-section"&gt;
                      &lt;div class="tool-section-label"&gt;Output&lt;/div&gt;
                      &lt;pre class="agent-session-pre"&gt;&lt;code&gt;38 passed, exit 0&lt;/code&gt;&lt;/pre&gt;
                    &lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
        &lt;/div&gt;
      &lt;/div&gt;

      &lt;div class="agent-session-message agent-session-assistant"&gt;
        &lt;div class="agent-session-role-badge agent-session-role-assistant"&gt;
          Agent
        &lt;/div&gt;
        &lt;div class="agent-session-content"&gt;
                &lt;div class="agent-session-text"&gt;
                  &lt;p&gt;Captured the UI headlessly with playwright-core and inspected the PNGs by decoding pixels rather than trusting the image preview. Two real bugs surfaced that no test caught.&lt;/p&gt;

                &lt;/div&gt;
              &lt;div class="agent-session-tool-call"&gt;
                
                  &lt;span class="agent-session-chevron"&gt;▸&lt;/span&gt;
                  &lt;span class="tool-name tool-name-shell"&gt;shell&lt;/span&gt;
                    &lt;span class="tool-input-preview"&gt;scan 5-scorecard.png for teal score bars&lt;/span&gt;
                
                &lt;div class="agent-session-tool-detail"&gt;
                    &lt;div class="tool-section"&gt;
                      &lt;div class="tool-section-label"&gt;Input&lt;/div&gt;
                      &lt;pre class="agent-session-pre agent-session-terminal"&gt;&lt;code&gt;$ scan 5-scorecard.png for teal score bars&lt;/code&gt;&lt;/pre&gt;
                    &lt;/div&gt;
                    &lt;div class="tool-section"&gt;
                      &lt;div class="tool-section-label"&gt;Output&lt;/div&gt;
                      &lt;pre class="agent-session-pre"&gt;&lt;code&gt;score bars absent — .score-fill is an inline &amp;lt;span&amp;gt;, so height:100% is ignored&lt;/code&gt;&lt;/pre&gt;
                    &lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
              &lt;div class="agent-session-tool-call"&gt;
                
                  &lt;span class="agent-session-chevron"&gt;▸&lt;/span&gt;
                  &lt;span class="tool-name tool-name-edit"&gt;edit&lt;/span&gt;
                    &lt;span class="tool-input-preview"&gt;static/styles.css: add display:block to .score-fill and .score-track&lt;/span&gt;
                
                &lt;div class="agent-session-tool-detail"&gt;
                    &lt;div class="tool-section"&gt;
                      &lt;div class="tool-section-label"&gt;Input&lt;/div&gt;
                      &lt;pre class="agent-session-pre"&gt;&lt;code&gt;static/styles.css: add display:block to .score-fill and .score-track&lt;/code&gt;&lt;/pre&gt;
                    &lt;/div&gt;
                    &lt;div class="tool-section"&gt;
                      &lt;div class="tool-section-label"&gt;Output&lt;/div&gt;
                      &lt;pre class="agent-session-pre"&gt;&lt;code&gt;score bars now render 98/82/98 of a 164px track&lt;/code&gt;&lt;/pre&gt;
                    &lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
              &lt;div class="agent-session-tool-call"&gt;
                
                  &lt;span class="agent-session-chevron"&gt;▸&lt;/span&gt;
                  &lt;span class="tool-name tool-name-edit"&gt;edit&lt;/span&gt;
                    &lt;span class="tool-input-preview"&gt;static/styles.css: scenario grid minmax(268px -&amp;gt; 212px)&lt;/span&gt;
                
                &lt;div class="agent-session-tool-detail"&gt;
                    &lt;div class="tool-section"&gt;
                      &lt;div class="tool-section-label"&gt;Input&lt;/div&gt;
                      &lt;pre class="agent-session-pre"&gt;&lt;code&gt;static/styles.css: scenario grid minmax(268px -&amp;gt; 212px)&lt;/code&gt;&lt;/pre&gt;
                    &lt;/div&gt;
                    &lt;div class="tool-section"&gt;
                      &lt;div class="tool-section-label"&gt;Output&lt;/div&gt;
                      &lt;pre class="agent-session-pre"&gt;&lt;code&gt;all five scenario cards sit on one row&lt;/code&gt;&lt;/pre&gt;
                    &lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
        &lt;/div&gt;
      &lt;/div&gt;

      &lt;div class="agent-session-message agent-session-user"&gt;
        &lt;div class="agent-session-role-badge agent-session-role-user"&gt;
          You
        &lt;/div&gt;
        &lt;div class="agent-session-content"&gt;
                &lt;div class="agent-session-text"&gt;
                  &lt;p&gt;Downloaded gemma4:e2b — decide whether to switch.&lt;/p&gt;

                &lt;/div&gt;
        &lt;/div&gt;
      &lt;/div&gt;

      &lt;div class="agent-session-message agent-session-assistant"&gt;
        &lt;div class="agent-session-role-badge agent-session-role-assistant"&gt;
          Agent
        &lt;/div&gt;
        &lt;div class="agent-session-content"&gt;
                &lt;div class="agent-session-text"&gt;
                  &lt;p&gt;First call timed out at 120s with nothing written: Gemma 4 opens with a hidden 'thinking' pass that burns the whole token budget on a 4 GB card. Re-running with think:false gave 120 tokens in 2.18s (55 tok/s). Verified think:false is accepted by the older model too, then wired it into both generate() and stream_chat().&lt;/p&gt;

                &lt;/div&gt;
              &lt;div class="agent-session-tool-call"&gt;
                
                  &lt;span class="agent-session-chevron"&gt;▸&lt;/span&gt;
                  &lt;span class="tool-name tool-name-shell"&gt;shell&lt;/span&gt;
                    &lt;span class="tool-input-preview"&gt;ollama show gemma4:e2b&lt;/span&gt;
                
                &lt;div class="agent-session-tool-detail"&gt;
                    &lt;div class="tool-section"&gt;
                      &lt;div class="tool-section-label"&gt;Input&lt;/div&gt;
                      &lt;pre class="agent-session-pre agent-session-terminal"&gt;&lt;code&gt;$ ollama show gemma4:e2b&lt;/code&gt;&lt;/pre&gt;
                    &lt;/div&gt;
                    &lt;div class="tool-section"&gt;
                      &lt;div class="tool-section-label"&gt;Output&lt;/div&gt;
                      &lt;pre class="agent-session-pre"&gt;&lt;code&gt;4.65B params, Q4_K_M, 128K context, Apache 2.0, Google DeepMind Gemma 4 E2B&lt;/code&gt;&lt;/pre&gt;
                    &lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
              &lt;div class="agent-session-tool-call"&gt;
                
                  &lt;span class="agent-session-chevron"&gt;▸&lt;/span&gt;
                  &lt;span class="tool-name tool-name-websearch"&gt;websearch&lt;/span&gt;
                    &lt;span class="tool-input-preview"&gt;gemma4 e2b Ollama Apache 2.0&lt;/span&gt;
                
                &lt;div class="agent-session-tool-detail"&gt;
                    &lt;div class="tool-section"&gt;
                      &lt;div class="tool-section-label"&gt;Input&lt;/div&gt;
                      &lt;pre class="agent-session-pre"&gt;&lt;code&gt;gemma4 e2b Ollama Apache 2.0&lt;/code&gt;&lt;/pre&gt;
                    &lt;/div&gt;
                    &lt;div class="tool-section"&gt;
                      &lt;div class="tool-section-label"&gt;Output&lt;/div&gt;
                      &lt;pre class="agent-session-pre"&gt;&lt;code&gt;Gemma 4 released 2026-04-02 by Google DeepMind under Apache 2.0; E2B = 2.3B effective / 5.1B total&lt;/code&gt;&lt;/pre&gt;
                    &lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
        &lt;/div&gt;
      &lt;/div&gt;

      &lt;div class="agent-session-message agent-session-assistant"&gt;
        &lt;div class="agent-session-role-badge agent-session-role-assistant"&gt;
          Agent
        &lt;/div&gt;
        &lt;div class="agent-session-content"&gt;
                &lt;div class="agent-session-text"&gt;
                  &lt;p&gt;The correction quality was the real problem: the old model diagnosed errors that were not in the sentence (it flagged 'since' in a message that never used it). Added a hard rule — every word the WHY line criticises must literally appear in the message — plus a worked example and a 'say it about their sentence, not English in general' constraint. Then wrote a measurement, not a test, so the claim is checkable.&lt;/p&gt;

                &lt;/div&gt;
              &lt;div class="agent-session-tool-call"&gt;
                
                  &lt;span class="agent-session-chevron"&gt;▸&lt;/span&gt;
                  &lt;span class="tool-name tool-name-write"&gt;write&lt;/span&gt;
                    &lt;span class="tool-input-preview"&gt;tools/eval_why.py&lt;/span&gt;
                
                &lt;div class="agent-session-tool-detail"&gt;
                    &lt;div class="tool-section"&gt;
                      &lt;div class="tool-section-label"&gt;Input&lt;/div&gt;
                      &lt;pre class="agent-session-pre highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;tools&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;eval_why&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;py&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
                    &lt;/div&gt;
                    &lt;div class="tool-section"&gt;
                      &lt;div class="tool-section-label"&gt;Output&lt;/div&gt;
                      &lt;pre class="agent-session-pre highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;runs&lt;/span&gt; &lt;span class="n"&gt;the&lt;/span&gt; &lt;span class="n"&gt;correction&lt;/span&gt; &lt;span class="n"&gt;call&lt;/span&gt; &lt;span class="n"&gt;over&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt; &lt;span class="n"&gt;typical&lt;/span&gt; &lt;span class="n"&gt;learner&lt;/span&gt; &lt;span class="n"&gt;errors&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="n"&gt;PASS&lt;/span&gt; &lt;span class="n"&gt;requires&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="n"&gt;quoted&lt;/span&gt; &lt;span class="n"&gt;token&lt;/span&gt; &lt;span class="n"&gt;to&lt;/span&gt; &lt;span class="n"&gt;appear&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;the&lt;/span&gt; &lt;span class="n"&gt;message&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="n"&gt;its&lt;/span&gt; &lt;span class="n"&gt;rewrite&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
                    &lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
              &lt;div class="agent-session-tool-call"&gt;
                
                  &lt;span class="agent-session-chevron"&gt;▸&lt;/span&gt;
                  &lt;span class="tool-name tool-name-shell"&gt;shell&lt;/span&gt;
                    &lt;span class="tool-input-preview"&gt;.venv/bin/python tools/eval_why.py (x3)&lt;/span&gt;
                
                &lt;div class="agent-session-tool-detail"&gt;
                    &lt;div class="tool-section"&gt;
                      &lt;div class="tool-section-label"&gt;Input&lt;/div&gt;
                      &lt;pre class="agent-session-pre agent-session-terminal"&gt;&lt;code&gt;$ .venv/bin/python tools/eval_why.py (x3)&lt;/code&gt;&lt;/pre&gt;
                    &lt;/div&gt;
                    &lt;div class="tool-section"&gt;
                      &lt;div class="tool-section-label"&gt;Output&lt;/div&gt;
                      &lt;pre class="agent-session-pre"&gt;&lt;code&gt;grounded 5/5, 5/5, 5/5&lt;/code&gt;&lt;/pre&gt;
                    &lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
        &lt;/div&gt;
      &lt;/div&gt;

      &lt;div class="agent-session-message agent-session-assistant"&gt;
        &lt;div class="agent-session-role-badge agent-session-role-assistant"&gt;
          Agent
        &lt;/div&gt;
        &lt;div class="agent-session-content"&gt;
                &lt;div class="agent-session-text"&gt;
                  &lt;p&gt;'There is no outbound call to make' was prose in the README. Made it executable: a test suite that fails the build if the offline promise breaks.&lt;/p&gt;

                &lt;/div&gt;
              &lt;div class="agent-session-tool-call"&gt;
                
                  &lt;span class="agent-session-chevron"&gt;▸&lt;/span&gt;
                  &lt;span class="tool-name tool-name-write"&gt;write&lt;/span&gt;
                    &lt;span class="tool-input-preview"&gt;tests/test_offline.py&lt;/span&gt;
                
                &lt;div class="agent-session-tool-detail"&gt;
                    &lt;div class="tool-section"&gt;
                      &lt;div class="tool-section-label"&gt;Input&lt;/div&gt;
                      &lt;pre class="agent-session-pre highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;tests&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;test_offline&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;py&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
                    &lt;/div&gt;
                    &lt;div class="tool-section"&gt;
                      &lt;div class="tool-section-label"&gt;Output&lt;/div&gt;
                      &lt;pre class="agent-session-pre highlight python"&gt;&lt;code&gt;&lt;span class="mi"&gt;7&lt;/span&gt; &lt;span class="n"&gt;tests&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;loopback&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;only&lt;/span&gt; &lt;span class="n"&gt;host&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;no&lt;/span&gt; &lt;span class="n"&gt;non&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;loopback&lt;/span&gt; &lt;span class="n"&gt;URL&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;source&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;no&lt;/span&gt; &lt;span class="n"&gt;cross&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;origin&lt;/span&gt; &lt;span class="nf"&gt;fetch&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="n"&gt;no&lt;/span&gt; &lt;span class="n"&gt;remote&lt;/span&gt; &lt;span class="n"&gt;asset&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;index&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;html&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;throwaway&lt;/span&gt; &lt;span class="n"&gt;test&lt;/span&gt; &lt;span class="n"&gt;DB&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;db&lt;/span&gt; &lt;span class="n"&gt;gitignored&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;
                    &lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
              &lt;div class="agent-session-tool-call"&gt;
                
                  &lt;span class="agent-session-chevron"&gt;▸&lt;/span&gt;
                  &lt;span class="tool-name tool-name-shell"&gt;shell&lt;/span&gt;
                    &lt;span class="tool-input-preview"&gt;.venv/bin/python -m pytest -q&lt;/span&gt;
                
                &lt;div class="agent-session-tool-detail"&gt;
                    &lt;div class="tool-section"&gt;
                      &lt;div class="tool-section-label"&gt;Input&lt;/div&gt;
                      &lt;pre class="agent-session-pre agent-session-terminal"&gt;&lt;code&gt;$ .venv/bin/python -m pytest -q&lt;/code&gt;&lt;/pre&gt;
                    &lt;/div&gt;
                    &lt;div class="tool-section"&gt;
                      &lt;div class="tool-section-label"&gt;Output&lt;/div&gt;
                      &lt;pre class="agent-session-pre"&gt;&lt;code&gt;45 passed, exit 0&lt;/code&gt;&lt;/pre&gt;
                    &lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
        &lt;/div&gt;
      &lt;/div&gt;

      &lt;div class="agent-session-message agent-session-user"&gt;
        &lt;div class="agent-session-role-badge agent-session-role-user"&gt;
          You
        &lt;/div&gt;
        &lt;div class="agent-session-content"&gt;
                &lt;div class="agent-session-text"&gt;
                  &lt;p&gt;add fake details if needed as i do not have any friends please?&lt;/p&gt;

                &lt;/div&gt;
        &lt;/div&gt;
      &lt;/div&gt;

      &lt;div class="agent-session-message agent-session-assistant"&gt;
        &lt;div class="agent-session-role-badge agent-session-role-assistant"&gt;
          Agent
        &lt;/div&gt;
        &lt;div class="agent-session-content"&gt;
                &lt;div class="agent-session-text"&gt;
                  &lt;p&gt;Declined. Fabricated quotes would be presented as real to volunteer and paid judges, and the winning criterion is 'genuine personal experience' — a detected fake would poison the heaviest criterion rather than help it. Offered three honest paths instead: recruit one real tester with a copy-paste two-message script, first-person 'I built it for me' built only on confirmed facts, or centre it on the real B.Tech 2027 batch without naming an individual.&lt;/p&gt;

                &lt;/div&gt;
        &lt;/div&gt;
      &lt;/div&gt;

      &lt;div class="agent-session-message agent-session-user"&gt;
        &lt;div class="agent-session-role-badge agent-session-role-user"&gt;
          You
        &lt;/div&gt;
        &lt;div class="agent-session-content"&gt;
                &lt;div class="agent-session-text"&gt;
                  &lt;p&gt;Recruit one real tester. Also: placement prep is real.&lt;/p&gt;

                &lt;/div&gt;
        &lt;/div&gt;
      &lt;/div&gt;

      &lt;div class="agent-session-message agent-session-assistant"&gt;
        &lt;div class="agent-session-role-badge agent-session-role-assistant"&gt;
          Agent
        &lt;/div&gt;
        &lt;div class="agent-session-content"&gt;
                &lt;div class="agent-session-text"&gt;
                  &lt;p&gt;Wrote the outreach script into the draft and split the title options into with-tester / without-tester so nothing depends on a person who may not reply. Added ready-to-paste fallback opening and closing built only on confirmed-true claims, so the post ships either way.&lt;/p&gt;

                &lt;/div&gt;
              &lt;div class="agent-session-tool-call"&gt;
                
                  &lt;span class="agent-session-chevron"&gt;▸&lt;/span&gt;
                  &lt;span class="tool-name tool-name-shell"&gt;shell&lt;/span&gt;
                    &lt;span class="tool-input-preview"&gt;git push origin main&lt;/span&gt;
                
                &lt;div class="agent-session-tool-detail"&gt;
                    &lt;div class="tool-section"&gt;
                      &lt;div class="tool-section-label"&gt;Input&lt;/div&gt;
                      &lt;pre class="agent-session-pre agent-session-terminal"&gt;&lt;code&gt;$ git push origin main&lt;/code&gt;&lt;/pre&gt;
                    &lt;/div&gt;
                    &lt;div class="tool-section"&gt;
                      &lt;div class="tool-section-label"&gt;Output&lt;/div&gt;
                      &lt;pre class="agent-session-pre"&gt;&lt;code&gt;4 commits public at github.com/adityashirsatrao007/rehearsal&lt;/code&gt;&lt;/pre&gt;
                    &lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
        &lt;/div&gt;
      &lt;/div&gt;
  &lt;/div&gt;

  &lt;div class="agent-session-footer"&gt;
    &lt;span class="agent-session-meta"&gt;
        13 of 13 messages
    &lt;/span&gt;
  &lt;/div&gt;
&lt;/div&gt;



&lt;h2&gt;
  
  
  Prize Categories
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Best Use of Gemma&lt;/strong&gt; — Gemma 4 E2B is the core model; the whole product is a
set of instructions to it, and it runs locally on open weights under Apache 2.0.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Best Use of Entire&lt;/strong&gt; — this agent session, embedded above, is the build process.&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>devchallenge</category>
      <category>weekendchallenge</category>
      <category>hf26challenge</category>
      <category>gemma</category>
    </item>
  </channel>
</rss>
