<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Abhash Chakraborty</title>
    <description>The latest articles on DEV Community by Abhash Chakraborty (@abhash_chakraborty).</description>
    <link>https://dev.to/abhash_chakraborty</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F2818938%2F42ffd28f-8e25-4530-b294-e4254d69ea16.png</url>
      <title>DEV Community: Abhash Chakraborty</title>
      <link>https://dev.to/abhash_chakraborty</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/abhash_chakraborty"/>
    <language>en</language>
    <item>
      <title>Fine-tuning a 0.8B model to beat hosted LLMs at one job</title>
      <dc:creator>Abhash Chakraborty</dc:creator>
      <pubDate>Mon, 05 Oct 2026 16:07:46 +0000</pubDate>
      <link>https://dev.to/abhash_chakraborty/fine-tuning-a-08b-model-to-beat-hosted-llms-at-one-job-4ajl</link>
      <guid>https://dev.to/abhash_chakraborty/fine-tuning-a-08b-model-to-beat-hosted-llms-at-one-job-4ajl</guid>
      <description>&lt;p&gt;I needed a model for one narrow decision. When a new fact about someone comes in, does it replace something already known? "I moved to Berlin" should retire "I live in Pune". "Watched two episodes of The Bear" should leave "works at Infosys" alone.&lt;/p&gt;

&lt;p&gt;It sounds trivial, but any assistant with long-term memory makes this call all the time, and both kinds of mistake hurt. Miss an update and it keeps repeating stale facts. Replace too eagerly and it throws away things that are still true.&lt;/p&gt;

&lt;p&gt;This post covers the models I compared, how I fine-tuned the best starting point, how I tested it, and what it took to train. The short version: a 0.8B open model, fine-tuned on 6,000 examples on a single T4, beat every hosted LLM I tested at this one job.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9kfyd0j285qzii50727h.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9kfyd0j285qzii50727h.webp" alt="Bar chart: the fine-tuned kev-mem scores 0.954 mean ROC AUC on memory-update decisions, ahead of Gemini 3.7 Flash (0.892), Gemma 4 31B (0.887), GPT-5.4 mini (0.873), Kev as released (0.852), Jev (0.834) and Laya (0.711)" width="799" height="453"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Every model on the same held-out tests. For my definition of an update, not a general ranking.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Why a System-One model
&lt;/h2&gt;

&lt;p&gt;The obvious approach is to prompt an LLM and parse what comes back. That works, but it's slow, every decision is an API call, and the answer is free text that you then have to interpret.&lt;/p&gt;

&lt;p&gt;TypeSafe's Jev takes a different approach. It's a "System One" model: instead of writing text, it answers typed questions with calibrated probabilities, like a choice from a list or a yes/no. For this task I send it the new fact and up to eight existing ones, and ask two questions about each existing fact:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;How does the new fact relate to it: duplicate, updates, elaborates, related or unrelated?&lt;/li&gt;
&lt;li&gt;Is the existing fact now outdated?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The score I use is the average of the two: &lt;code&gt;p_update = (P(updates) + P(outdated)) / 2&lt;/code&gt;. At 0.70 or above, the old fact is replaced.&lt;/p&gt;

&lt;h2&gt;
  
  
  The candidates
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Jev&lt;/strong&gt; (hosted, by TypeSafe). Solid, but tuned for general decisions rather than my definition of an update.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Laya&lt;/strong&gt;, an open 421M ModernBERT model that answers in Jev's format. Fast and light, but off the shelf it replaced too many everyday events.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Kev&lt;/strong&gt;, Jared Palmer's open-weight model, compatible with Jev. I tried two sizes. Kev-0.8B was the best of everything as released on stated updates. Kev-4B scored lower on that, and on the T4 it ran out of memory on longer requests and took 5.5 seconds per decision.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For reference I also ran a zero-shot NLI model (DeBERTa-v3-large) and several hosted LLMs, each prompted with the same definition of an update. Kev-0.8B was the clear starting point: the best as released, and the only size that both trains and runs on a 16 GB T4.&lt;/p&gt;

&lt;h2&gt;
  
  
  The data
&lt;/h2&gt;

&lt;p&gt;Each training example is one request: a new fact, up to eight existing facts, and the typed questions, with a label for every question. I built them from two pools:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;6,818 logged Jev decisions.&lt;/strong&gt; Training on these is distillation: Kev learns to make Jev's calls on real requests.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;10,219 decisions with ground truth.&lt;/strong&gt; I simulated 48 lives of a year each. People move city, change jobs, partners and managers, switch gyms, cars, courses, diets and laptops, reschedule appointments, mention crucial facts and make small talk. Because the simulator knows what actually changed, every pair has the right answer: same attribute with a new value is an update, the same value is a duplicate, anything else is related or unrelated. That's 53,870 fact pairs, 2,298 of them real updates.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The 6,000 training examples took every decision that contains a real update first, then filled up from the shuffled mix. Nothing from the test data went in.&lt;/p&gt;

&lt;h2&gt;
  
  
  Stage 1: supervised fine-tuning
&lt;/h2&gt;

&lt;p&gt;LoRA on Qwen3.5-0.8B-Base, starting from Kev's released weights: rank 16, alpha 32, dropout 0.05, on every linear and linear-attention projection, plus a small decision head. One epoch over the 6,000 examples at a learning rate of 5e-5, with label smoothing of 0.05.&lt;/p&gt;

&lt;p&gt;The GPU was a single NVIDIA T4 with 16 GB. It has no bf16, so everything ran in fp32, with batch 1, gradient accumulation of 8 and gradient checkpointing. The requests are short (87 tokens on average, 138 at the 95th percentile), so nothing was truncated. That came to 750 optimizer steps and 5.52M tokens in 6.55 hours, peaking at 5.09 GB of GPU memory.&lt;/p&gt;

&lt;h2&gt;
  
  
  Stage 2: reward-ranked RL (ReST)
&lt;/h2&gt;

&lt;p&gt;Supervised training teaches the format and the definition, but it doesn't know which mistakes matter more. So I let the stage-1 model make the real decision, replace or keep at the 0.70 threshold, on 2,000 training decisions (10,574 pairs), and scored every pair:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;+1 for catching a real update&lt;/li&gt;
&lt;li&gt;−1 for missing one&lt;/li&gt;
&lt;li&gt;−2 for replacing a fact that was still true&lt;/li&gt;
&lt;li&gt;+0.25 for correctly keeping a fact&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The mean reward was 0.265. There were 15 wrong replacements and 65 missed updates, spread over 52 decisions with a negative reward. Those 52 were relabelled with the truth (repeated twice when they included a −2) and mixed with replayed stage-1 examples, for 366 records in total. A second, gentler pass at a learning rate of 1e-5 took 46 steps, 0.34M tokens and 23 minutes, peaking at 4.98 GB.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxihm3elhjp9fjabmun45.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxihm3elhjp9fjabmun45.webp" alt="The training run on one NVIDIA T4: stage 1, supervised fine-tuning on 6,000 decisions, 750 steps, 5.52M tokens, 6.55 hours, 5.09 GB peak; stage 2, ReST on 366 records, 46 steps, 0.34M tokens, 23 minutes, 4.98 GB peak; 885M parameters of which 11.3M trained" width="800" height="400"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Both stages on one T4. About 12 GPU-hours for the whole project, baseline tests included.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;I picked the final checkpoint on six separate validation lives (400 decisions), where it scored 0.9991 AUC against 0.9983 for stage 1. The test sets were never used to choose anything.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I tested it
&lt;/h2&gt;

&lt;p&gt;Two held-out tests, both kept out of training:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A two-year simulated life&lt;/strong&gt; with template text: 122 decisions, 428 pairs, 186 updates, 45 of them stated outright.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A live run&lt;/strong&gt; where an LLM wrote the facts from a three-month simulated life, so they read more like real conversation: 127 decisions, 498 pairs, 104 updates, 36 stated. This is the harder and more realistic one.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A stated update says it directly ("I moved to Berlin"). An implied one doesn't, like a new address in another city with no mention of moving. I report both.&lt;/p&gt;

&lt;p&gt;The main metric is ROC AUC. Take one real update and one non-update at random: AUC is the chance the model scores the real update higher. 0.5 is a coin flip and 1.0 is perfect. It isn't accuracy, so 0.954 doesn't mean "right 95% of the time". I also check precision and recall at the 0.70 threshold, because that's where the model actually acts.&lt;/p&gt;

&lt;h2&gt;
  
  
  Results
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fva096zy9wjkhmvttxn96.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fva096zy9wjkhmvttxn96.webp" alt="Line chart: on the live test, Kev's ROC AUC for stated updates goes from 0.752 as released to 0.929 after SFT and 0.945 after ReST, above Jev's 0.781; for any update it goes from 0.708 to 0.873 and 0.872" width="800" height="467"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The live test. Most of the gain comes from SFT; ReST makes the model more careful about replacing.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;On the live test, stated updates went from 0.752 as released to 0.929 after stage 1 and 0.945 after stage 2. For any update, including implied ones, it went from 0.708 to 0.873 and 0.872. Jev scores 0.781 and 0.570 on the same test. On the two-year life the stated score went from 0.952 to 0.962, and for any update from 0.583 to 0.840.&lt;/p&gt;

&lt;p&gt;At the 0.70 threshold on the live test, precision was 1.00: it never replaced a fact that was still true. The cost is a recall of 0.14, so it only replaces when it's very sure. Kev and Laya as released never crossed the threshold at all. For this job, a cautious model is the right default, and a lower threshold chosen on validation data recovers more of the misses.&lt;/p&gt;

&lt;p&gt;Averaged over both tests on stated updates, with every hosted LLM given the same definition:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;my fine-tuned Kev (kev-mem): &lt;strong&gt;0.954&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;Gemini 3.7 Flash: 0.892&lt;/li&gt;
&lt;li&gt;Gemma 4 31B: 0.887&lt;/li&gt;
&lt;li&gt;GPT-5.4 mini: 0.873&lt;/li&gt;
&lt;li&gt;Kev as released: 0.852&lt;/li&gt;
&lt;li&gt;Jev: 0.834&lt;/li&gt;
&lt;li&gt;DeBERTa-v3-large NLI, zero-shot: 0.815&lt;/li&gt;
&lt;li&gt;Laya: 0.711&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Size and cost
&lt;/h2&gt;

&lt;p&gt;The model has about 885M parameters: 873.4M in the base, 10.8M in the LoRA adapter and 0.5M in the decision head, so only 11.3M were trained. The adapter is 43 MB. It runs in about 4 GB of GPU memory and takes 0.81 seconds per decision on the T4, without any of the optimised kernels. Counting the baseline tests and a second stage-1 run, the whole project used about 12 GPU-hours.&lt;/p&gt;

&lt;h2&gt;
  
  
  Caveats
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;This is my use case and my definition of an update. Jev and Kev are general-purpose models and the hosted LLMs were only prompted, so a lot of the gain comes from specialising. That's the point, but it isn't a general ranking.&lt;/li&gt;
&lt;li&gt;These are single training runs, with no confidence intervals yet.&lt;/li&gt;
&lt;li&gt;Jev answered 234 of the 249 test decisions; the rest failed after retries.&lt;/li&gt;
&lt;li&gt;The test lives are simulated, apart from the LLM-written text of the live run.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Try it
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://huggingface.co/abhash-chakraborty/kev-mem-0.8b" rel="noopener noreferrer"&gt;kev-mem-0.8b on Hugging Face&lt;/a&gt;, served through Kev's own server, so any Jev client works with it&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.kaggle.com/models/abhashchakraborty/kev-mem" rel="noopener noreferrer"&gt;kev-mem on Kaggle&lt;/a&gt; and the &lt;a href="https://www.kaggle.com/code/abhashchakraborty/kev-mem-system-1-benchmark" rel="noopener noreferrer"&gt;benchmark notebook&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;The &lt;a href="https://www.kaggle.com/benchmarks/tasks/abhashchakraborty/memory-update-detection" rel="noopener noreferrer"&gt;memory-update detection benchmark&lt;/a&gt; for hosted LLMs&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Thanks to TypeSafe for Jev, Jared Palmer for Kev, the Laya authors for open-sourcing their work (Apache-2.0), and the Qwen team for Qwen3.5.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://abhashchakraborty.tech/writing/fine-tuning-a-small-model-for-memory-updates" rel="noopener noreferrer"&gt;abhashchakraborty.tech&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>machinelearning</category>
      <category>llm</category>
      <category>finetuning</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Building a portfolio that feels fast</title>
      <dc:creator>Abhash Chakraborty</dc:creator>
      <pubDate>Mon, 05 Oct 2026 12:05:04 +0000</pubDate>
      <link>https://dev.to/abhash_chakraborty/building-a-portfolio-that-feels-fast-29o1</link>
      <guid>https://dev.to/abhash_chakraborty/building-a-portfolio-that-feels-fast-29o1</guid>
      <description>&lt;p&gt;I started rebuilding my portfolio this summer with a short brief: it should feel like mine, it should open fast, and I should be able to change any of it without a deploy. The first two pull against each other. The third turns what could have been a static site into one with a backend.&lt;/p&gt;

&lt;p&gt;This is a write-up of how the pieces fit together, the problems that took longest, and what I settled on. Most of it applies to any content site that wants to be quick without being plain.&lt;/p&gt;

&lt;h2&gt;
  
  
  The shape of the system
&lt;/h2&gt;

&lt;p&gt;Almost every request is someone reading, so I optimised for reads and accepted a bit more work on writes. Pages are React Server Components rendered on Vercel and cached as static HTML. A visitor's browser never talks to the database. It gets HTML, a few small interactive islands, and images from a separate storage domain.&lt;/p&gt;

&lt;p&gt;Content lives in Convex: my profile, projects, posts, likes and the game leaderboard. Images and files live in Cloudflare R2 on their own domain. A private editor, which I call the Studio, writes to both. Product analytics end up in PostHog, but only after passing through the site's own pipeline.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwvsymq9h9fgl2di3ff04.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwvsymq9h9fgl2di3ff04.webp" alt="System overview: the visitor, Next.js on Vercel with static pages and client islands, Convex, PostHog, Cloudflare R2 and the Studio" width="800" height="507"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Reads come from cached HTML, writes come from the Studio, and events pass through an outbox before they leave.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  A rich design on a light page
&lt;/h2&gt;

&lt;p&gt;The design leans on texture: brushed silver plates, a monochrome portrait that picks up colour on hover, and a holographic glint on project cards. Each of those could easily have cost a library and a few hundred kilobytes.&lt;/p&gt;

&lt;p&gt;The rule I followed was simple. If something can be HTML and CSS, it is. JavaScript runs only in small client components that need state: the header, the command palette, the live status strip and the terminal. Each game is its own chunk and loads when you start it. Everything else ships as markup.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzhphln1xatidhjekqa6j.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzhphln1xatidhjekqa6j.webp" alt="Design system board: silver, ink and green colour tokens, the rainbow holo foil palette, the Geist type scale, components and corner radii" width="800" height="560"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The tokens behind the look: eight colours, the holo foil that only appears on hover, one type family and three radii.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The rainbow foil shows how to keep an effect cheap. It's a conic gradient seen through a striped mask, and hovering only moves one pre-drawn layer with a transform, so nothing repaints while the pointer moves. On touch screens, where there's no hover, the same layers play as a slow pulse.&lt;/p&gt;

&lt;p&gt;Scrolling was the subtle part. A few sections animate as they scroll into view, which means the browser marks styles dirty on every frame. Early on, the scroll rail read element positions inside its scroll handler, which forced a full layout each frame and made the top of the page stutter. The fix was to measure once on load and on resize, cache the numbers, and only ever animate opacity and transforms so the compositor does the work.&lt;/p&gt;

&lt;p&gt;Order matters as much as size. The hero portrait and the status request are preloaded so they download alongside the HTML, and anything optional, like analytics, waits until the browser is idle. On a cold load from India, the first byte arrives in 83 ms and the largest paint lands at about 0.3 seconds.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzygd7h54qvu9w1n4t8ra.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzygd7h54qvu9w1n4t8ra.webp" alt="Request waterfall for one cold load of the homepage, with first paint at 276 ms and largest paint at 304 ms" width="799" height="457"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;One cold load in Edge with an empty cache, in milliseconds from the start of navigation.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Editing without redeploying
&lt;/h2&gt;

&lt;p&gt;I wanted to write and fix things from a browser, but I didn't want page views to depend on a database call. The answer is a cache that the editor is allowed to break.&lt;/p&gt;

&lt;p&gt;All public reads go through one small content layer. Each query is cached under a shared tag and revalidated hourly. When I save something, the save clears that tag, so the next request renders fresh HTML while every other request keeps coming from cache. If the database is unreachable, the profile and projects fall back to copies that ship with the code, so the site stays up even when the backend doesn't.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwniotd0c3hjpyd5v1ut6.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwniotd0c3hjpyd5v1ut6.webp" alt="Publishing flow in four steps: the editor, a server action that checks the session and sanitizes HTML, a Convex transaction, and the site clearing its content cache" width="799" height="333"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;What one save does: sanitize once, write once, and invalidate only the cached content that changed.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Posts and project write-ups use a Tiptap editor. It stores two things: the editor's structured document, which is what I edit, and HTML, which is what the site renders. The HTML is sanitized on the server against an allowlist on every save, so the site can render it directly without trusting the editor.&lt;/p&gt;

&lt;p&gt;Charts, metrics, flow diagrams and galleries were harder. They're React components, and I didn't want stored HTML to be able to render arbitrary components. In storage each one is an empty placeholder carrying its type and a small JSON payload. When a page renders, the HTML is split at those placeholders, the payload is validated, and the real component goes in its place. A malformed block renders nothing instead of breaking the page.&lt;/p&gt;

&lt;h2&gt;
  
  
  Previews that never leak
&lt;/h2&gt;

&lt;p&gt;The Studio shows the real page next to the editor, not a lookalike. That uses the framework's draft mode, which skips the cache and lets pages read unpublished content. The risk is obvious: draft mode is carried by the browser, and anything a browser carries can be copied.&lt;/p&gt;

&lt;p&gt;So draft mode alone isn't enough. A page only reads drafts when the request also carries a valid editor session, and the public queries never return unpublished posts at all. Previews are also scoped to the tab that asked for one, so opening the site in another tab after an editing session shows exactly what visitors see.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fybr04g4bjjv9usu6udca.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fybr04g4bjjv9usu6udca.webp" alt="The Studio's Site editor with the homepage sections on the left and the live page on the right" width="800" height="507"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The Site editor: each section's controls on the left, the real page in draft mode on the right.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  A terminal that owns the keyboard
&lt;/h2&gt;

&lt;p&gt;The site has a terminal. Pages behave like folders, so &lt;code&gt;cd projects&lt;/code&gt; takes you there, and it has six games: Snake, Space Defender, Type Racer, 2048, Minesweeper and a daily 2048 challenge.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftiw8re99ljobi5czwac9.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftiw8re99ljobi5czwac9.webp" alt="The terminal's game picker showing six games" width="800" height="507"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The game picker. Each card acts out its game on hover.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The hardest bug came from a browser extension. Some copy-paste "unlocker" extensions stop every key press before the page sees it. In Edge that meant Space scrolled the page and the games ignored the arrow keys. Listening on the window in the capture phase fixed it, because those listeners run before anything attached to the document. Focus needed the same care: games take focus without scrolling, and the terminal never calls &lt;code&gt;scrollIntoView&lt;/code&gt;, which moves the page as well as the container.&lt;/p&gt;

&lt;p&gt;Scores taught me about leaving. A run used to count only when it ended, so closing the tab mid-game threw it away. Now each game saves its score so far when you exit, switch tabs or close the window. Posting lives in a module that outlives the terminal, keeps a verification token ready while you play, and sends the last score with a keepalive request so it survives the page closing. The server still checks every score against the game's rules.&lt;/p&gt;

&lt;h2&gt;
  
  
  Analytics without tracking people
&lt;/h2&gt;

&lt;p&gt;I want to know which pages get read and where people give up, not who they are. That rules out most off-the-shelf setups, so the site runs its own small pipeline and only forwards anonymised events.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5bbxb0uob0o6ut8ndnyp.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5bbxb0uob0o6ut8ndnyp.webp" alt="Analytics pipeline: the browser, a same-origin endpoint that hashes the visitor, a Convex transaction, and delivery to PostHog" width="799" height="313"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;From a page view to a chart. Nothing that identifies a person is stored along the way.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The browser sends a page view once the page is idle, then the time the page was actually on screen and how far down it was read. Read depth comes from four invisible markers watched by an IntersectionObserver, so nothing measures the layout while you scroll. On the way out, a beacon carries the last update, so closing the tab doesn't lose it.&lt;/p&gt;

&lt;p&gt;On the server a visitor is a one-way hash. By default it's built from values that change every day, so it can count visits without following anyone from one day to the next, and the inputs themselves are never stored. Visitors who choose to be remembered get a random id instead, which is what makes returning-visitor numbers possible. Each event is written to the database together with the counters it affects, in one transaction. A scheduled job then delivers it to PostHog and backs off and retries if PostHog is having a bad day.&lt;/p&gt;

&lt;p&gt;Counting correctly took more thought than collecting. A reload shouldn't add a reader, so each post keeps one row per reader per day, and later updates only add the difference. The number shown under a post counts each reader once. Shares count when someone actually opens the shared link rather than when a button is pressed, because copying a link isn't sharing it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd keep
&lt;/h2&gt;

&lt;p&gt;Most of the speed came from deciding what not to do on the hot path: no database on reads, no layout work while scrolling, nothing optional before the first paint. Most of the reliability came from writing things down before acting on them: content into a cache that's cleared on purpose, events into an outbox before they leave.&lt;/p&gt;

&lt;p&gt;If I started again I'd keep the same order: settle a few visual rules, make reading fast, then add the fun parts where they don't get in the way of either.&lt;/p&gt;

&lt;p&gt;You can browse the &lt;a href="https://abhashchakraborty.tech/projects" rel="noopener noreferrer"&gt;projects&lt;/a&gt; or open the &lt;a href="https://abhashchakraborty.tech/terminal" rel="noopener noreferrer"&gt;terminal&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://abhashchakraborty.tech/writing/building-a-portfolio-that-feels-fast" rel="noopener noreferrer"&gt;abhashchakraborty.tech&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>webdev</category>
      <category>nextjs</category>
      <category>performance</category>
      <category>architecture</category>
    </item>
  </channel>
</rss>
