<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: G.Santhosh Reddy</title>
    <description>The latest articles on DEV Community by G.Santhosh Reddy (@gsanthosh_reddy_4a67d556).</description>
    <link>https://dev.to/gsanthosh_reddy_4a67d556</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4156114%2Fe650678d-0668-425c-96b2-69f9d170b4b3.jpg</url>
      <title>DEV Community: G.Santhosh Reddy</title>
      <link>https://dev.to/gsanthosh_reddy_4a67d556</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/gsanthosh_reddy_4a67d556"/>
    <language>en</language>
    <item>
      <title>An uncalibrated classifier is worse than no classifier</title>
      <dc:creator>G.Santhosh Reddy</dc:creator>
      <pubDate>Fri, 02 Oct 2026 02:35:22 +0000</pubDate>
      <link>https://dev.to/gsanthosh_reddy_4a67d556/an-uncalibrated-classifier-is-worse-than-no-classifier-4i2</link>
      <guid>https://dev.to/gsanthosh_reddy_4a67d556/an-uncalibrated-classifier-is-worse-than-no-classifier-4i2</guid>
      <description>&lt;p&gt;&lt;em&gt;What I learned building a website that rewrites itself for whoever is reading it, on ₹0 of infrastructure.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Project Blog : &lt;a href="https://santhosh-reddy.vercel.app/en/blog/8" rel="noopener noreferrer"&gt;https://santhosh-reddy.vercel.app/en/blog/8&lt;/a&gt;&lt;br&gt;
Project Breakdown : &lt;a href="https://santhosh-reddy.vercel.app/en/project/8" rel="noopener noreferrer"&gt;https://santhosh-reddy.vercel.app/en/project/8&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;My portfolio gets very different readers. A recruiter clicking through from a LinkedIn job post wants to know&lt;br&gt;
whether I'm hireable for an internship. A developer arriving from a GitHub repo wants architecture and code.&lt;br&gt;
A static page serves both of them the same thing, so it's a printed brochure.&lt;/p&gt;

&lt;p&gt;So I built &lt;strong&gt;Evolve&lt;/strong&gt;: every slot on the page (hero title, intro line, project order) is a &lt;em&gt;gene&lt;/em&gt;, every visitor&lt;br&gt;
gets a &lt;em&gt;genome&lt;/em&gt;, and each audience has its own population of genomes that competes, retires losers and breeds&lt;br&gt;
winners. The thing that decides "who is this?" is &lt;strong&gt;OpenJev&lt;/strong&gt;, a free, Jev-style typed-decision API I host on&lt;br&gt;
top of &lt;a href="https://github.com/NandhaKishorM/laya" rel="noopener noreferrer"&gt;Laya&lt;/a&gt;, an open-source small model for structured decisions.&lt;/p&gt;

&lt;p&gt;It's live on &lt;a href="https://santhosh-reddy.vercel.app" rel="noopener noreferrer"&gt;santhosh-reddy.vercel.app&lt;/a&gt;, and the code is at&lt;br&gt;
&lt;a href="https://github.com/SanthoshReddy352/Evolve-x-OpenJev" rel="noopener noreferrer"&gt;github.com/SanthoshReddy352/Evolve-x-OpenJev&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The most useful result wasn't the bandit or the genetics. It was finding out that &lt;strong&gt;routing visitors with a&lt;br&gt;
confidently wrong classifier is worse than not segmenting at all&lt;/strong&gt;, and that fixing it was a calibration&lt;br&gt;
problem, not an accuracy problem.&lt;/p&gt;

&lt;h2&gt;
  
  
  How it works
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;visitor ─► page + 3.8 KB SDK ─► evolve-edge Worker ─► Durable Object (populations, posteriors, lineage)
                                      │ waitUntil — never blocks the visitor
                                      └─► openjev-gateway Worker ─► D1 cache ─► tunnel ─► Laya student (Oracle VM)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Assignment.&lt;/strong&gt; Each genome keeps a Beta posterior over a session reward in [0, 1]. Clicking a project,
opening GitHub, downloading the resume and reaching contact all add weight; a bounce subtracts.
Thompson sampling picks the genome for each visit, per audience segment. 10% of visits are held out on the
original page so lift can be measured honestly.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Evolution.&lt;/strong&gt; Every 30 minutes, each segment estimates P(best) for its genomes, retires ones that have had
a fair hearing (≥ 20 sessions) and are clearly losing (P(best) &amp;lt; 2%), and refills the population with
children of the two best: uniform crossover per slot plus 10% mutation. Children inherit a weak prior
from their parents.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Who's visiting.&lt;/strong&gt; On a first visit the edge renders what it can observe into one sentence ("Referrer:
linkedin.com. Campaign parameters: utm_source=linkedin… Landing page: /. Device: desktop.") and asks
OpenJev two typed questions: &lt;em&gt;audience&lt;/em&gt; (recruiter / developer / student / founder) and &lt;em&gt;intent&lt;/em&gt;
(hire / explore projects / learn / contact). That call runs in &lt;code&gt;waitUntil&lt;/code&gt;: the visitor gets the global
population instantly, and their segment is ready for the next page view. &lt;strong&gt;Visitors never wait on the model.&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The finding
&lt;/h2&gt;

&lt;p&gt;Before training anything I wrote a simulator: 20 synthetic worlds × 5,000 visits, where each audience&lt;br&gt;
secretly prefers different genes. Conversion in the last 20% of visits:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;policy&lt;/th&gt;
&lt;th&gt;conversion&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;classic A/B test, 6 fixed variants&lt;/td&gt;
&lt;td&gt;18.8%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;one global evolving population&lt;/td&gt;
&lt;td&gt;27.3%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Evolve, segmented by Laya (85% accurate, calibrated)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;39.9%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;oracle (knows every visitor)&lt;/td&gt;
&lt;td&gt;58.0%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Segmenting doubles conversion over A/B testing. But then I swept the classifier's quality:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Laya routing&lt;/th&gt;
&lt;th&gt;conversion&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;32% accurate, confidently wrong (zero-shot)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;25.2%&lt;/strong&gt;, worse than no segmentation (27.3%)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;60%&lt;/td&gt;
&lt;td&gt;29.2%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;85%, uncalibrated&lt;/td&gt;
&lt;td&gt;38.4%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;85%, calibrated&lt;/td&gt;
&lt;td&gt;39.9%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;95%, calibrated&lt;/td&gt;
&lt;td&gt;42.4%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A bad classifier doesn't just fail to help. It splits your traffic into populations that each learn from the&lt;br&gt;
wrong people, so every segment converges more slowly &lt;em&gt;and&lt;/em&gt; toward the wrong page. And zero-shot Laya was&lt;br&gt;
exactly that classifier: on my 53-row hand eval it scored 0.32 on audience (random is 0.25), it labelled&lt;br&gt;
&lt;strong&gt;all 21&lt;/strong&gt; live synthetic contexts "developer", and it never abstained on genuinely ambiguous visitors.&lt;/p&gt;

&lt;p&gt;Calibration is what lets the system be wrong safely. If the model says "recruiter, 0.52", Evolve can fall back&lt;br&gt;
to the global population instead of polluting the recruiter population. That only works if 0.52 &lt;em&gt;means&lt;/em&gt; 0.52.&lt;/p&gt;

&lt;h2&gt;
  
  
  Fixing it for ₹0
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Training data with honest targets.&lt;/strong&gt; I didn't hand-label thousands of visits. I wrote a seeded generative&lt;br&gt;
model of portfolio visitors: sample a latent (audience, intent), then sample every observable the edge can see&lt;br&gt;
conditioned on it (referrer host from category tables, UTM parameters, landing page, language, device). Because&lt;br&gt;
the model is known, every row's gold label is the &lt;strong&gt;exact Bayes posterior&lt;/strong&gt; over audiences and intents, so Laya&lt;br&gt;
is trained toward honest uncertainty instead of hard labels. 6,000 rows; 4 epochs on Kaggle's free 2×T4; 14 minutes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Evaluated on referrers it never saw.&lt;/strong&gt; Twelve referrer hosts were held out of training entirely:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;unseen referrers, n = 300&lt;/th&gt;
&lt;th&gt;zero-shot&lt;/th&gt;
&lt;th&gt;fine-tuned&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;audience accuracy&lt;/td&gt;
&lt;td&gt;0.397&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.817&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;audience ECE (lower is better)&lt;/td&gt;
&lt;td&gt;0.23&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.071&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;intent accuracy&lt;/td&gt;
&lt;td&gt;0.583&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.963&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;abstains on ambiguous visitors&lt;/td&gt;
&lt;td&gt;0 / 7&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;3 / 7&lt;/strong&gt; (0 / 53 on clear ones)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The same LinkedIn job-post visit that zero-shot Laya called "developer (0.65)" now comes back&lt;br&gt;
&lt;strong&gt;recruiter (0.91), intent hire&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Distilled to fit a free VM.&lt;/strong&gt; Fine-tuned Laya needs ~1.6 GB of RAM and ~1.4 s per decision on my laptop&lt;br&gt;
CPU, and my laptop isn't a server. So I distilled it into a MiniLM-L6 student (22M parameters, one linear head&lt;br&gt;
per question, one fitted temperature per head so its confidence stays calibrated), exported it to int8 ONNX,&lt;br&gt;
and run it on Oracle's free 1 GB micro VM:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;unseen referrers&lt;/th&gt;
&lt;th&gt;laptop Laya&lt;/th&gt;
&lt;th&gt;Oracle student (int8)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;audience&lt;/td&gt;
&lt;td&gt;0.817&lt;/td&gt;
&lt;td&gt;0.823&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;intent&lt;/td&gt;
&lt;td&gt;0.963&lt;/td&gt;
&lt;td&gt;0.937&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;per decision / memory&lt;/td&gt;
&lt;td&gt;~1.4 s / ~1.6 GB&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~100 ms / 93 MB&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The VM sustains about 8 fresh decisions per second. Repeat visitor contexts never reach it: the gateway&lt;br&gt;
caches decisions in D1 (11 ms p50), so the real ceiling is Cloudflare's free 100k requests/day, which is roughly&lt;br&gt;
12–15k visits a day.&lt;/p&gt;

&lt;h2&gt;
  
  
  The ₹0 stack
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Cloudflare Workers + Durable Objects (SQLite) + D1&lt;/strong&gt;: the edge, the per-site population state, keys,
rate limits and the decision cache.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Kaggle&lt;/strong&gt; free GPUs for fine-tuning and distillation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Oracle Cloud&lt;/strong&gt; free micro VM for the student, exposed with a Cloudflare quick tunnel that re-registers
itself with the gateway on boot. The gateway health-checks every origin every 5 minutes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Vercel&lt;/strong&gt; for the portfolio itself.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;80 tests across six packages. No card on file anywhere.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd say honestly
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;The simulator numbers are from my own simulator, and the hand eval was written by the same person who wrote
the generative tables, which flatters the recipe. The held-out-referrer split is the number I trust.&lt;/li&gt;
&lt;li&gt;Real-traffic lift is &lt;strong&gt;not measured yet&lt;/strong&gt;. A student portfolio gets tens of visits a day; the holdout needs a
few hundred sessions before a lift number means anything. It's running, and I'll update this post when it lands.&lt;/li&gt;
&lt;li&gt;Genetic-algorithm site optimization and segment personalization both exist commercially. What I think is worth
sharing is the narrower lesson: &lt;strong&gt;if you route by a model's prediction, its calibration matters more than its
accuracy, and a typed-decision model that can say "I don't know" is what makes adaptive pages safe.&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;em&gt;Code: &lt;a href="https://github.com/SanthoshReddy352/Evolve-x-OpenJev" rel="noopener noreferrer"&gt;github.com/SanthoshReddy352/Evolve-x-OpenJev&lt;/a&gt;&lt;br&gt;
· Laya: &lt;a href="https://github.com/NandhaKishorM/laya" rel="noopener noreferrer"&gt;github.com/NandhaKishorM/laya&lt;/a&gt; (Apache-2.0)&lt;/em&gt;&lt;/p&gt;

&lt;h1&gt;
  
  
  machinelearning
&lt;/h1&gt;

&lt;h1&gt;
  
  
  webdev
&lt;/h1&gt;

&lt;h1&gt;
  
  
  opensource
&lt;/h1&gt;

&lt;h1&gt;
  
  
  cloudflare
&lt;/h1&gt;

</description>
      <category>frontend</category>
      <category>showdev</category>
      <category>webdev</category>
    </item>
  </channel>
</rss>
