What I learned building a website that rewrites itself for whoever is reading it, on ₹0 of infrastructure.
Project Blog : https://santhosh-reddy.vercel.app/en/blog/8
Project Breakdown : https://santhosh-reddy.vercel.app/en/project/8
My portfolio gets very different readers. A recruiter clicking through from a LinkedIn job post wants to know
whether I'm hireable for an internship. A developer arriving from a GitHub repo wants architecture and code.
A static page serves both of them the same thing, so it's a printed brochure.
So I built Evolve: every slot on the page (hero title, intro line, project order) is a gene, every visitor
gets a genome, and each audience has its own population of genomes that competes, retires losers and breeds
winners. The thing that decides "who is this?" is OpenJev, a free, Jev-style typed-decision API I host on
top of Laya, an open-source small model for structured decisions.
It's live on santhosh-reddy.vercel.app, and the code is at
github.com/SanthoshReddy352/Evolve-x-OpenJev.
The most useful result wasn't the bandit or the genetics. It was finding out that routing visitors with a
confidently wrong classifier is worse than not segmenting at all, and that fixing it was a calibration
problem, not an accuracy problem.
How it works
visitor ─► page + 3.8 KB SDK ─► evolve-edge Worker ─► Durable Object (populations, posteriors, lineage)
│ waitUntil — never blocks the visitor
└─► openjev-gateway Worker ─► D1 cache ─► tunnel ─► Laya student (Oracle VM)
- Assignment. Each genome keeps a Beta posterior over a session reward in [0, 1]. Clicking a project, opening GitHub, downloading the resume and reaching contact all add weight; a bounce subtracts. Thompson sampling picks the genome for each visit, per audience segment. 10% of visits are held out on the original page so lift can be measured honestly.
- Evolution. Every 30 minutes, each segment estimates P(best) for its genomes, retires ones that have had a fair hearing (≥ 20 sessions) and are clearly losing (P(best) < 2%), and refills the population with children of the two best: uniform crossover per slot plus 10% mutation. Children inherit a weak prior from their parents.
-
Who's visiting. On a first visit the edge renders what it can observe into one sentence ("Referrer:
linkedin.com. Campaign parameters: utm_source=linkedin… Landing page: /. Device: desktop.") and asks
OpenJev two typed questions: audience (recruiter / developer / student / founder) and intent
(hire / explore projects / learn / contact). That call runs in
waitUntil: the visitor gets the global population instantly, and their segment is ready for the next page view. Visitors never wait on the model.
The finding
Before training anything I wrote a simulator: 20 synthetic worlds × 5,000 visits, where each audience
secretly prefers different genes. Conversion in the last 20% of visits:
| policy | conversion |
|---|---|
| classic A/B test, 6 fixed variants | 18.8% |
| one global evolving population | 27.3% |
| Evolve, segmented by Laya (85% accurate, calibrated) | 39.9% |
| oracle (knows every visitor) | 58.0% |
Segmenting doubles conversion over A/B testing. But then I swept the classifier's quality:
| Laya routing | conversion |
|---|---|
| 32% accurate, confidently wrong (zero-shot) | 25.2%, worse than no segmentation (27.3%) |
| 60% | 29.2% |
| 85%, uncalibrated | 38.4% |
| 85%, calibrated | 39.9% |
| 95%, calibrated | 42.4% |
A bad classifier doesn't just fail to help. It splits your traffic into populations that each learn from the
wrong people, so every segment converges more slowly and toward the wrong page. And zero-shot Laya was
exactly that classifier: on my 53-row hand eval it scored 0.32 on audience (random is 0.25), it labelled
all 21 live synthetic contexts "developer", and it never abstained on genuinely ambiguous visitors.
Calibration is what lets the system be wrong safely. If the model says "recruiter, 0.52", Evolve can fall back
to the global population instead of polluting the recruiter population. That only works if 0.52 means 0.52.
Fixing it for ₹0
Training data with honest targets. I didn't hand-label thousands of visits. I wrote a seeded generative
model of portfolio visitors: sample a latent (audience, intent), then sample every observable the edge can see
conditioned on it (referrer host from category tables, UTM parameters, landing page, language, device). Because
the model is known, every row's gold label is the exact Bayes posterior over audiences and intents, so Laya
is trained toward honest uncertainty instead of hard labels. 6,000 rows; 4 epochs on Kaggle's free 2×T4; 14 minutes.
Evaluated on referrers it never saw. Twelve referrer hosts were held out of training entirely:
| unseen referrers, n = 300 | zero-shot | fine-tuned |
|---|---|---|
| audience accuracy | 0.397 | 0.817 |
| audience ECE (lower is better) | 0.23 | 0.071 |
| intent accuracy | 0.583 | 0.963 |
| abstains on ambiguous visitors | 0 / 7 | 3 / 7 (0 / 53 on clear ones) |
The same LinkedIn job-post visit that zero-shot Laya called "developer (0.65)" now comes back
recruiter (0.91), intent hire.
Distilled to fit a free VM. Fine-tuned Laya needs ~1.6 GB of RAM and ~1.4 s per decision on my laptop
CPU, and my laptop isn't a server. So I distilled it into a MiniLM-L6 student (22M parameters, one linear head
per question, one fitted temperature per head so its confidence stays calibrated), exported it to int8 ONNX,
and run it on Oracle's free 1 GB micro VM:
| unseen referrers | laptop Laya | Oracle student (int8) |
|---|---|---|
| audience | 0.817 | 0.823 |
| intent | 0.963 | 0.937 |
| per decision / memory | ~1.4 s / ~1.6 GB | ~100 ms / 93 MB |
The VM sustains about 8 fresh decisions per second. Repeat visitor contexts never reach it: the gateway
caches decisions in D1 (11 ms p50), so the real ceiling is Cloudflare's free 100k requests/day, which is roughly
12–15k visits a day.
The ₹0 stack
- Cloudflare Workers + Durable Objects (SQLite) + D1: the edge, the per-site population state, keys, rate limits and the decision cache.
- Kaggle free GPUs for fine-tuning and distillation.
- Oracle Cloud free micro VM for the student, exposed with a Cloudflare quick tunnel that re-registers itself with the gateway on boot. The gateway health-checks every origin every 5 minutes.
- Vercel for the portfolio itself.
80 tests across six packages. No card on file anywhere.
What I'd say honestly
- The simulator numbers are from my own simulator, and the hand eval was written by the same person who wrote the generative tables, which flatters the recipe. The held-out-referrer split is the number I trust.
- Real-traffic lift is not measured yet. A student portfolio gets tens of visits a day; the holdout needs a few hundred sessions before a lift number means anything. It's running, and I'll update this post when it lands.
- Genetic-algorithm site optimization and segment personalization both exist commercially. What I think is worth sharing is the narrower lesson: if you route by a model's prediction, its calibration matters more than its accuracy, and a typed-decision model that can say "I don't know" is what makes adaptive pages safe.
Code: github.com/SanthoshReddy352/Evolve-x-OpenJev
· Laya: github.com/NandhaKishorM/laya (Apache-2.0)
Top comments (0)