What I Built
OneInbox is a daily email digest that sorts my inbox into four buckets:
- Action required: needs a reply, or has a deadline within 7 days
- Worth a look: useful, but nothing to do
- FYI: receipts, confirmations, security notices
- Noise: promos, newsletters, social
The idea came from a complaint in a hackathon Discord about how tiring it is to keep checking inboxes. I have the same problem, so I built it for my own inbox.
The core is a LoRA fine-tune of Qwen3.5-4B (Apache 2.0). Each morning it reads the last 24 hours of mail and produces a page I can read on my phone in under a minute. Every item has buttons to mark it Accurate or Wrong, silence the sender, or always show it.
Demo
Live: https://oneinbox-web.onrender.com/digest?token=hacktober2026
It runs in demo mode on 25 synthetic emails, classified by the fine-tuned model. The first load can take about 30 seconds, because Render's free tier sleeps when idle. Due dates are in September because the sample emails are synthetic, and timestamps are UTC.
Code
- GitHub: https://github.com/DkHrR/oneinbox (MIT)
- Model: https://huggingface.co/dKhRr/oneinbox-qwen3.5-4b-lora
How I Built It
Data. I didn't want to put a real inbox into a hackathon project, so I generated 600 synthetic emails (150 per label) with Qwen3.6-35B-A3B on Tinker. Sender names match the type of mail, deadlines are computed from each email's received date, and a second pass checks every email. I audited the set: no deadline mismatches, and very few subjects shouting "URGENT". I split it into 540 train and 60 validation emails.
Training. A Tinker LoRA fine-tune of Qwen3.5-4B on the full JSON answer (label, summary, deadline). Training cost about $0.19 to $0.31.
Results on the 60 held-out synthetic emails:
| Variant | Labels correct | Macro-F1 | Deadline match | Cost / 1k emails |
|---|---|---|---|---|
| Base 4B, zero-shot | 52/60 | 0.86 | 23/28 | $0.157 |
| Base 4B, 3-shot | 56/60 | 0.93 | 25/28 | $0.278 |
| Fine-tuned 4B | 58/60 | 0.97 | 25/28 | $0.141 |
| Base 35B-A3B, zero-shot | 51/60 | 0.85 | 25/28 | $0.229 |
What this shows, and what it doesn't:
- The fine-tune beats 3-shot prompting by two emails. That is a small gain, not a landslide. The bigger point is that it needs no prompt examples, so it costs about half as much per 1,000 emails.
- It is not faster: p50 latency is about 2.85s, against 2.55s for 3-shot.
- Deadline accuracy is tied with 3-shot.
- The test set is synthetic, and its gold labels came from the generator, not from a human. I reviewed the fine-tuned model's mistakes: some were real misses (an off-by-one bank deadline, a missed hospital deadline) and some were debatable labels. Treat these numbers as a sanity check, not a real-world accuracy claim.
- Pricing uses Tinker's published base-model rates, because I found no separate rate for fine-tuned sampling.
Memory (Backboard). Sender preferences like "always show" and "silence" are stored as Backboard memories and applied by simple code rules each run. I use Backboard only for memory. Accurate/Wrong feedback is logged there but does not change filtering yet. It is saved for future retraining.
Hosting (Render + GitHub Actions). The digest page is a small Flask service on Render's free tier. The daily run is a GitHub Actions workflow at 07:00 IST, and it posts the finished digest to the service. Render cron needs a paid plan, so I moved the schedule to GitHub Actions to keep the project at $0 to run. The web service keeps nothing on disk, so the digest is saved in Backboard memory and survives restarts.
What I'd Fix Next
No outside person has tried OneInbox yet, so I had an AI reviewer play a first-time user and tell me what felt wrong. This is not user feedback, but its points matched gaps I could already see:
- No way to open the original email. The digest tells you what needs action, but you still have to go find the email yourself. In IMAP mode I can link each item to a mail search for that sender and subject. The demo can't, because its emails are synthetic.
- Accurate / Wrong buttons don't change anything yet. They only log feedback, as noted above. Next step is to use that log to adjust sender rules, or to feed it into a retraining run. Until then they are a data-collection tool, not a control.
- Sender names can hide the real address. The full address is only in a tooltip, which doesn't work on a phone. Showing the domain in small text under the name matters for security alerts, where spoofed senders are the whole risk.
- Trusting an AI with deadlines. A wrong date on a bill is worse than no digest. The fine-tuned model matched 25/28 deadlines in my test, not 28/28, so I want the original email one tap away, which brings back the first point.
Why Open Innovation Matters
An inbox is some of the most private text a person has, so I didn't want a closed model deciding what matters in it. With an open-weight 4B model:
- The weights, my LoRA adapter, and the code are public, and the base model is Apache 2.0. Anyone can inspect, fork, or retrain them.
- A small model is cheap enough to run daily without a subscription.
- My feedback buttons are designed to feed a future fine-tune on someone's own corrections.
Prize Categories
- Tinker: generated the dataset, fine-tuned, and evaluated everything on Tinker.
-
Render: the digest page runs as a Render free-tier web service.
render.yamlis in the repo. - Backboard: sender preferences, feedback logs, and digest persistence use Backboard memory.


Top comments (0)