This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
SideQuest turns "we should study together" into one plan. One person creates a study circle and shares a link. Everyone adds what works for them, privately: how much they can spend, when they're free, where they like to study, and whether they need somewhere quiet or step-free. The group sees the options and the votes, never each other's numbers. The organiser confirms, and anyone in the circle can save a calendar invite or play a short spoken invitation.
The part I cared about most: nobody should have to post their budget, or the fact that they can't manage stairs, to the whole group.
Demo
Live app: https://sidequest-lzrz.onrender.com
Explore a sample circle works without an account. (It's on Render's free tier, so the first load after a quiet spell can take a minute.) To try the real thing, create a circle with Live study spaces, open Your preferences, and type the way you'd text: kal shaam 5 se 8 free hu, 200 se zyada nahi, library ya cafe.
Code
SideQuest
Study circles for school and college students, with private preferences and one shared plan.
SideQuest helps classmates turn "we should revise together" into a time and place. Students share availability, study-space budgets and requirements privately, vote on a shortlist and confirm a session with a calendar invitation.
Live app | Architecture | Testing | Sponsor evidence
How a circle works
flowchart LR
A[Create a study circle] --> B[Invite classmates]
B --> C[Save private preferences]
C --> D[Find study spaces]
D --> E[Vote on a shortlist]
E --> F[Organiser confirms]
F --> G[Download calendar invite]
- Name the session, such as Data structures revision, and choose a city/date.
- Share the invite link. Each student joins with a separate participant session.
- Save a budget in INR, availability, preferred space types, quietness and step-free requirements.
- The organiser finds options once everyone has saved preferences.
- Students vote; the organiser acknowledges unresolved details…
How I Built It
A 4B model that reads how students actually text
Nobody fills in forms in a group chat. They write "kal shaam 5 se 8 free hu, 200 se zyada nahi, library ya cafe, lift chahiye". SideQuest turns that, or a voice note, into a draft of the six private fields. The student checks the draft before anything is saved, and the planner's hard rules never depend on the model.
I started with Qwen3.5-4B as it comes. On 50 hand-written test messages (English, Hinglish, typos, a couple of "ignore previous instructions") it got 34 completely right. The misses weren't exotic. Half of them read a student's "4 to 6" as four in the morning, even though the prompt spells out the rule. It turned "stairs are no problem" into needs step-free access. When a message didn't mention money, it set the budget to ₹0 instead of keeping the student's saved value. And given "kal shaam 5 se 8 free hu, 200 se zyada nahi", it ticked all five kinds of study space, though the message never says where.
So I fine-tuned it on Tinker: 800 synthetic messages in English and Hinglish, LoRA rank 16, 75 steps.
| Model (same prompt) | Hand-written, 50 | Unseen generated, 100 |
|---|---|---|
| Qwen3.5-4B, as it comes | 34 | 58 |
| Qwen3.5-4B, three examples in the prompt | 35 | 65 |
| Qwen3.6-27B, as it comes | 50 | 91 |
| Qwen3.5-4B, fine-tuned on Tinker | 50 | 99 |
The tuned 4B model matches a model nearly seven times its size on the hand-written set, and it's the model running in the app now, served through Tinker's OpenAI-compatible endpoint. One file holds the prompt, and both the trainer and the server read it, so production sends exactly the tokens the model was trained on. I compared the Python and JavaScript renderings byte for byte before shipping.
The limits, plainly: every test message is synthetic. I wrote the 50 hand-written ones separately from the training generator, but nobody's real messages are in there. 50/50 is not "never wrong", which is why the model only ever drafts.
Voice in, voice out, with ElevenLabs
If typing is too much, upload a voice note. ElevenLabs Scribe transcribes it, the same tuned model drafts it, and the student reviews it. When the plan is confirmed, ElevenLabs reads it back as a short invitation. Both need an explicit consent tick. The demo film's narration is ElevenLabs too.
"Is it quiet at 5 PM?" TabPFN on Google popular times
SerpApi finds real study spots in the city: libraries, reading rooms, cafes, coworking spaces, parks. But "I need somewhere quiet" is hard to check. Google shows how busy a place usually gets for only 39 of the 175 spots I found, and for just 3 of the 30 libraries.
So I ran TabPFN v2, with its open weights on my laptop's CPU, on those 39 places' hourly history. It uses only what a search result already gives: place type, rating, review count, location, opening hours, weekday and hour. I scored it only on venues it had never seen (5-fold cross-validation, grouped by venue):
| On unseen venues | Avg. error (0–100) | Quiet hours found | Quiet calls right |
|---|---|---|---|
| Average by place type and hour | 14.8 | 42% | 60% |
| Gradient boosting | 14.4 | 52% | 62% |
| TabPFN v2 | 14.5 | 63% | 59% |
On average error it's a tie. Where TabPFN pulled ahead is the part that matters to a group that needs quiet: it found more of the genuinely quiet hours, and its "quiet" calls were right about as often as the others'. It's a modest edge from 39 venues, not a solved problem, and I'm treating it that way.
Its forecasts now sit beside each option ("Usually quiet around 5 PM · TabPFN forecast from similar places"). When anyone in the circle needs quiet, quieter places rank first among equally good matches. Busyness is crowding, not noise, so a forecast only orders options and never removes one. Coworking spaces get no forecast, because only one had history to learn from.
The rest
- Mastra runs both workflows: discovery (search, then the constraint planner) and interpretation (the tuned model).
- The planner is ordinary code: shared free time, budgets, quiet and step-free requirements. Unknown venue facts stay visibly unknown.
- Render hosts the React app and Express API together. MongoDB Atlas stores circles and the busyness forecasts, and a confirmed plan survived a redeploy.
- Every person's credential is hashed. Peers can't fetch each other's exact preferences.
- GitHub Actions type-checks, builds and runs the tests on every push to main.
Why Open Innovation Matters
I could fix the model instead of hoping. A closed API would have given me the 34/50 model and a prompt to fiddle with. Open weights let me teach a small model the conventions of how students in India write ("shaam 5", "200 se zyada nahi", "lift chahiye"), measure the change on held-out messages, and ship that exact adapter.
Small is enough once it's specific. A 4B model tuned for one job matched a 27B general model on this task.
TabPFN runs on a laptop. The busyness model is TabPFN v2's open weights on my CPU, with no account and no service of its own. Each cross-validation fold took about a minute.
Every open piece is measured. The model before and after fine-tuning, and the forecasts against simple baselines. The results files are in the repo, misses included.
Prize Categories
- Best Use of Tinker: fine-tuned Qwen3.5-4B, 34/50 → 50/50 on hand-written held-out messages, serving live.
- Best Use of TabPFN: busyness forecasts for venues with no Google history, evaluated on unseen venues.
- Best Use of ElevenLabs: Scribe transcription into the open model, spoken invitations, film narration.
- Best Use of Render: the app, API and model-backed workflows run on Render.
- Best Use of Mastra: workflows orchestrating the open model and the search tool.
- Best Use of MongoDB Atlas: the database behind an app built on an open-weight model.
- Best Use of SerpApi: live study-space search.
Top comments (0)