DEV Community

Cover image for Finally, an AI that won't oversell my friend's resume. It shows the proof instead.
Aniket Dhingra
Aniket Dhingra

Posted on

Finally, an AI that won't oversell my friend's resume. It shows the proof instead.

Hacktoberfest Weekend Challenge: Build for a Friend Submission 🤝

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

What I Built

My friend Pranshu is job hunting. Like everyone, he's staring at a job post with twenty requirements and a two-page resume, wondering which line actually proves what. The AI tools he could use either rewrite everything with confident claims that aren't his, or hand back a "92% match" that explains nothing.

So I built ProofPrep, a resume coach with one rule: no claim without a quote.

  1. Add details: Upload a resume, paste the job post.
  2. Check the evidence: Each requirement becomes a card with the exact resume line behind it, highlighted. No line? The card says "no evidence found" (not "you can't do this", because a missing line is not a missing skill). He accepts, rejects, or adds context in his own words.
  3. *Improve the resume:. Rewrites use only facts he approved, new numbers or tools get flagged "Check this detail", and new bullets go where they fit, not into the Volunteering section. Export a clean PDF.
  4. Practice the interview: Questions come only from approved facts. Pick one bullet and get grilled on it from three angles, like an interviewer who actually read your resume.

No score. No hiring odds. Nobody gets a number to be scared of.

What Pranshu said:

Being in control of every edit to my resume, and getting interview prep built around my specific bullet points, gives me the confidence to stand behind my own work.

His feedback also reshaped the app: longer job posts, a Full Stack posting that exposed a bad evidence match (now fixed with per-requirement search), step-by-step guidance, and a spelling slip I'd rather not discuss ("practise").

Demo

Code

ProofPrep

Help a friend get shortlisted and interview-ready for one specific job, using only what their resume can actually prove. Everything runs on your own laptop through Ollama; no API keys, no cloud, no cost.

ProofPrep reads a resume and a job description, links each job requirement to a verbatim quote from the resume (or says plainly that it found none), lets the candidate accept, reject or add context to every suggestion, proposes resume edits grounded only in approved facts, exports the approved version as a PDF, and runs adaptive mock interviews built from the same approved facts.

It never produces a score, a match percentage or a hiring probability. A missing resume line is treated as "not shown", never as "can't do it".

Run it

Requirements: Node 22+, Ollama, ~6 GB of disk for two small models.

ollama pull qwen3:4b-instruct-2507-q4_K_M   # evidence
…
Enter fullscreen mode Exit fullscreen mode

How I Built It

Two small open models through Ollama, each doing the job it's actually good at:

  1. Qwen3 4B finds evidence and writes grounded rewrites.
  2. Gemma 3 4B writes interview questions, feedback and follow-ups.

I didn't plan two models. I started with Gemma for everything, then measured both on the same real resume and job post:

Gemma happily treated unrelated tech lines as proof of "collaboration" and made up four quotes. My app rejected all four. So Gemma got the job where a draft is fine (interview practice), and Qwen got the job where accuracy matters. (One resume, one run, pipeline tuned on it. A measurement, not a benchmark.)

The guardrails don't rely on the model behaving:

  • JSON schemas stop the model skipping a requirement or citing a line it wasn't shown.
  • Every "supported" verdict needs a verbatim quote, and the app checks it exists. A fake quote becomes "could not assess". A half-fake one is cut back to the real part.
  • Rewrites can't add facts without a flag.
  • A Mastra workflow runs the four human decision points with real suspend/resume. Close the tab, restart the server, and it picks up where it stopped.

Plus React + Vite + TypeScript, node:sqlite, pdf.js, and browser print-to-PDF. 26 mocked tests cover the software's guarantees; they say nothing about model accuracy, which is what the eval script is for.

Why Does Open Innovation Matter?

  • A resume is personal, and it never leaves the laptop. Extraction, both models, storage and export all run locally, no external fonts or CDNs. After the one-time model download, it works with the Wi-Fi off. A closed API can't give me "nothing is sent anywhere".
  • Free inference meant I could measure. I re-ran the pipeline over and over and compared two models on identical input. A metered API would have made me ration that.
  • I could swap models per task. The best result came from not picking one model. Both are one ollama pull away, and env variables swap either.
  • Where closed would win: raw judgment. A 4B model on a laptop is slower and less sharp than a frontier API, and a long job post takes a few minutes. So I built around it with schemas, quote checks and a human in the loop. (No fine-tuning here. The gains came from model choice and pipeline design.)

My Agent Session

Claude Code Session
You

This session is being continued from a previous conversation that ran out of context. The summary below covers the earlier portion of the conversation.

Summary:

1. Primary Request and Intent

Project: ProofPrep, Pranshu's Hacktoberfest 2026 DEV challenge entry ("Build for a Friend"; deadline Oct 5, 2026 12:29 PM IST). It is a local-first app that:
- links job-description requirements to verbatim resume quotes;
- lets the candidate accept, reject, or add context to each match;
- proposes grounded resume edits;
- exports a PDF;
- runs adaptive mock interviews.

Models: Qwen3 4B (qwen3:4b-instruct-2507-q4_K_M) for evidence analysis and tailoring, Gemma 3 4B for interview questions and feedback.

Constraints:
- ₹0 budget; Ollama only.
- Never invent qualifications or metrics.
- No ATS score or hiring probability.
- A missing resume line does not mean the candidate lacks the skill.
- Distinguish passing mocked tests from real model accuracy.
- Make edits and run checks directly instead of asking the user to copy patches.

Requests handled this session (all done, committed to the device):
- Friend-feedback round: JD limit, practise→practice, Full Stack evidence cross-check and pipeline fixes, simple UI with help and guides, disclaimers removed, premium UI with animations.
- The resume text inserted at the end of the Volunteering section (new bullets), improved PDF export template, and bullet-drill question quality.
- The model deciding where a new bullet goes, or which line to replace.
- A fix for random bold/large text in the PDF export.

CURRENT request (latest user message, implemented but not yet verified or committed): "now create me a very nice asthetic premium looking start page, like the home screen, which visually draws the judges. and also if you could add a dark and light mode toggle on the top left. and remove the private on this computer part next to the ready status, just keep the ready status. i dotn want anything which shows on the project that this is running locally."

2. Key Technical Concepts

  • Ollama /api/chat with JSON-schema format. num_ctx is now 8192; per-call timeout 180s, total 900s. Pipeline evidence-review-v6.
  • Pipeline:
    • Chunked JD requirement extraction (verbatim quotes, up to 40).
    • Per-requirement resume search (schema selectionSchemaFor([r.id], sourceIds)).
    • Per-requirement review (reviewSchemaFor).
    • verifiedQuote: contiguous match with 2+ words, or one 5+ letter term when the line has ≤14 words. If the full quote fails, salvage the longest contiguous fragment of 2+ words and warn.
  • LIMITS = {resume: 12000, jobDescription: 12000} (server and src/lib/api.ts). Server body limits 160000.
  • node:sqlite (DB migration UPDATE applications SET stage='practice' WHERE stage='practise'); Mastra workflow with suspend/resume (zod resumeSchema strips unknown keys, so new fields must be added to the schema).
  • Tailoring:
    • ungroundedTerms flags new numbers/tools/names ("Check this detail", unchecked by default).
    • No-op rewrites are dropped.
    • Shared applyEdits(resume, accepted) is used by both the /versions route and the workflow's approve-edits step.
  • Placement (choosePlacement): the model picks a bullet ID from an enum with insert_after or replace. Replace is downgraded if word overlap < 0.35. Fallback fallbackPlacement uses word overlap with bullet text plus entry name. bulletOptions(resume) lists content bullets (excluding "Built using" lines) with their entry and section.
  • Extraction (extract.mjs): pdf.js lines, blank lines stripped for PDFs, joinWrapped merges wrapped bullet lines (lowercase start, mid-phrase endings, digit continuation after an unfinished bullet).
  • Interview focus mode: three separate per-angle calls with a single-question schema {question, probes}. Each call sees only the focus fact and previous questions ([...asked] copy), with overlap and repeat checks and a template fallback.
  • Print export: src/lib/resumeHtml.ts (resumeToHtml, RESUME_CSS, internal joinWrapped) renders name header, contact lines, section headings with rules, entry titles with right-aligned dates, bullets, skills rows, "Built using" tech lines, summary paragraphs, and split " • " runs (volunteering).
  • Frontend: React 19 + Vite 7 + TS, plain CSS with tokens, hash routing, animations that respect prefers-reduced-motion, no external fonts/CDNs. Theme via data-theme on <html>.
  • Operational:
    • The commit tool can re-send cached files, so stage under a fresh /mnt/user-data/outputs/<dir>/ each time and verify with device_list_dir.
    • Device path base is C:\Users\[user]\OneDrive\Documents\proofprep (Windows, backslashes).
    • There is no device_bash; deletions are impossible, so the stale src\screens\Practise.tsx stays on the device and the user was told to delete it.
    • Playwright uses an explicit executablePath.
    • /tmp/e2e/restart.sh restarts fake-ollama (port 11997) and the server (port 3001).

3. Files and Code Sections (all under [REDACTED]/proofprep, mirrored to the device)

Committed to device (r2–r5): server/analysis.mjs, analysis.test.mjs, index.mjs, db.mjs, tailor.mjs, interview.mjs, workflow.mjs, extract.mjs, src/App.tsx (older version), src/styles.css (older), src/components/Guide.tsx (older), src/lib/api.ts, src/lib/resumeHtml.ts, src/screens/{Applications,Inputs,ApplicationShell,Review,Tailor,Practice}.tsx, dist/index.html + assets (last committed: index-AELXCRA8.js, index-DI4i5nkp.css), README.md, HANDOFF.md. The last full device state corresponds to r5; round-4 changes below are NOT yet on the device.

Key code in server/tailor.mjs:
- bulletOptions(resume) returns [{id, text, entry, section}].
- choosePlacement({callModel, resume, text, bullets, signal}) returns {action, lineId, reason, by:'model'|'rules'}.
- applyEdits handles rewrites (keeping the bullet marker), placed items (insert after the chosen bullet, past wrapped continuation lines, or replace), and legacy additions with section/anchor via insertIntoSection.
- The /tailor route adds placement/placements to new-bullet proposals.

Current-round changes (uncommitted, built OK):
- src/App.tsx (rewritten):
- Hash routing: #/ → <Landing/>, #/applications → <Applications/>, #/new → <Applications newApp/>, #/app/:id/:tab → ApplicationShell.
- useTheme() stores pp-theme in localStorage (try/catch) and sets documentElement.dataset.theme plus the theme-color meta.
- Topbar order: theme-toggle button (sun/moon icon, left), brand link #/, spacer, status pill (Ready / Engine offline / Server unreachable / Checking…), an "Applications" ghost link when not on landing, and a "How it works" button.
- <HelpDrawer open onClose />, with no health prop.
- The topbar gets class on-landing (transparent) on the landing page.
- src/screens/Landing.tsx (new). Sections:
- Hero with pill "Evidence-first interview prep", headline "Walk into the interview knowing what your resume can prove." (gradient text), lead copy, CTAs "Get started" (#/new) and "See how it works" (scrolls to #how), and trust list.
- A floating product-preview stack (pv-main evidence card, pv-q interview question card, "Quote verified" and "Step 2 of 4" badges).
- 4-step cards, a 6-feature grid, a "paper" resume export showcase, a gradient CTA band, and a footer.
- useReveal() IntersectionObserver adds .in to .reveal elements.
- src/styles.css:
- Added the :root[data-theme="dark"] token set and new light tokens (--topbar --field --field-focus --mark --glow-a/b --info-line --info-ink --warn-line --warn-ink --bad-line --bad-ink --diff-add --diff-del --card-edge, color-scheme).
- Replaced hard-coded colors with tokens.
- Added .df/.df-add/.df-del, a select style and .theme-toggle.
- Appended the landing CSS block (.landing, .lp-bg orbs and grid, .lp-hero, .lp-copy h1, .grad-text, .pv-*, .lp-steps/.lp-step, .lp-grid/.lp-feature, .lp-showcase, .paper*, .lp-cta, .lp-foot, responsive and reduced-motion rules, .reveal).
- src/components/Guide.tsx:
- Icon union extended with sun/moon/lock/pdf/save.
- HelpDrawer signature is now ({ open, onClose }: { open: boolean; onClose: () => void; health?: unknown }).
- FAQ item "Is my information shared?" replaces "Where is my data?".
- The "Under the hood" section is removed.
- Practice step text: "Sessions are saved automatically."
- src/screens/Applications.tsx:
- The big hero is replaced by a compact page-head (title, "Home" link, "New application" / "Start your first application" button).
- Cancel and links use #/applications.
- The how-cards still show when the list is empty.
- ApplicationShell.tsx, Tailor.tsx, Practice.tsx: go('#/') → go('#/applications'), href="#/" → href="#/applications".
- ApplicationShell.tsx banner: "The analysis engine isn’t running yet. Start it and this button will unlock."
- Tailor.tsx: diff spans use className={w.k === 'same' ? '' :df df-${w.k}}; the inline select style was removed.
- index.html: added <meta name="theme-color" content="#f6f7f4" /> and an inline pre-paint script that sets data-theme from localStorage or prefers-color-scheme.
- Remaining intentional mentions: the Practice screen error banner "The interview model isn’t installed yet. Run ollama pull …" (error state only), and the Inputs text "The AI engine isn’t running yet; you can save now and continue later." (error state).

Test state: server/analysis.test.mjs has 26 passing tests as of the last run (the landing and theme changes touch only the frontend). tsc passed and npm run build succeeded after the landing work, producing dist/assets/index-C1Xdu2hq.css and index-vzBLN0xC.js.

4. Errors and fixes

  • Playwright browser not found: use executablePath: '/opt/pw-browsers/chromium-1194/chrome-linux/chrome'.
  • getByRole('button', {name:/How it works/}) failed because aria-label="Open help" overrides the name: use getByLabel('Open help').
  • Re-ran E2E when the DB wasn't fresh: restart.sh removes /tmp/e2e/test.sqlite*.
  • Fake model hard-coded to an old JD: rewrote its requirements, selection and review responses to be JD- and schema-aware.
  • Review test expected old behaviours: updated tests (the 3-word rule changed, call counts changed from 5 to 7, the extra-keys test changed).
  • previousQuestions mutation made the interview test see [3,3,3]: pass a copy [...asked].
  • New bullets still landed in Volunteering via the UI: the Mastra workflow had its own duplicate apply logic. Refactored to a shared applyEdits and extended the zod resume schema with section, anchor and then placement.
  • Placement was stripped by zod: added placement: z.object({action: z.enum([...]), lineId: z.string()}).nullable().optional() to the schema.
  • PDF export showed bold/large wrapped fragments: added joinWrapped and the sentence-end fragment rule to resumeHtml.ts.
  • First fallback test expectation (S8) was wrong; corrected to S7.
  • Edit-tool "file changed on disk" notices were from my own edits and were not a problem.

5. Problem Solving

Solved and committed: the pipeline limits and false negatives (mock-verified only), rename, UI redesign, placement, PDF template, drill questions, and wrapped-line merging at extraction. Not verifiable from the cloud: real Qwen/Gemma accuracy on the new search and placement, and the real drill output; the user must rerun with the real models and re-upload the PDF so extraction merges wrapped lines. Ongoing: verification and delivery of the landing page, dark mode and de-"local" copy (see sections 7 and 8).

6. All user messages

  • (Earlier, summarized) Handoff-continuation request; "what do u need me to do now"; eval outputs; Gemma prize screenshot and "wait a second what is qwen?…"; "lets go with option 2. i cant miss the prized category…"; storage complaint about gemma3n:e4b; pasted eval outputs; "whats the final decision then?…"; honest write-up angle plus UI/architecture question; stack choices (React+Vite+TS, Mastra yes, print-to-PDF); DB questions; then the six-item friend-feedback message (JD limit; practise→practice; cross-check Full Stack evidence; easier UX with instructions; remove local/model disclaimers; premium UI with animations).
  • "the UI is better now. there is somthing more to work on: 1. i have attached the screenshot. it basically added the suggested edits at the end of the resume in the which has the volunteering section which does not make any sense 2. the resume template which is used when the user clicks on "export pdf" is not good looking at all like can you make it better in layout wise and format wise. there is no formatting, no spacing. so improve that. 3. the bullet point specific practice, the first question is relatable to the correct bullet point but the other 2 questions are just repeating and also not relatable."
  • (Two screenshots) "if a new line is added as an edit, then i just adds at the last of the section, which is not good, can we do something where the model decides automatically where in the resume the line would look nice and appealing or replacing a line with the suggested edit, like think on its own?"
  • (Screenshot) "i love the foprmaatting of the resume but as u can see it is randomly highlighting words and increasing font sizes creating confusion between project titles and normal text. fix this"
  • (Screenshot of the home screen) "now create me a very nice asthetic premium looking start page, like the home screen, which visually draws the judges. and also if you could add a dark and light mode toggle on the top left. and remove the private on this computer part next to the ready status, just keep the ready status. i dotn want anything which shows on the project that this is running locally."

7. Pending Tasks

(All for the latest request.)
- Visual verification: restart /tmp/e2e/restart.sh, write a Playwright script that screenshots the landing page in light and dark, the #/applications and #/new pages, an inner page in dark mode (review/tailor), and the help drawer; check the toggle works and persists, and no console errors.
- Fix any layout or dark-mode contrast issues found (e.g., banners, chips, stepper, nextbar, drawer, textareas, diffs, .guide/.hero unchanged gradients).
- Grep the UI for leftover "local/computer/private/Ollama" text (the error-state mentions are the only intended ones) and confirm the status pill shows only "Ready".
- Run node --test server/analysis.test.mjs (expect 26 pass) and npx tsc --noEmit -p .; npm run build.
- Stage changed files under a fresh /mnt/user-data/outputs/r6/ and device_commit_files to C:\Users\[user]\OneDrive\Documents\proofprep\…:
- index.html
- src/App.tsx
- src/styles.css
- src/components/Guide.tsx
- src/screens/Landing.tsx (new)
- Applications.tsx
- ApplicationShell.tsx
- Tailor.tsx
- Practice.tsx
- dist/index.html
- dist/assets/index-vzBLN0xC.js and index-C1Xdu2hq.css (or whatever the final build names are)

Verify sizes with device_list_dir.
- Then tell the user, concisely, what changed (landing page, theme toggle top-left, status shows only "Ready", local mentions removed), what couldn't be verified, and to restart with node server\index.mjs. Optionally update the README layout/limits notes.

8. Current Work

Immediately before this summary request I had finished writing and building the landing page, theme toggle and copy changes: npx tsc --noEmit -p . passed and npm run build produced dist/assets/index-C1Xdu2hq.css (29.95 kB) and dist/assets/index-vzBLN0xC.js (284.47 kB). I had sent the user a progress message ("Building the new landing page, dark/light toggle (top left) and removing every 'runs on this computer' mention. Writing the page styles now, then I'll screenshot both themes."). The next step was screenshot verification in both themes.

9. Optional Next Step

Continue exactly where I left off: verify the new landing page and dark mode visually. Restart the stand-in stack (bash /tmp/e2e/restart.sh), then run a Playwright script (explicit chromium executablePath) that screenshots the landing page (#/) in light and dark, the applications and inner pages in dark, and the help drawer; fix any contrast or layout problems; run tests, tsc and build; stage under a fresh dir (r6); commit to the device; verify sizes; and report to the user. This follows the user's latest request: "now create me a very nice asthetic premium looking start page… add a dark and light mode toggle on the top left… remove the private on this computer part next to the ready status, just keep the ready status. i dotn want anything which shows on the project that this is running locally."

If you need specific details from before compaction (like exact code snippets, error messages, or content you generated), read the full transcript at: /root/.claude/projects/-home-claude/d18c25d3-a4ac-5dcb-ab3f-1fa309353fde.jsonl
Continue the conversation from where it left off without asking the user any further questions. Resume directly — do not acknowledge the summary, do not recap what was happening, do not preface with "I'll continue" or similar. Pick up the last task as if the break never happened.

Agent
Agent
Agent
Agent
Agent
Agent
Agent
Agent
Agent

Quick fix: the landing topbar is transparent, so content slides under it when scrolling. I'll make it solid once the page scrolls.

Agent
Agent
Agent
Agent

Header is fixed. Committing to the device.

Agent
Agent
Agent

The new start page, the dark/light toggle and the "Ready"-only status are all in your project folder. I checked them in headless Chromium against a stand-in for the AI engine, not your real models. There were no console errors, and the 26 server tests still pass.

  • Start page: the app now opens on a new home page with a headline, "Get started" and "See how it works" buttons, and floating preview cards of an evidence match and an interview question. Below that are a four-step section, six feature cards, a resume-export showcase and a closing call to action. It animates in and respects reduced-motion settings. "Applications" is now a separate page, and "Home" and the logo return to the start page.
  • Theme toggle: a sun/moon button sits at the top left, before the logo. It remembers your choice and defaults to your system setting. I checked that it persists after a reload. The whole app follows the theme, including the applications list, the help drawer and the diffs.
  • Status pill: it now just says "Ready", with no "private on this computer". I removed the other "runs locally" wording from the help text and banners too. The only mentions left are error messages when the AI engine is off or a model isn't installed, such as "run ollama pull…".

I screenshotted the start page and the applications list in both themes, plus the help drawer in dark. I did not screenshot the review, tailor and practice pages in dark mode. They use the same theme colours, but look through them once.

To see it, close the old server and run node server\index.mjs again, then hard-refresh. The old src\screens\Practise.tsx is still there and still needs deleting by hand.

You

yes i love the design, the colors and the project is alos working fine. now lets think about the przied categories what about that

Agent
Agent
Agent
Agent
Agent
Agent

I checked the official challenge page. It's 21:13 IST now and the deadline is 12:29 PM IST tomorrow, so you have about 15 hours.

Categories ProofPrep can enter

  1. Best Use of Gemma ($200). This is your main target. The rule is just "use Gemma, Google's open-weight model", and Gemma 3 writes your interview questions and feedback. It is judged on partner-technology use plus the standard criteria. Your honest angle fits well: you measured where Gemma is reliable and gave it only the jobs it handles well.
  2. Best Use of Mastra ($100). You use a real Mastra workflow with suspend and resume, so the human approves before the workflow continues. That matches "build a workflow". It costs almost nothing extra: one paragraph in the write-up plus a screenshot of the workflow.
  3. Overall ($250). Writing quality is weighted most heavily, then relevance, creativity and technical execution. The honest, measured write-up is your best chance here, and the new design helps with the screenshots.

Categories I'd skip

  • ElevenLabs (voice) and Backboard (compare open-weight models) sound close, but they need those partners' tools. You'd have to build something new tonight, which isn't worth the risk.
  • Sentry is possible if you show agent traces. It's a stretch, so only try it if you have time left.
  • Render, DigitalOcean, TabPFN, Tinker, Arduino and the rest don't fit.

What I'd do tonight

  • Pick one primary category in the post. The submission template asks for a "technical approach", so lead with Gemma. Mention Mastra in its own short section.
  • Put your own numbers in the write-up. Include your Gemma versus Qwen results, why you split the work, and what you tested on the real models versus the mocks. I only verified the pipeline against a stand-in model, so don't claim real-model accuracy I haven't seen.
  • Run one real end-to-end pass with your actual resume and a real job post. Use that for the demo video and screenshots. If the real output has a flaw, your write-up should say so.
  • Use the exact tags. They are #devchallenge, #weekendchallenge and #hf26challenge, with the template sections, code link and demo.
  • Submit by about 9 AM IST to leave a buffer.

I don't know whether one post can win several categories at once. The page doesn't say, so check the rules or ask the DEV team before counting on it.

Want me to draft the post (Gemma-first, with a Mastra section), or run through the demo and screenshot plan first?

Sources: Hacktoberfest Weekend Challenge: Build for a Friend

You

whats sentry and how long it will take to integrate, like tell me what would sentry do in this project?

Agent
Agent
Agent
Agent

Sentry is a hosted monitoring service. It started as error tracking and now has "Agent Tracing", which records each step of an AI run as a timeline: which model was called, with what input and output, how long it took, how many tokens it used, and where it failed. The prize asks you to "show your agent's work" with traces or screenshots.

What it would show in ProofPrep

For one analysis run you'd see a trace like this:

  • Extract requirements, in 3 chunks.
  • Search the resume for each requirement.
  • Review each requirement.
  • Suspend and wait for your approval.
  • Resume, tailor, place bullets.

Each step would show its Qwen or Gemma call and its duration, so you could point to the slow step. Failures would show up too: a quote that failed verification and fell back to a shorter fragment, a timeout, or an empty model reply. It would also back up your "measured where Gemma is reliable" story with real traces.

What it costs

  • Time. I'd estimate 2 to 3 hours, and more if something goes wrong. That covers a Sentry account and project, installing the packages, a preload file, and the Mastra integration. The docs list version requirements, including a recent @mastra/core, that I'd have to check against your install. Most of the model calls go through plain fetch to Ollama rather than Mastra agents, so I'd also write custom spans around each call. Then there are the traces and screenshots to capture.
  • Privacy. By default the SDK sends the inputs and outputs of model calls to Sentry's cloud, which means your resume and job text. You can turn that off with a dataCollection option. Even with it off, the help text saying nothing leaves your machine would have to change. I'd also drop the "runs locally" angle that your write-up and prize pitch lean on. That's a real conflict with the project's identity.
  • Setup. You need a Sentry account and DSN. I couldn't confirm pricing or a free tier from the docs, and the ₹0 budget is a hard rule for you.
  • Risk. It adds a new dependency to a working project, 15 hours before the deadline.

My recommendation: skip it. It's $100, you'd be competing against projects built around agent traces, and it works against your privacy story. The Gemma, Mastra and Overall prizes together are worth more, and an hour spent on the write-up and a real end-to-end run pays off more than Sentry. If you want a similar effect for free, I can add a small "run timeline" panel in the app showing each step's duration and model, using data we already have. It would take about 30 to 45 minutes. It's good for demo screenshots, though it isn't Sentry and wouldn't qualify for that prize.

Do you want the timeline panel, or should I move on to the write-up draft?

Sources: Sentry: Mastra agent tracing (Node), Sentry: Set up agent tracing (Node), Sentry: Instrument AI agents

You

ok lets finally work on submitting the draft. below is the template given by the judges:

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

What I Built

<!-- What does it do, and who is the friend or loved one you built it for? What problem does it solve for them? -->

Demo

<!-- Share a deployed link or a video demo. -->

Code

<!-- Show us the code! You can embed a GitHub repo directly into your post. -->

How I Built It

<!-- Which open-source AI did you use (open-weight models, agent harnesses, frameworks, local inference), and how is your project built around it? -->

Why Does Open Innovation Matter?

<!-- Why does open innovation matter for what you built? What did it make possible that a closed API wouldn't? -->

My Agent Session

<!-- Optional, but judges love it. Save your session with DevRelay and embed it with the agent_session tag (see the challenge page), or link to it. -->

Prize Categories

<!-- Which partner categories are you entering? List every one that applies, or remove this section. -->

<!-- Team Submissions: Please pick one member to publish the submission and credit teammates by listing their DEV usernames directly in the body of the post. -->

<!-- Thanks for participating! -->

Also adding a friend feedback/review would create better chances. In your post, tell us why open innovation matters for what you built. Does it run on a laptop with no internet? Keep someone's data off a server they don't control? Let you fine-tune, swap models, or change how your agent behaves? Cost nothing to run? Tell us where your open-based approach worked better than a closed one.

also give me a very unique and catchy title for the post, like a one liner which instantly raises curiosity amongst users and stops scrolling and have to then open my post reacd it engage with it, etc

Agent
Agent
Agent
Agent
Agent

The draft is in proofprep-post.md, following the judges' template. I filled in the facts I could verify from the project and left brackets for what only you know. It's 12:00 ish IST by now, so you have under half an hour, depending on when you read this. Check the clock first.

Title options (my pick is first)

  1. A 4B model invented 4 quotes on a resume. My app caught every one. This is the strongest because it has a concrete number, a villain and a payoff, and it ties to Gemma. Both figures come from your README's eval, so check them.
  2. I built a resume tool that's forbidden from flattering you
  3. Your resume doesn't need a better rewrite. It needs receipts.
  4. The AI said my friend had "excellent collaboration skills." Their resume said nothing of the sort. Only use this one if it really happened.
  5. I let two small AI models fight over one resume. Only one was allowed to speak.

What you must fill in before posting

  • Friend's name and role type. I didn't have them.
  • A real friend quote or two, and their reaction to the final version. I did not invent a review. The bullet list of changes comes from the feedback you relayed earlier. The quote slots are empty on purpose, since a fake testimonial would hurt you if a judge asks.
  • Demo link. A 60 to 90 second video or GIF; the suggested flow is in the draft.
  • GitHub URL for the repo embed, and make sure the repo is public.
  • Agent session section. Paste a DevRelay embed or delete the section.

Things to check before you publish

  • The eval numbers (8 of 11 wrong, 4 fabricated quotes, 2 of 9 for Qwen) are from your README, a single run on one resume and job post. The draft says so openly, which also helps with credibility.
  • The post says "26 tests". Run npm test on your machine to confirm the count, since the README still says 24.
  • The post says it works with the Wi-Fi off after the model download. The app has no external fonts or CDNs, but try it once with the internet disconnected so you can stand behind that.
  • It lists only Gemma and Mastra, the two categories that truly fit. Sentry isn't included.

I didn't claim fine-tuning or any real-model testing beyond the eval. If you ran the full UI against real Qwen and Gemma since then, add one sentence about what you saw, because that is stronger than anything I can write.

You
  1. Friend name: Pranshu. do not mention job role it can be generic, based on the user's interest.
  2. real friend review: The manual charge on making the edits to the resume and also the interview prep based on specific bullet points, gives me confidence to back up my work. you can frame this in a better sentence
  3. demo link: will send youtube video soon
  4. github url: https://github.com/DevAni27/proofprep
  5. help me with the agen session section at last after solving all above 4 problems
Agent

Prize Categories

  • Best Use of Gemma: Gemma 3 4B runs interview practice, and I measured where it is and isn't reliable and designed around it.
  • Best Use of Mastra: the human decision points are a Mastra workflow with suspend/resume that survives a restart.

Top comments (2)

Collapse
 
sanika_luktuke_26923f98b9 profile image
Sanika Luktuke •

Incredibly useful!

Collapse
 
aaryandeshmukh2005 profile image
Aaryan Deshmukh •

needed this