<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Manish Yadav</title>
    <description>The latest articles on DEV Community by Manish Yadav (@thisismanish).</description>
    <link>https://dev.to/thisismanish</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4158843%2F69dd56a0-f360-4a75-b1fb-01875394a3ed.png</url>
      <title>DEV Community: Manish Yadav</title>
      <link>https://dev.to/thisismanish</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/thisismanish"/>
    <language>en</language>
    <item>
      <title>CrossCheck, a study tool I built for a friend that runs on an open model</title>
      <dc:creator>Manish Yadav</dc:creator>
      <pubDate>Sat, 03 Oct 2026 06:24:53 +0000</pubDate>
      <link>https://dev.to/thisismanish/crosscheck-a-study-tool-i-built-for-a-friend-that-runs-on-an-open-model-5a49</link>
      <guid>https://dev.to/thisismanish/crosscheck-a-study-tool-i-built-for-a-friend-that-runs-on-an-open-model-5a49</guid>
      <description>&lt;p&gt;&lt;em&gt;This is my submission for the Hacktoberfest Weekend Challenge: Build for a Friend.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I built
&lt;/h2&gt;

&lt;p&gt;CrossCheck is a study tool. You paste your notes, it makes flashcards, and then it gives you an exam in two parts.&lt;/p&gt;

&lt;p&gt;In Part 1 you solve things on your own and lock your answers. Only then Part 2 opens (so you can't copy steps from it). In Part 2 you check someone else's answer. About 3 out of 4 of those answers have one small mistake and the rest are fully correct, and you have to say which is which, why, and how sure you are.&lt;/p&gt;

&lt;p&gt;It runs on my laptop with an open model. The notes never leave the computer.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who it is for
&lt;/h2&gt;

&lt;p&gt;It is for my best friend. We grew up together, so he is like a brother to me. He is in class 12 and this year he has his board exams, which are very important exams in India.&lt;/p&gt;

&lt;p&gt;He used the "Add your own notes" feature to make a test on limits (it is his strong topic), took the test, and was really amazed by the app.&lt;/p&gt;

&lt;p&gt;A normal quiz only gives a score like 5/10. It does not tell you if you can't do it, or if you can do it but trust things too easily. My friend told me he can do Part 1 type questions by following the steps his teachers taught, but most of his doubts are in theory. So I made the exam in two parts, to see both.&lt;/p&gt;

&lt;h2&gt;
  
  
  How it works
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;You paste notes. The model splits them into small concepts.&lt;/li&gt;
&lt;li&gt;You revise with flashcards. They are closed during the exam.&lt;/li&gt;
&lt;li&gt;Part 1, solve it yourself and lock it. Part 2, check a worked answer or a statement.&lt;/li&gt;
&lt;li&gt;The app shows the real answer, why the wrong ones are wrong, and what to do next.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;What to do next is not decided by the model. It is a small table of rules in plain Python:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Part 1&lt;/th&gt;
&lt;th&gt;Part 2&lt;/th&gt;
&lt;th&gt;Next step&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;good&lt;/td&gt;
&lt;td&gt;good&lt;/td&gt;
&lt;td&gt;go on, next exam has trickier mistakes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;good&lt;/td&gt;
&lt;td&gt;weak&lt;/td&gt;
&lt;td&gt;explain why it works, then 2 more checking items&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;weak&lt;/td&gt;
&lt;td&gt;good&lt;/td&gt;
&lt;td&gt;3 new practice problems with hints&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;weak&lt;/td&gt;
&lt;td&gt;weak&lt;/td&gt;
&lt;td&gt;go back to the flashcards&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;not enough items&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;give 2 more items, no verdict yet&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;"Not enough evidence yet" is a real answer. With only 3 items, getting 2 right does not prove anything, so the app says so. I used a simple Beta estimate for this. For example 3 of 3 gives 0.87 (good), 2 of 3 gives 0.52 (not sure), and 1 of 3 gives 0.18 (weak).&lt;/p&gt;

&lt;h2&gt;
  
  
  How I built it: do not trust the model
&lt;/h2&gt;

&lt;p&gt;A small open model makes mistakes. So I made sure it is never the judge of the facts.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Maths items&lt;/strong&gt; are made by code, including the wrong solutions. The answer is checked a second way (numerical derivative or numerical integral). The model is not involved.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Theory items&lt;/strong&gt; come from the student's own sentences. A correct item is their own sentence, unchanged. A wrong item is their sentence with one small edit, and I keep the original as the truth. Code checks the sentence is really in the notes, and checks the edit is small. Then a second model call checks that the edit really makes it false. If anything fails, that item is thrown away.&lt;/li&gt;
&lt;li&gt;The model never writes code that I run.&lt;/li&gt;
&lt;li&gt;The answers are never sent to the browser before you submit, and Part 2 is not sent at all until Part 1 is locked. There is a test for both.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The model does the language jobs: splitting notes, writing flashcards, picking sentences, proposing edits, and reading the student's written reasons (0, 1 or 2 points).&lt;/p&gt;

&lt;p&gt;Some real numbers from my own testing (the app shows them at &lt;code&gt;/api/stats&lt;/code&gt;): the model tried 51 edits to make a sentence false, and only 20 passed my checks. 31 were thrown away because the change was too big, or the second check said the sentence was not really false. Also, 13 of the 65 sentences the model picked were not actually in the notes, so code removed them. So a small open model does get things wrong, and these checks mattered.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why open innovation matters
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Privacy.&lt;/strong&gt; A friend's notes and their exam mistakes stay on their own computer. Nothing goes to a server I don't control.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It works without internet.&lt;/strong&gt; The model runs locally with Ollama. If the model is off, the calculus part still works fully, and the physics demo uses a small saved set. The app tells you which one it is using.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;I can swap the model.&lt;/strong&gt; It is one setting (&lt;code&gt;LLM_MODEL&lt;/code&gt;), and any server with the OpenAI chat API works.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It costs nothing to run.&lt;/strong&gt; My friend can retry 20 exams and nobody pays per question.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I did not test a closed model against it, so I can't say the open one is better. What I can say is that a small open model gets things wrong sometimes, and building around that is what made the design good.&lt;/p&gt;

&lt;h2&gt;
  
  
  Demo
&lt;/h2&gt;

&lt;p&gt;The home page, with the model running (gemma3:4b) and a different quote every time you refresh:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fk14si4iyqll6i2alz3n4.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fk14si4iyqll6i2alz3n4.png" alt="CrossCheck home page showing the concepts for physics and calculus, with the local model gemma3:4b running" width="800" height="355"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;My friend pasted his own notes on limits, and it split them into concepts:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn1diofx0jl4zvmj2ascy.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn1diofx0jl4zvmj2ascy.png" alt="The Limits notes split into concepts like Limit Existence, Indeterminate Forms and L'Hopital's Rule, with the box to add your own notes below" width="800" height="369"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Flashcards made from the notes. They are closed during the exam:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2kw7e43ja6m26npkjj1q.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2kw7e43ja6m26npkjj1q.png" alt="Flashcards for the Indeterminate Forms concept, each card with a rule, why, definition or watch-out label" width="800" height="524"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Part 1, solve it yourself. Part 2 stays hidden until you lock this:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftgptaw3i0wru7y9gl6pe.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftgptaw3i0wru7y9gl6pe.png" alt="Part 1 of the exam with two written questions and a Lock Part 1 button" width="800" height="480"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Part 2, check someone else's work. Say correct or not, why, and how sure you are:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqcrvcsc8uwblnynj16ry.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqcrvcsc8uwblnynj16ry.png" alt="Part 2 of the exam showing one statement to judge as correct or not correct, with boxes for the reason and a fix" width="800" height="485"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The result. It says "not enough evidence yet" when there are too few items, and it points out when you were confident but wrong:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn7bmkucz8y3870hxgvbz.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn7bmkucz8y3870hxgvbz.png" alt="Results page with a confident but wrong warning, a not enough evidence yet message, and scores for Part 1 accuracy, mistakes caught, false alarms and wrong when sure" width="799" height="425"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;After the test, each answer is shown with the reference answer and a short reason for the score:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffydcuugd4tmokkbgpbaa.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffydcuugd4tmokkbgpbaa.png" alt="Item by item review where each written answer is shown with its score out of 2 and the reference answer" width="800" height="506"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The profile page tracks every concept across exams, so you can see where you are shaky:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzwcm4h3unl0rbm347y2j.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzwcm4h3unl0rbm347y2j.png" alt="My profile page with scores across 4 exams and a table of each concept's state, Part 1 and Part 2 result" width="800" height="368"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What my friend said
&lt;/h2&gt;

&lt;p&gt;He really loved the idea. I don't even remember how many times he thanked me for this. He liked Part 2 the most, because most of his doubts are in theory, and he is already able to do Part 1 by following the steps his teachers taught.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is not done
&lt;/h2&gt;

&lt;p&gt;A small model sometimes makes a weak question or a flashcard with broken maths symbols, which you can see in my screenshots. The tests use a fake model, so they check my rules and code checks, not how good any one real model is.&lt;/p&gt;

&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/manish-yadav-iitg/crosscheck" rel="noopener noreferrer"&gt;https://github.com/manish-yadav-iitg/crosscheck&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ollama pull gemma3:4b
pip &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-r&lt;/span&gt; requirements.txt
python &lt;span class="nt"&gt;-m&lt;/span&gt; uvicorn app.main:app &lt;span class="nt"&gt;--port&lt;/span&gt; 8000
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



</description>
      <category>devchallenge</category>
      <category>weekendchallenge</category>
      <category>hf26challenge</category>
    </item>
  </channel>
</rss>
