DEV Community

Arthur031221
Arthur031221

Posted on

cardsmith: offline flashcards from your own PDFs, checked against the source

cardsmith: generate a deck from a text file, then study it

I study from long PDFs and slide decks often enough that flashcards would help, but making them by hand from a chapter of reading takes long enough that I usually skip the step and reread instead. Rereading is a weaker way to study than active recall, so I built cardsmith to close that gap.

What it does

Point cardsmith at a PDF, a PPTX, or a plain text file. It splits the document into chunks of roughly 120 to 900 words, sends each chunk to a local model through Ollama, and asks the model to write flashcards using only facts in that chunk. Every card comes back with a verbatim source_quote, the exact sentence the card is based on.

That is the part I care most about. Most AI flashcard generators hand you cards with no way to tell whether the model made something up. cardsmith checks whether the quote it returned is an actual substring of the source chunk, and if it is not, the card is flagged in the preview before you save anything. You can still edit or delete any card before it goes into your deck.

Cards live in a local SQLite database with a plain SM-2 scheduler, the same algorithm Anki started from. Four grade buttons record how well you knew each card and schedule the next review. When you want to study in Anki itself, export to a real .apkg file with a genanki-built deck.

What I measured

On a 3,248 word public domain biology chapter, cardsmith generated 152 cards in about 13 minutes on a MacBook Air. The mechanical quote check matched 142 of 152 cards to an exact substring of the source text. I then hand rated a 40 card sample against the source and found 37 factually accurate, 1 inaccurate, and 2 unusable, for 92.5 percent accuracy. The full method and raw cards are in the repository's eval directory.

What is rough

Card quality depends on the 4B model I default to, which occasionally paraphrases a quote instead of copying it, which is why the grounding check exists rather than trusting the model's claim. There is no OCR, so a scanned PDF with no text layer will not work without running it through an OCR tool first. It is single user and single machine by design, with no sync and no account.

Try it

ollama pull qwen3:4b
uvx --from git+https://github.com/Arthur031221/cardsmith cardsmith
Enter fullscreen mode Exit fullscreen mode

The repository is at https://github.com/Arthur031221/cardsmith, MIT licensed. I would welcome feedback, especially on the grounding check and on what would make the accuracy number more trustworthy.

Top comments (0)