DEV Community

Cover image for PrepFlow — From Study Material to an Actual Study Plan
Piyush
Piyush

Posted on

PrepFlow — From Study Material to an Actual Study Plan

Hacktoberfest Weekend Challenge: Build for a Friend Submission 🤝

What if AI didn't just explain your notes, but actually figured out what you should study next?

This is my submission for the Hacktoberfest Weekend Challenge: Build for a Friend.

I built PrepFlow because students don't usually have a shortage of study material.

They have the opposite problem.

Too much of it.

PDFs. Notes. Syllabi. Previous-year questions. Multiple subjects. Weak topics. Different confidence levels. And a deadline that keeps getting closer.

The difficult question isn't:

“What does this PDF say?”

It's:

“Given everything I need to learn, how should I spend the limited time I have left?”

That's the problem PrepFlow tries to solve.


🧠 The Idea

PrepFlow takes this:

                    BEFORE PREPFLOW

       ┌─────────┐   ┌─────────┐   ┌─────────┐
       │  PDFs   │   │  Notes  │   │  Syllabus│
       └────┬────┘   └────┬────┘   └────┬────┘
            │             │             │
            └─────────────┼─────────────┘
                          │
                          ▼
                  ┌───────────────┐
                  │     STUDENT   │
                  │               │
                  │ "What do I    │
                  │ study today?" │
                  └───────────────┘
Enter fullscreen mode Exit fullscreen mode

and turns it into:

                     WITH PREPFLOW

 PDF / TEXT
     │
     ▼
┌──────────────┐
│   EXTRACT    │
│ page-aware   │
│ text         │
└──────┬───────┘
       ▼
┌──────────────┐
│    CHUNK     │
│ deterministic│
│ boundaries   │
└──────┬───────┘
       ▼
┌──────────────┐
│    GEMMA     │
│ understand   │
│ material     │
└──────┬───────┘
       ▼
┌──────────────┐
│   VERIFY     │
│ evidence +   │
│ structured   │
│ output       │
└──────┬───────┘
       ▼
┌──────────────┐
│   PRIORITIZE │
│ importance   │
│ weakness     │
│ urgency      │
│ difficulty   │
└──────┬───────┘
       ▼
┌──────────────┐
│    PLAN      │
│ capacity +   │
│ prerequisites│
│ revision     │
└──────┬───────┘
       ▼
┌──────────────────────────┐
│ TODAY                    │
│                          │
│ 1. Dynamic Programming  │
│ 2. Graph Traversal      │
│ 3. Revise Greedy        │
│                          │
│ 145 / 160 min planned   │
└──────────────────────────┘
Enter fullscreen mode Exit fullscreen mode

The important part is that AI doesn't generate the final timetable.

That distinction shaped almost the entire architecture.


🎯 The Core Principle

I split the system into two responsibilities:

AI should do Software should do
Understand unstructured material Calculate available time
Identify topics Calculate priority
Identify subtopics Enforce capacity
Suggest prerequisites Validate evidence
Estimate semantic difficulty Persist state
Connect topics to source evidence Schedule tasks
Interpret weak-topic hints Schedule revision
Produce structured analysis Handle progress
Recover missed work

Why?

Because I don't want an LLM deciding whether:

“You have 120 minutes available, so I'll give you 157 minutes of work.”

That's not intelligence.

That's a bug.

So PrepFlow follows a simple rule:

AI interprets. Deterministic software decides.


🔄 The Full PrepFlow Pipeline

The repository currently implements the complete path from material ingestion through planning, daily execution, revision, and progress.

┌─────────────────────┐
│  STUDY SETUP        │
│                     │
│ Exam date           │
│ Hours/day            │
│ Study days           │
│ Weak topics          │
│ Confidence           │
│ Buffer               │
└──────────┬──────────┘
           │
           ▼
┌─────────────────────┐
│ MATERIAL INGESTION   │
│                     │
│ PDF / pasted text   │
│ Page preservation   │
└──────────┬──────────┘
           │
           ▼
┌─────────────────────┐
│ DETERMINISTIC       │
│ CHUNKING             │
└──────────┬──────────┘
           │
           ▼
┌─────────────────────┐
│ GEMMA               │
│                     │
│ Topics              │
│ Subtopics           │
│ Difficulty          │
│ Importance          │
│ Prerequisites       │
│ Evidence            │
└──────────┬──────────┘
           │
           ▼
┌─────────────────────┐
│ VALIDATION          │
│                     │
│ JSON → Zod          │
│ Evidence matching   │
│ Page derivation     │
└──────────┬──────────┘
           │
           ▼
┌─────────────────────┐
│ KNOWLEDGE MODEL     │
│                     │
│ Subjects            │
│ Topics              │
│ Prerequisites       │
│ Evidence            │
└──────────┬──────────┘
           │
           ▼
┌─────────────────────┐
│ PRIORITY ENGINE     │
│                     │
│ Importance          │
│ Weakness            │
│ Foundation          │
│ Difficulty          │
│ Urgency             │
│ Emphasis            │
│ Evidence            │
│ Revision            │
└──────────┬──────────┘
           │
           ▼
┌─────────────────────┐
│ CAPACITY ENGINE     │
│                     │
│ Available minutes   │
│ Buffer              │
│ Revision reserve    │
│ Exam deadline       │
└──────────┬──────────┘
           │
           ▼
┌─────────────────────┐
│ PLANNER             │
│                     │
│ MUST                │
│ SHOULD              │
│ IF TIME             │
│ DEFER               │
└──────────┬──────────┘
           │
           ▼
┌─────────────────────┐
│ DAILY EXECUTION     │
│                     │
│ Start               │
│ Complete            │
│ Skip                │
│ Progress            │
└──────────┬──────────┘
           │
           ▼
┌─────────────────────┐
│ REVISION + RECOVERY │
│                     │
│ +1 / +3 / +7        │
│ Missed tasks        │
│ Rebalance           │
│ Replan              │
└─────────────────────┘
Enter fullscreen mode Exit fullscreen mode

🤖 Why This Isn't "Chat With Your PDF"

There are thousands of projects that can answer:

“Summarize chapter 4.”

That wasn't the product I wanted to build.

PrepFlow's model produces structured knowledge, not the final answer the student follows.

The application then transforms that knowledge into an executable workflow.

                 GENERIC PDF CHAT

PDF ──────► LLM ──────► Answer


                 PREPFLOW

PDF
 │
 ▼
Extraction
 │
 ▼
Chunking
 │
 ▼
Gemma
 │
 ▼
Structured Topics
 │
 ▼
Evidence Verification
 │
 ▼
Priority Engine
 │
 ▼
Capacity Engine
 │
 ▼
Planner
 │
 ▼
Daily Tasks
 │
 ▼
Progress
 │
 ▼
Revision
Enter fullscreen mode Exit fullscreen mode

The difference is subtle but important:

The output isn't text.

The output is a decision system.


🔍 Evidence Is First-Class

One of the parts I cared about most was preventing the model from inventing study topics.

If Gemma says:

“Dynamic Programming is highly important.”

PrepFlow doesn't simply trust it.

The model must provide source evidence.

The system then checks that evidence against the extracted source text and derives the PDF page programmatically.

Gemma says:

Topic:
Dynamic Programming

Evidence:
"Dynamic programming solves problems
by combining solutions to overlapping
subproblems."

             │
             ▼

       SOURCE MATERIAL

             │
             ▼

     Does this exact evidence
       exist in the source?

        ┌────┴────┐
       YES        NO
        │          │
        ▼          ▼
    ACCEPT       REJECT
        │
        ▼
   derive PDF page
Enter fullscreen mode Exit fullscreen mode

This gives each topic a traceable relationship back to the material.

The repository documents the same architecture: model output is parsed, schema-validated, then evidence-verified before topics are accepted.


🧮 How PrepFlow Decides What Matters

Not every topic deserves equal time.

PrepFlow combines multiple bounded factors:

Factor What it represents
Importance How important the topic is
Weakness How weak the student is
Foundation Whether other topics depend on it
Difficulty How demanding it is
Urgency How close the exam is
Emphasis Explicit emphasis in the source
Evidence Strength of supporting evidence
Revision Need for later review

This produces a bounded priority score rather than asking the model to invent one.

                  TOPIC PRIORITY

Importance ────────┐
Weakness ──────────┤
Foundation ────────┤
Difficulty ────────┤
Urgency ───────────┤
Emphasis ──────────┼──► PRIORITY SCORE ──► PLAN
Evidence ──────────┤
Revision ──────────┘
Enter fullscreen mode Exit fullscreen mode

⏱️ The Planner Knows When You Don't Have Enough Time

This was one of the most important product decisions.

Suppose the student has:

40.5 hours of estimated work

but only:

24 hours available

A typical AI timetable might confidently distribute everything across the available days.

PrepFlow doesn't.

It tells the truth.

ESTIMATED WORK
████████████████████████████████████████ 40.5h

AVAILABLE TIME
████████████████████████ 24.0h

                         ───────────────
                         16.5h OVERLOAD
Enter fullscreen mode Exit fullscreen mode

Then it prioritizes the work:

Tier Meaning
🔴 MUST Highest-value work to protect
🟠 SHOULD Important if capacity permits
🟡 IF TIME Useful but lower priority
⚪ DEFER Cannot realistically fit

The planner therefore answers two questions:

  1. What should I study?
  2. What should I stop pretending I can finish?

That second question is surprisingly important.


🧱 Prerequisites Matter

A study plan shouldn't tell someone to learn:

“Advanced Dynamic Programming”

before:

“Dynamic Programming fundamentals.”

PrepFlow models prerequisite relationships.

Arrays
  │
  ▼
Recursion
  │
  ▼
Dynamic Programming
  │
  ├──────────────► Knapsack
  │
  └──────────────► Longest Common Subsequence
Enter fullscreen mode Exit fullscreen mode

The planner can therefore account for foundational topics instead of treating every topic as an independent checkbox.


📅 A Plan Is Not Finished When It Is Generated

Real students miss tasks.

That's normal.

So PrepFlow treats planning as a feedback loop:

          ┌──────────────┐
          │ INITIAL PLAN │
          └──────┬───────┘
                 │
                 ▼
          ┌──────────────┐
          │     TODAY    │
          └──────┬───────┘
                 │
        ┌────────┼────────┐
        ▼        ▼        ▼
     COMPLETE   SKIP    MISS
        │        │        │
        └────────┴────────┘
                 │
                 ▼
        ┌─────────────────┐
        │ RECOVERY /      │
        │ REBALANCE       │
        └────────┬────────┘
                 │
                 ▼
          NEW PLAN VERSION
Enter fullscreen mode Exit fullscreen mode

Completed work is preserved.

Unfinished work can be recovered or re-planned.

This is much closer to how real studying works than generating one static timetable.


🔁 Revision Is Built Into the Plan

Studying a topic once isn't enough.

PrepFlow schedules revision sessions after learning using spaced intervals.

Conceptually:

LEARN
  │
  ├──── +1 study day ────► REVISION 1
  │
  ├──── +3 study days ────► REVISION 2
  │
  └──── +7 study days ────► REVISION 3
Enter fullscreen mode Exit fullscreen mode

The planner also respects the exam boundary and available study days.


🏗️ Architecture

The repository is a TypeScript monorepo containing the React frontend, Express/TypeScript API, and shared schemas/types.

PrepFlow/
│
├── apps/
│   ├── api/
│   │   ├── domain/
│   │   │   ├── planning/
│   │   │   └── analysis/
│   │   │
│   │   ├── infrastructure/
│   │   │   ├── prisma/
│   │   │   ├── PDF extraction
│   │   │   └── AI providers
│   │   │
│   │   └── HTTP API
│   │
│   └── web/
│       ├── React
│       ├── TypeScript
│       ├── Vite
│       └── Tailwind
│
├── packages/
│   └── shared/
│       └── Zod schemas + shared types
│
├── docs/
│   ├── analysis.md
│   ├── planning.md
│   └── decisions/
│
└── PostgreSQL + Prisma
Enter fullscreen mode Exit fullscreen mode

🧠 Why the AI Layer Is Replaceable

I didn't want the entire application to become dependent on one model.

The architecture therefore separates:

                APPLICATION

                    │
                    ▼
             ┌─────────────┐
             │ AI PROVIDER │
             │ INTERFACE   │
             └──────┬──────┘
                    │
          ┌─────────┼─────────┐
          ▼         ▼         ▼
       Ollama     Gemini     Fake
       + Gemma     + Gemma   Provider
Enter fullscreen mode Exit fullscreen mode

The model can change without rewriting the planning engine, database layer, or frontend.

The repository currently exposes ollama, gemini, and fake provider modes.


🛡️ AI Security and Trust Boundaries

The system treats uploaded study material as untrusted input.

USER MATERIAL
     │
     │ untrusted
     ▼
┌──────────────┐
│ nonce +      │
│ delimiters   │
└──────┬───────┘
       ▼
     GEMMA
       │
       ▼
┌──────────────┐
│ tolerant     │
│ JSON parsing │
└──────┬───────┘
       ▼
┌──────────────┐
│ Zod schema   │
│ validation   │
└──────┬───────┘
       ▼
┌──────────────┐
│ source       │
│ evidence     │
│ verification │
└──────┬───────┘
       ▼
     ACCEPT
Enter fullscreen mode Exit fullscreen mode

The model receives no application secrets or tools.

Its output is never treated as trusted application state.

The repository also documents bounded retries, analysis concurrency limits, upload limits, CORS controls, and environment-based secrets.


🧪 Testing the Important Parts

I didn't want the project to only work in a happy-path demo.

The current test suite covers multiple layers:

Layer What is tested
Shared Schemas and shared logic
API Business logic and routes
Web Frontend behavior
HTTP End-to-end API routing
Planning Priority, capacity, scheduling
Randomized invariants Planner safety properties
PostgreSQL Real persistence
Concurrency Rebalance/version behavior
PDF extraction Real document parsing

Current validation:

423 automated tests passing

and

20/20 real PostgreSQL integration tests passing

The integration tests include persistence and concurrent plan-rebalancing scenarios.


📊 What the Test Numbers Actually Mean

AUTOMATED TESTS

Shared       ████████████████████  22
API          ████████████████████ 358
Web          ████████████████████ 43
                                  ───
TOTAL                              423
Enter fullscreen mode Exit fullscreen mode

Real database integration:

PostgreSQL integration

Passed     ████████████████████ 20
Failed                           0
Enter fullscreen mode Exit fullscreen mode

I also validated the PDF extraction pipeline against a real academic PDF and compared the extracted page structure against the source.


🧰 Technology Stack

Layer Technology Why
Frontend React + TypeScript Component-based UI
Build Vite Fast development/build
Styling Tailwind CSS Consistent UI system
Backend Node.js + Express Lightweight API
Language TypeScript Shared type safety
Validation Zod Runtime contracts
Database PostgreSQL Relational persistence
ORM Prisma Type-safe DB access
PDF PDF.js Page-aware extraction
AI Gemma Open-weight material analysis
Local AI Ollama Local model execution
Testing Vitest Fast automated testing
CI GitHub Actions Automated verification

🌍 Why Open Innovation Matters

Study material can be surprisingly sensitive.

It can contain:

  • Personal information
  • University documents
  • Assignments
  • Private notes
  • Exam preparation
  • Proprietary material

That's why I wanted PrepFlow's architecture to support local/open-weight AI, rather than making a cloud LLM the only possible path.

Gemma sits behind an AI-provider abstraction instead of being hardcoded into the application's business logic.

That gives the project a useful property:

The intelligence can evolve without rebuilding the product around a single AI provider.


🆚 What PrepFlow Is — And Isn't

PrepFlow is PrepFlow isn't
AI-assisted planning A generic chatbot
Evidence-backed analysis Blind LLM output
Deterministic scheduling LLM-generated timetable
Capacity-aware “Everything fits” fantasy
Revision-aware One-time checklist
Explainable Black-box prioritization
Open-weight AI compatible Locked to one provider
Built around execution Just another note summarizer

💡 The Engineering Lesson

The biggest thing I learned wasn't how to connect an AI model to a backend.

It was learning where not to use AI.

A tempting architecture would have been:

PDF
 │
 ▼
LLM
 │
 ▼
"Here is your study plan."
Enter fullscreen mode Exit fullscreen mode

It's easy.

It's also difficult to trust.

PrepFlow instead looks more like:

                 AI
                  │
                  ▼
        UNDERSTAND THE MATERIAL
                  │
                  ▼
             STRUCTURED DATA
                  │
                  ▼
        ┌────────────────────┐
        │ DETERMINISTIC CODE │
        │                    │
        │ Validate           │
        │ Prioritize         │
        │ Calculate          │
        │ Schedule           │
        │ Revise             │
        │ Recover            │
        └─────────┬──────────┘
                  │
                  ▼
             REAL PLAN
Enter fullscreen mode Exit fullscreen mode

That separation gives me much more confidence in the system.


🚀 What's Next

PrepFlow is intentionally an MVP.

The next areas I'd explore are:

Area Possible improvement
Documents OCR for scanned PDFs
AI More provider/model evaluation
Knowledge Better prerequisite inference
Planning Calendar-aware scheduling
Progress More sophisticated learning models
AI runtime Additional local/open-weight providers
UX Topic editing and plan refinement
Validation Larger real-student testing

The current repository roadmap already separates the completed repository/database, ingestion, AI analysis, and planning milestones from later topic-review, polish, deployment, and real-user testing.


🔗 Try It

Source Code

github.com/piyusshhjangid/PrepFlow

Local Development

npm install
npm run db:migrate
npm run db:seed

npm run dev:api
npm run dev:web
Enter fullscreen mode Exit fullscreen mode

Then open the web application.

The repository also documents the local Gemma/Ollama and hosted Gemma provider setup.


❤️ Why “Build for a Friend” Matters

Building for a friend changes the question.

Instead of:

“What AI feature can I add?”

you start asking:

“What is actually making this person's life harder?”

For PrepFlow, the answer was:

Students don't necessarily need more study material.

They need help turning the material they already have into a realistic sequence of actions.

So that's what I built.

Not another chatbot.

Not another PDF summarizer.

A system that tries to answer one deceptively difficult question:

What should I study next?


Built for the Hacktoberfest Weekend Challenge — Build for a Friend.

devchallenge #weekendchallenge #hf26challenge #gemma

Top comments (0)