This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend.
What I Built
A friend of mine has a large collection of study PDFs spread across different folders. The problem wasn't a lack of study material. The problem was finding the right material quickly and knowing what each document actually contained.
So I built StudyShelf AI, a local-first study organizer that uses an open-weight AI model to turn a messy collection of study PDFs into an organized, searchable study shelf.
The idea is simple:
Instead of manually opening dozens of PDFs to figure out what they contain, upload them to StudyShelf AI and let local AI organize them.
For each document, the application generates:
- Subject/category
- Relevant tags
- Priority
- Short summary
The result is a much easier way to understand and search through a large study library.
Demo
Video demo:
Watch the StudyShelf AI demo
The demo shows the complete workflow:
- Running the Qwen model locally through LM Studio
- Opening the StudyShelf AI application
- Uploading a study PDF
- Extracting the PDF text locally
- Sending the extracted text to the local AI model
- Generating the category, tags, priority and summary
- Adding the document to the searchable study shelf
Code
GitHub repository:
MANAV-MISHRA-BYTES/StudyShelf-AI
(https://github.com/MANAV-MISHRA-BYTES/StudyShelf-AI)
The complete source code, requirements, setup instructions and project documentation are available in the repository.
How I Built It
StudyShelf AI is built with Python and Streamlit.
The core pipeline is:
Study PDF
↓
PyMuPDF
↓
Local text extraction
↓
LM Studio
↓
Qwen3.8 9B
↓
Category + Tags + Priority + Summary
↓
Searchable Study Shelf
↓
CSV Export
Tech Stack
- Python
- Streamlit
- PyMuPDF
- Pandas
- Requests
- LM Studio
- Qwen3.8 9B Q5_K_M
PyMuPDF handles local text extraction from PDF files.
Streamlit provides the user interface.
Requests communicates with the local LM Studio API.
Pandas is used to organize and export the generated document index.
LM Studio runs the open-weight model locally and provides the local API used by the application.
For this project, I used Qwen3.8 9B Q5_K_M through LM Studio.
Why Does Open Innovation Matter?
Open-weight AI was important to this project because the application deals with study documents that may contain private notes, assignments, personal annotations and other information that a student may not want to upload to a third-party AI service.
With StudyShelf AI, the core document organization workflow can remain on the user's computer:
Study PDF
↓
Local text extraction
↓
Local LM Studio server
↓
Open-weight Qwen model
↓
Document organization
There is no requirement for a cloud AI API for the core organization task.
This gives the user more control over their study material and avoids sending the documents to an external AI service just to categorize and summarize them.
It also means the model can be changed or upgraded without rebuilding the entire application around a proprietary API.
For this project, open AI was not just a requirement of the challenge. Local inference directly supported the privacy goal of the application.
What I Learned
One of the biggest lessons from this project was that a useful AI application does not necessarily need to be huge.
I initially considered building a much larger study assistant with RAG, a database, authentication, question answering and several other features.
But that would have added complexity without solving the immediate problem better.
Instead, I focused on one useful workflow:
Take a messy collection of study PDFs and make it understandable.
That kept the application small enough to build during the challenge while keeping the AI component central to the solution.
Current Limitations
StudyShelf AI is currently an MVP, so there are several things it does not do yet:
- Scanned or image-only PDFs are not currently OCR'd.
- It does not yet provide full RAG-based question answering.
- The organized library is not currently persisted in a database.
- The application expects a local LM Studio-compatible server.
- Search currently operates on generated document metadata rather than performing semantic search across the full document contents.
These limitations also provide a clear path for future versions.
What's Next?
Some improvements I would like to add in future versions are:
- OCR support for scanned documents
- Persistent document libraries
- Semantic search
- RAG-based question answering
- Flashcard generation
- Quiz generation
- Automatic revision scheduling
- Duplicate document detection
- Better document previews
- Support for additional local open-weight models
Hacktoberfest
StudyShelf AI was built for the Hacktoberfest Weekend Challenge: Build for a Friend.
The project addresses a real study-organization problem for a friend while using an open-weight AI model and local inference at the core of the solution.
The challenge encouraged building something useful for a real person rather than simply building a large project for the sake of complexity.
I therefore focused on one small but practical problem: making a large collection of study material easier to understand and navigate.
Prize Categories
No partner category is being claimed for this submission.
My Agent Session
No agent session is being submitted.
Thanks for checking out StudyShelf AI.
Top comments (0)