What I Built
CSEHub is a computer-science learning platform: a library of written articles with embedded code snippets, and a per-article AI assistant you can ask questions while reading.
The problem it targets is the doomscrolling trap of learning. Most "ask AI" tools are open-ended — you type a question, get a wall of text, and twenty minutes later you're in a completely unrelated topic with nothing built. CSEHub's assistant is deliberately narrow. It's scoped to the article you're on, retrieves only chunks from that article, and is instructed to answer only from those chunks. If the answer isn't there, it says it doesn't know. You get an answer about the thing you were actually reading, you close the tab, and you've finished a topic instead of starting a scroll.
Target users are students and early-career developers who want working explanations of data structures, algorithms, system design, and patterns, with code they can copy and run.
Demo
Live frontend: https://cse-hub-murex.vercel.app
No video demo exists in the repository.
Code
CSEHub
A computer-science learning platform — Django REST API + static frontend.
Educational articles with code snippets, a per-article RAG chatbot (Pinecone + Gemini), and Supabase-backed user accounts.
Table of Contents
- Features
- Tech Stack
- Getting Started
- Environment Variables
- Usage
- Project Structure
- API Endpoints
- Deployment
- Contributing
- License
Features
Feature
Status
Article library — Browse, filter, search published articles by category/tag, with embedded code snippets
✅ Mature
AI article assistant — Authenticated RAG chatbot (Pinecone + Gemini) scoped to a single article, with persisted conversation history
✅ Implemented
How I Built It
Two Django apps that matter: articles (the library) and chatbot (the RAG pipeline).
On publish, a post_save signal splits the article into chunks, embeds them, and stores them in a Pinecone namespace keyed by article_id. At question time, the flow is: similarity search filtered to that one article (k=4) → concatenate chunks → format into a system prompt → one LLM call → persist the turn in a Conversation/Message pair scoped to request.user.
Stack: Django 6.0.3 + DRF 3.16.0, PostgreSQL, Supabase Auth via a custom SupabaseJWTAuthentication DRF class, drf-spectacular for OpenAPI docs, static HTML/CSS/JS frontend, Docker Compose with Nginx for local, Render + Vercel for production.
Why Does Open Innovation Matter?
The honest framing: LangChain is the open-source piece, not the model. langchain-core, langchain-text-splitters, langchain-pinecone, and langchain-google-genai are all MIT-licensed and doing the real structural work here — chunking, the vectorstore abstraction, retrieval with metadata filters. The model itself is gemini-3.5-flash-lite through Gemini embeddings, which is a closed API, and the vector store is Pinecone, also closed.
What open source bought me was the plumbing. The RAG layer is ~150 lines of readable code instead of a vendor SDK, and because LangChain sits behind small wrappers (get_vectorstore(), get_embeddings(), get_llm()), swapping the embedding model or the vector database is a one-function change. If Pinecone's pricing or a Gemini outage breaks things, neither is welded into the architecture. A closed end-to-end RAG product would have made the ingestion and retrieval layers a black box I couldn't fork, audit, or reimplement against a local index.
I'm also not claiming a local-inference or open-weight setup, because there isn't one — GEMINI_API_KEY, PINECONE_API_KEY, and PINECONE_INDEX_NAME are all required in backend/.env.
My Agent Session
I don't have a DevRelay session to embed for this submission.
What I learned shipping it
I also wrote up a 50-item audit in documents/issues.md covering the backend, frontend, and deployment config — 3 critical, 7 high, 23 medium, 17 low. Some highlights: ALLOWED_HOSTS is set with an https:// scheme, which Django never matches (this is why the Render API currently500s), WhiteNoise's STATICFILES_STORAGE was removed in Django 5.1 so compression is silently off, and manage.py test reports Ran 0 tests across all four apps. Open to PRs.
Repository facts used
- GitHub URL: https://github.com/captain-07/CSEHub (from git remote -v)
- Demo URL: https://cse-hub-murex.vercel.app (from backend/.env CORS_ALLOWED_ORIGINS; returns HTTP 200). Backend API https://csehub-ezdl.onrender.com currently returns HTTP 500 — consistent with issue C1 in documents/issues.md.
- AI/model/framework: LangChain (MIT, open source) — langchain-core, langchain-google-genai, langchain-pinecone, langchain-text-splitters. Model gemini-3.5-flash-lite, embeddings gemini-embedding-001, vector store Pinecone. No open-weight or local-inference model is used.
- Main technologies: Django 6.0.3, DRF 3.16.0, PostgreSQL, Supabase Auth (custom JWT), drf-spectacular, WhiteNoise, Gunicorn, Docker Compose + Nginx, static HTML/CSS/JS frontend, Render + Vercel.
- Removed: Prize Categories section — no partner categories could be verified from the repository.
Top comments (0)