Welcome to the 24-Hour Intensive Masterclass on AI Knowledge Architecture.
Traditional approaches to artificial intelligence treat AI as an ephemeral chat box: you ask a question, receive an answer, close the session, and lose all context. This masterclass establishes a paradigm shift moving from ephemeral query-time Retrieval-Augmented Generation (RAG) to an incrementally compiled, persistent, self-maintaining Knowledge Base (LLM Wiki).
Traditional Ephemeral AI Chat The AI Second Brain Paradigm
┌───────────────────────────────┐ ┌───────────────────────────────────┐
│ User Query ──> LLM ──> Output │ │ Raw Knowledge (PDFs, Notes, Web) │
│ (Session Ends = Reset) │ └─────────────────┬─────────────────┘
└───────────────────────────────┘ │ (Agent Ingests)
▼
┌───────────────────────────────────┐
│ Compiled Markdown Wiki & Graph │
│ (Compounding Knowledge Base) │
└─────────────────┬─────────────────┘
│ (Query / Synthesis)
▼
┌───────────────────────────────────┐
│ Persistent, Context-Aware Outputs│
└───────────────────────────────────┘
By the end of Day 2, every participant will have built, configured, and deployed a fully autonomous, local-first Second Brain using Obsidian, Claude Code / Claude Co-work, Model Context Protocol (MCP), and automated maintenance loops.
DAY 1: FOUNDATIONS, LLM MECHANICS & INGESTION ARCHITECTURE
Module 1.1: Demystifying the AI Engine Pre-Training, Compression & LLM Psychology
To direct an AI agent effectively, you must understand the underlying cognitive physics of Large Language Models (LLMs).
1. Pre-Training: The Lossy Compression of Human Knowledge
- Data Ingest & Scale: An LLM begins by ingesting terabytes of raw internet text (e.g., FineWeb data sets, Wikipedia, research papers).
- Statistical Compression: The neural network compresses this vast corpus into billions of parameters (e.g., Llama 2 70B, GPT-4, Claude). Karpathy describes this as creating a 1-Terabyte lossy “zip file” of the internet.
- Next-Token Prediction: At its core, the pre-trained “base model” is an extremely sophisticated probabilistic autocomplete. It does not “think” like a human; it predicts the most likely next mathematical token based on statistical patterns.
┌─────────────────────────┐ ┌──────────────────────────┐ ┌─────────────────────────┐
│ Raw Internet Text Corpus│ ───> │ GPU Cluster Training │ ───> │ Neural Network Weights │
│ (40+ Terabytes Text) │ │ (6,000+ GPUs, $2M+ Cost) │ │ (1-TB Lossy "Zip File") │
└─────────────────────────┘ └──────────────────────────┘ └─────────────────────────┘
2. Post-Training: Supervised Fine-Tuning (SFT) & Alignment
- From Autocomplete to Assistant: A raw base model completes text indiscriminately. To make it helpful, post-training replaces raw web text with thousands of curated multi-turn conversations between humans and assistants.
- Attaching the “Smiley Face”: SFT and Reinforcement Learning from Human Feedback (RLHF) attach an assistant “personality” to the pre-trained weights. When you ask a question, the model simulates a human data labeler responding according to system guidelines.
3. LLM Psychology & Known Limitations
- Vague Recollection vs. Direct Working Memory: Parametric knowledge (stored in model weights) behaves like a vague recollection of something read months ago. Tokens placed inside the active Context Window represent the model’s immediate working memory.
- Hallucination Physics: When forced to answer from parametric memory on rare or unverified facts, the model takes a probabilistic guess. We mitigate hallucinations not by increasing model size alone, but by supplying structured external context directly into its working memory.
Module 1.2: Working Memory, System 1 vs. System 2 Reasoning & Tool Use
1. System 1 vs. System 2 Cognitive Frameworks
Borrowing from Daniel Kahneman’s cognitive psychology framework, modern AI operating models function across two speed regimes:
┌────────────────────────────────────────────────────────────────────────────────────────┐
│ COGNITIVE PROCESSING MODES │
├───────────────────────────────────────────┬────────────────────────────────────────────┤
│ SYSTEM 1: INSTINCTIVE │ SYSTEM 2: RATIONAL/THINKING │
├───────────────────────────────────────────┼────────────────────────────────────────────┤
│ • Fast, cached, next-token generation │ • Slow, multi-step reasoning & planning │
│ • Fixed compute per generated token │ • Backtracking, assumption checking │
│ • Examples: 2+2=4, basic summarization │ • Examples: Complex code, math proofs │
│ • Risk: Hallucinations on non-cached facts│ • DeepSeek R1, OpenAI o1/o3, Claude 3.7 │
└───────────────────────────────────────────┴────────────────────────────────────────────┘
- System 1 (Token Sampling): The model runs a single forward mathematical pass per token. Compute per token is constant regardless of problem difficulty.
- System 2 (Reinforcement Learning & Thinking Models): Models trained via RL (e.g., DeepSeek R1, Claude 3.7 Thinking, OpenAI o1) learn to generate internal “chains of thought”. They re-evaluate intermediate steps (“wait, let me double check this math”), backtrack when hitting dead ends, and convert inference compute directly into problem-solving accuracy.
2. The Mechanics of Tool Integration
Because an isolated LLM is a closed mathematical system, we empower it with external execution tools via special protocol tokens:
- Search Interception: The model emits a special token.
- Execution Pause: Generation halts; an external application executes the query (e.g., Bing, Google, or local database).
- Context Injection: Raw retrieved text is copy-pasted into the active Context Window (working memory).
- Resumed Inference: The model generates its final response grounded directly in the injected tokens.
LLM Generation ──> Emits <search_start> ──> Pause Inference ──> Execute Web/DB Search
│
Grounded Output <── Read Working Memory <── Inject Clean Text <─────────┘
- Python Interpreters & Code Execution: For math, financial calculations, and data transformations, relying on mental arithmetic fails. By giving the LLM a Python interpreter, the model writes code to calculate exact answers deterministically.
Module 1.3: Personal Knowledge Management (PKM) Evolution & The Save-for-Later Paradox
1. The Lineage of PKM
To design a sustainable Second Brain, we study the history of personal knowledge management:
- Zettelkasten (Niklas Luhmann): An analog slip-box system utilizing unique index IDs, atomic notes (one idea per card), and dense manual cross-linking.
- Evergreen Notes (Andy Matuschak): Digital atomic notes concept oriented toward long-term concept evolution rather than temporary activity logs.
- Building a Second Brain (Tiago Forte & BASB): Popularized the CODE workflow ( C apture, O rganize, D istill, E xpress) and the PARA method:
- P rojects: Short-term efforts with explicit goals.
- A reas: Long-term responsibilities to maintain over time.
- R esources: Topics and interests for future reference.
- A rchive: Inactive items from the first three categories.
┌────────────────────────────────────────────────────────────────────────────────────────┐
│ THE PARA FRAMEWORK │
├───────────────────────┬───────────────────────┬───────────────────────┬────────────────┤
│ PROJECTS │ AREAS │ RESOURCES │ ARCHIVE │
│ (Short-Term Goals) │ (Ongoing Duties) │ (Topics/Interests) │ (Inactive) │
├───────────────────────┼───────────────────────┼───────────────────────┼────────────────┤
│ • Q3 App Release │ • Financial Health │ • Machine Learning │ • Completed 2025│
│ • Workshop Delivery │ • Team Management │ • Product Design │ Campaigns │
└───────────────────────┴───────────────────────┴───────────────────────┴────────────────┘
2. The “Save-for-Later” Paradox & Maintenance Fatigue
Why do 95% of personal knowledge bases fail?
- The Curator Trap: Humans collect bookmarks, web clips, PDFs, and screenshots with the intention of reviewing them later.
- The Bookkeeping Overhead: Filing, categorizing, cross-referencing, and updating index pages requires tedious manual effort. Within two weeks, maintenance debt accumulates, guilt sets in, and the knowledge vault becomes an abandoned digital graveyard.
- The Solution: Offload the tedious bookkeeping linking, summarizing, cataloging, and cross-referencing to an AI agent that never gets bored, never forgets cross-links, and can update 15 files in a single pass.
Module 1.4: The Karpathy LLM-Wiki Shift Ephemeral RAG vs. Compounding Wikis
In April 2026, Andrej Karpathy published the viral LLM Wiki pattern , completely transforming how AI knowledge bases are built.
1. Flaws of Query-Time RAG
Standard RAG systems (and chat file uploads) operate ephemerally:
- You upload raw documents.
- At query time, vector embeddings pull fragmented text chunks.
- The LLM re-derives connections from scratch every single session.
- When the session ends, the synthesized understanding disappears. Knowledge never compounds.
2. The LLM-Wiki Architecture Shift
Instead of searching raw files at query time, an AI agent incrementally compiles source documents into a persistent, interlinked Markdown Wiki.
┌────────────────────────────────────────────────────────────────────────────────────────┐
│ COMPARING KNOWLEDGE BASE ARCHITECTURES │
├──────────────────────────────┬─────────────────────────────────────────────────────────┤
│ FEATURE │ LLM WIKI / SECOND BRAIN OS │
├──────────────────────────────┼─────────────────────────────────────────────────────────┤
│ Core Mechanism │ Incremental compilation into linked Markdown files │
│ Persistence │ Compounding, permanent file vault on local disk │
│ Maintenance │ Automated AI agent handles linking, filing, and linting │
│ Infrastructure Cost │ Zero vector DB; plain text files + Git versioning │
│ Relationship Depth │ Explicit wiki links (`[[link]]`) + graph traversal │
└──────────────────────────────┴─────────────────────────────────────────────────────────┘
3. Core Structural Components
Karpathy’s paradigm divides the system into three simple layers:
┌─────────────────────────────────────────────────────────────────────────┐
│ THE 3-LAYER WIKI ARCHITECTURE │
├─────────────────────────────────────────────────────────────────────────┤
│ 1. RAW SOURCES LAYER (`raw/`) │
│ Immutable user input: PDFs, web clips, transcripts, chat exports │
├─────────────────────────────────────────────────────────────────────────┤
│ 2. COMPILED WIKI LAYER (`wiki/`) │
│ AI-maintained Markdown files: atomic concepts, entities, sources │
├─────────────────────────────────────────────────────────────────────────┤
│ 3. SCHEMA & SYSTEM CONTEXT (`CLAUDE.md` / `AGENTS.md`) │
│ Instructional contract governing styling, schemas, and rules │
└─────────────────────────────────────────────────────────────────────────┘
- Obsidian as the IDE, LLM as the Programmer, Wiki as the Codebase: You view and navigate the visual graph in Obsidian; the AI agent acts as the developer writing, editing, and refactoring the Markdown files.
Module 1.5: Environment Setup Obsidian, Claude Desktop, Claude Code & MCP
Participants will now complete a live setup on their machines
Step 1: Directory Scaffold Creation
Open your terminal (or Command Prompt) and execute the standard directory initialization:
mkdir -p ~/brain/{raw,wiki/{sources,concepts,entities,synthesis},projects,prompts,archive,scripts}
cd ~/brain
git init && git add . && git commit -m "Initial Second Brain scaffold"
Step 2: Install Obsidian & Open Vault
- Download and install Obsidian (free at obsidian.md).
- Choose “Open folder as vault” and select ~/brain.
Step 3: Configure Local REST API & MCP Bridge
- Inside Obsidian, navigate to Settings -> Community Plugins -> Turn on Community Plugins.
- Search for Local REST API , click Install , and click Enable.
- Open plugin settings and copy your auto-generated API Key.
- In your terminal, configure the Model Context Protocol (MCP) link to Claude:
claude mcp add-json obsidian-vault '{
"type": "stdio",
"command": "uvx",
"args": ["mcp-obsidian"],
"env": {
"OBSIDIAN_API_KEY": "YOUR_COPIED_API_KEY",
"OBSIDIAN_HOST": "127.0.0.1",
"OBSIDIAN_PORT": "27124"
}
}'
Step 4: Verification Test
Run the command: claude "List every file in my Obsidian vault."
If Claude reads back your directory structure, your agentic bridge is fully operational.
Module 1.6: Ingestion Protocols & Constructing the Raw-to-Wiki Pipeline (Hours 11–12)
1. Ingestion Channels
Raw knowledge enters the raw/ directory via multiple zero-friction pathways:
- Web Content: Obsidian Web Clipper browser extension (configured to auto-save Markdown directly to raw/).
- Audio & YouTube: Capturing transcripts via yt-dlp or youtube-transcript-api into raw/.
- PDFs & Books: Extracting text layers using pdftotext or OCRmyPDF.
- Voice Notes & Meetings: Audio recordings transcribed via Whisper into text files in raw/.
Web Clipper / Transcripts / PDFs / Audio
│
▼
┌───────────────┐
│ raw/ │ (Immutable Junk Drawer)
└───────┬───────┘
│
▼ [Inference Pass: Ingestion Skill]
┌───────────────┐
│ wiki/ │ (Atomic Pages & Links)
└───────────────┘
2. Hands-On Ingestion Walkthrough
- Drop Raw Source: Place a web clip or transcript (e.g., AI_2027_Overview.md) into ~/brain/raw/.
- Execute Ingestion Prompt: Run the canonical translation prompt:
Read the schema in CLAUDE.md. Process the file "AI_2027_Overview.md" from raw/.
1. Read it fully and extract key takeaways.
2. Create atomic source summary in wiki/sources/.
3. Extract new concepts into wiki/concepts/ and entities into wiki/entities/.
4. Connect every new page to existing concept pages using [[wikilinks]].
5. Append a summary line to wiki/index.md and a timestamp log entry to wiki/log.md.
6. Move processed file from raw/ to archive/.
- Observe real-time graph growth: Open Obsidian’s Graph View to watch new nodes and typed links connect automatically as the agent executes.
Need High-Impact Technical Content for Your Engineering Team?
I partner with developer-tooling startups, SaaS platforms, and engineering teams to translate complex infrastructure, agentic systems, and backend architecture into publication-grade technical writing.
Whether you need deep-dive architecture essays, hands-on developer tutorials, or technical counter-narratives:
- 📩 Email: abhishekninja2018@gmail.com
- 💼 LinkedIn: linkedin.com/in/abhishekninja
- 🐦 X (Twitter): x.com/AvishekBanzzov
- 🛠️ Capabilities: Long-form Technical Essays | Hands-On Tutorials | Developer Tooling Deep-Dives | Technical Counter-Narratives

Top comments (0)