Building MyZubster: Docker, Local AI, RAG and the First On-Chain Proof on Ethereum Sepolia
Today I reached an important milestone in the development of MyZubster.
What started as a local experiment around AI-assisted knowledge retrieval has evolved into a working development stack combining Docker, Ollama, Qdrant, an API layer, Open WebUI, GitHub and Ethereum Sepolia.
The interesting part wasn't simply getting the individual components running. The goal was to connect them into a system that could ingest knowledge, retrieve the correct context, expose it to an AI interface, and begin attaching independently verifiable proofs to knowledge artifacts.
- Building the local AI stack The first step was creating a reproducible local environment. The development stack uses Docker Compose to orchestrate several components:
- Ollama for local language-model execution
- Qdrant as the vector database
- an API layer for retrieval and application logic
- Open WebUI as the user-facing AI interface
- the MyZubster application and knowledge-ingestion pipeline This allowed the project to move from isolated scripts toward an actual integrated environment. The repository containing the work is: github.com/nicolaususnicola-lgtm/myzubster-mvp
- Building the RAG knowledge pipeline The next challenge was knowledge ingestion. Documents have to be transformed into chunks before embeddings can be stored in Qdrant and later retrieved by the AI system. During testing, I found a problem in the original chunking logic: paragraph and line boundaries could produce poor chunks and potentially interfere with retrieval. The ingestion script was therefore modified to improve the chunking strategy while:
- preserving useful overlap;
- respecting paragraph and line boundaries;
- avoiding loops;
- skipping empty documents;
- keeping chunks suitable for retrieval. After the changes, the ingestion process completed successfully. In one of the verification runs, 46 knowledge chunks were loaded into Qdrant and one empty document was skipped. That gave us a concrete test that the modified ingestion pipeline was functioning.
- Testing retrieval through the API Storing embeddings isn't enough. The important question is whether the system can retrieve the right knowledge when a user asks a question. The RAG pipeline was therefore tested through the API endpoint. One particularly useful discovery involved the amount of context supplied to the model. By increasing AI_CONTEXT_LIMIT from 1 to 5, the system was able to recover the relevant context and return the expected GitHub repository information. This was a small configuration change, but an important lesson: retrieval quality depends not only on embeddings and chunking, but also on how much retrieved context is actually made available downstream.
- Moving from AI knowledge to verifiable knowledge The next experiment was more ambitious. Instead of asking users to simply trust that a knowledge artifact existed at a particular moment, MyZubster needed a way to create an independently inspectable proof. For this first experiment I created a Solidity smart contract called: MyZubsterProof.sol The contract records a bytes32 value representing a hash associated with a knowledge artifact. The objective is deliberately simple: the blockchain stores the proof, not the complete knowledge content. This distinction matters. Sensitive or large content does not need to be published on-chain. A cryptographic digest can instead provide a reference that can later be compared with the original artifact.
- First deployment on Ethereum Sepolia The contract was deployed to the Ethereum Sepolia test network. The deployment transaction is publicly inspectable through Etherscan, and the contract source was subsequently verified there with an Exact Match result. That gives the project its first independently inspectable on-chain component. For the development process, this was an important transition: local experiment → public cryptographic evidence The corresponding Solidity source and documentation were also committed to the project's GitHub repository.
- Connecting the evidence back to MyZubster The next step was documenting the work as a MyZubster knowledge card. The card contains references to:
- the project repository;
- the local Docker test;
- the chunking improvement commit;
- the Sepolia proof commit;
- the verified Sepolia contract;
- the Sepolia deployment transaction. An interesting bug appeared here too. The profile builder limits a knowledge card to 12 sources, and initially some sources had been entered using separate lines for title, URL and description. The parser treated those lines as separate sources. The expected format was actually: name | URL | note with one source per line. After normalizing the source list, the card contained six sources and could be successfully saved as a private draft before publication. It's a tiny implementation detail, but exactly the kind of thing that appears when a prototype starts becoming a real product. What exists now At this stage, MyZubster has demonstrated several pieces of the architecture working together: Docker environment → knowledge ingestion → Qdrant → RAG retrieval → AI interface/API → GitHub evidence → Ethereum Sepolia proof → MyZubster knowledge card There is still work to do. In particular, the complete Open WebUI integration deserves additional end-to-end verification, and the relationship between knowledge cards and on-chain hashes can be automated further. But the important milestone is that the project now has both a local AI knowledge pipeline and its first publicly inspectable blockchain proof. Next steps The next development phase will focus on tightening the connection between these layers. Instead of manually creating blockchain proofs, the long-term flow could become: Knowledge Card → canonical representation → SHA-256 → on-chain attestation → transaction reference → verification That would allow a MyZubster knowledge artifact to carry not only sources and AI-readable context, but also a cryptographic proof that can be independently checked. The goal isn't to put knowledge on a blockchain. The goal is to make claims about digital knowledge more inspectable, attributable and verifiable. And this first Sepolia deployment is a small but concrete step in that direction.
Top comments (0)