This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
DemurrageDesk is a local RAG assistant for working with shipping line and terminal documents.
I built it around a problem I could easily imagine a friend working in logistics dealing with: finding demurrage and detention rules buried inside long tariff PDFs.
These documents can contain rules such as:
7 free days, days 8β14 at one rate, and day 15 onward at a higher rate.
Finding the right rule is only half the problem. The more dangerous part is applying the wrong container size, mixing up import and export terms, missing a tier boundary, or simply letting an LLM do the arithmetic.
DemurrageDesk is designed around the assumption that the model can be wrong.
You ask a question in plain English, the system retrieves the relevant document sections, and the answer includes the source file, page number, and the exact sentence used as evidence.
Then, when a charge needs to be estimated, the model extracts the relevant rules into structured data β but Python performs the actual calculation.
The goal isn't to make the LLM magically reliable.
The goal is to build a system that checks it.
Demo
π₯ Watch the DemurrageDesk demo
The video shows:
- Asking a question about a shipping tariff
- Retrieving the relevant source and page
- Verifying the citation
- Extracting the demurrage rate tiers
- Calculating the final charge deterministically with Python
Code
π’ DemurrageDesk: Port Terms Assistant
DemurrageDesk is a professional-grade, local RAG (Retrieval-Augmented Generation) assistant designed to parse complex port and shipping terminal documents. It allows users to query free-time rules, demurrage rates, and procedures, providing strictly cited answers and deterministic charge calculations.
Everything runs locally. Your documents never leave your machine.
β¨ Key Features
- Strict Citations: The assistant provides direct quotes from documents. A verification layer ensures that the AI didn't hallucinate the quote.
- Deterministic Calculator: Instead of letting the LLM do math (which is error-prone), the AI extracts the rules into a structured format, and a pure-Python engine calculates the final charges.
- Provider Filtering: Organize documents by shipping line or terminal; filter queries using the sidebar to avoid cross-provider confusion.
- Privacy First: Built with Ollama, ChromaDB, and Streamlit for a 100% local execution environment.
π οΈ Tech Stack
- LLM: Gemma (via Ollama)
- Embeddings: nomic-embed-textβ¦
The repository contains the complete implementation along with the calculator tests and setup instructions.
How I Built It
I wanted the entire pipeline to run locally because shipping tariffs, contracts, and terminal documents can contain sensitive commercial information.
The stack is:
- Gemma 3 through Ollama for local generation
- nomic-embed-text through Ollama for embeddings
- ChromaDB for local vector search
- PyMuPDF for PDF text extraction
- Pydantic for structured model output
- Streamlit for the interface
- Plain Python for the actual demurrage calculation
The application stores the provider/terminal as metadata during ingestion, which allows the UI to filter retrieval by shipping line or terminal rather than mixing documents from different providers.
The important part: evidence before answers
The RAG pipeline doesn't simply ask the model for an answer.
The response schema requires citations containing a source, page number, and a complete sentence copied from the retrieved text. The application then normalizes the citation and checks whether that quoted sentence actually exists in the retrieved document chunk. Only verified citations count toward found=True.
That means the model cannot simply invent a convincing-looking citation and have the application accept it.
The system prompt also explicitly tells the model to respect container size, import/export direction, and the applicable shipping line or terminal β and not to perform the charge calculation itself.
The LLM extracts the rule. Python does the math.
For a charge calculation, the model extracts structured information such as:
- free days
- starting day of each tier
- ending day of each tier
- rate
- currency
- notes about the rule
The result is represented as a structured Rule/Tier model rather than free-form text.
The actual calculation is deliberately kept outside the LLM.
The Python calculator resolves open-ended tiers automatically. For example:
- Days 1β7: free
- Days 8β14: $40/day
- Day 15 onward: $80/day
The calculator converts the open-ended first tier into days 8β14 and then calculates each applicable day deterministically.
This gives me a much clearer failure boundary:
LLM: retrieve and interpret the document.
Python: perform arithmetic.
I also tested the failure cases
One of the useful things about building this was discovering that a seemingly correct RAG application can still produce a dangerously wrong number.
The calculator tests include:
- days inside free time β
$0 - days 8β10 at
$40/dayβ$120 - crossing into the second tier at day 16
- open-ended tiers
- tiers supplied in reverse order
For example, the test suite explicitly checks that 10 days held with 7 free days produces $120, and that 16 days correctly crosses into the $80 tier.
The Streamlit UI also lets the user inspect and edit the extracted free days and rate tiers before the final calculation, rather than hiding the intermediate rule.
Why Does Open Innovation Matter?
This project is a good example of where local and open AI can be more useful than simply calling a hosted API.
Shipping tariffs and contracts can be commercially sensitive. Sending those documents to a third-party AI service may not be acceptable in some workflows.
With DemurrageDesk, the generation model, embedding model, vector database, document parsing, retrieval, and calculation all run locally.
The models can be pulled through Ollama, and the application uses ChromaDB as its local vector store.
More importantly, open/local AI gave me control over the entire reliability layer around the model.
I could decide:
- what evidence the model must provide
- how citations are verified
- when an answer is considered βfoundβ
- what information gets converted into structured rules
- and exactly how the final charge is calculated
I don't have to trust the model to do everything correctly.
I can put deterministic software around the parts where deterministic software is better.
That's the part of open innovation I found most interesting: the model doesn't have to be perfect if the system around it is designed to catch its mistakes.
Limitations
This is intentionally a focused prototype rather than a billing authority.
There are still important limitations:
- PDF tables can extract poorly.
- Scanned PDFs require OCR preprocessing.
- Small local models can still misunderstand complex language.
- Calendar-day vs. working-day rules require human confirmation.
- The extracted rule should be reviewed before acting on a calculated charge.
The README therefore explicitly treats human verification as part of the workflow.
The next step would be to put DemurrageDesk in front of someone who actually works with shipping tariffs every day and see which failure cases I haven't anticipated yet.
Prize Categories
- Best Use of Gemma
- Most Helpful Tool for a Friend
Top comments (0)