DEV Community

Mawaddah Dawood
Mawaddah Dawood

Posted on

🔐 Vault — Privacy-First Local AI for Sensitive Legal Documents

Hacktoberfest Weekend Challenge: Build for a Friend Submission 🤝

My Prompt & Theme: Why I Built It for a Friend.
This project is built for the Hacktoberfest 2026 Weekend Challenge: Build for a Friend.
I built Vault specifically for a close relative who works as a lawyer. In the legal sector, professionals handle highly confidential documents, court case files, and private client infrastructure agreements daily. They urgently need AI power to summarize long text files and extract dates, but ethically and legally, they cannot upload these documents to cloud-based AI services like ChatGPT due to strict data privacy laws and third-party data leak risks.
Vault explores an alternative: running an open-weight model locally so document contents can be processed directly on the user's own machine.
The complete implementation, code environment settings, and demonstration files are available transparently on my GitHub repository:
👉 GitHub: https://github.com/mawaddah-dawood/Vault
The Tech Stack & Architecture
To prioritize privacy, the core engine of Vault utilizes:
• AI Core: Google's gemma2:2b open-weight model.
• Local Inference Framework: Ollama configured for local processing.
• Frontend UI: Streamlit built using clean Python environments.
• Security Control: A specialized Privacy Dashboard communicating the application's local-processing status.
Local Processing Flow:
Confidential Text Document ➡️ Local RAM Ingestion ➡️ Ollama / Gemma 2 Local Inference ➡️ Structured Security & Legal Output (No External AI API Calls).
Creativity & Key Features
What makes Vault unique is that it is designed with a Cybersecurity Mindset:

  1. The Privacy Dashboard: An embedded interface UI that communicates the application's local-processing status and ensures no external AI APIs are configured.
  2. Potentially Sensitive Data Flags: The system leverages local AI capabilities to identify and flag potentially sensitive information such as phone numbers, email addresses, and other confidential terms, alerting the user to handle them securely. Real Execution & Test Results (Gemma 2 Output) Note: All names, contact details, case references, and dates shown below are entirely synthetic and created strictly for demonstration purposes. Testers can use the demo/sample_legal_case.txt file available in the repository. 🛡️ Application Interface & Local Loading Status:

Here is the exact live output generated locally by Vault when auditing a confidential legal case study file:

  1. Smart Summary

This legal case note outlines a dispute between Johnathan Doe and his opponent over a possible breach of a non-disclosure agreement (NDA) signed in early 2025. The opponent alleges Johnathan Doe accessed confidential information outside the designated network. The case involves a key hearing on October 15, 2026.
• Key Evidence Identified: Logs showing no data left the network between March and June 2025, and reports detailing local infrastructure compliance.

  1. Key Legal Information & Dates • Case Reference: IL-2026-9941 • Client Name: Johnathan Doe • Hearing Date: October 15, 2026 • NDA Sign Date: Early 2025 • Deadline for Evidence Submission: October 10, 2026
  2. Potentially Sensitive Data Flags 🛡️ • Personal Contact Information Detected: Johnathan Doe's phone number (+1-555-019-2834) and email address (j.doe@privatemail.com) were found in the plain text document. • Action: Flagged appropriately by the local model to remind the user to secure communication channels. Why Open-Source AI Matters Here Cloud-based AI can be unsuitable for workflows where sensitive documents cannot be transmitted to third-party services. Vault demonstrates the ultimate power of open innovation: • Local Data Sovereignty: Because document contents are processed locally rather than sent to an external AI API, the application does not provide those contents to a remote AI provider for processing. • Cost-Efficiency: Runs seamlessly on a standard 8GB RAM computer with absolutely zero operation or API token costs. How to Run Locally
  3. Install Ollama from ollama.com.
  4. Run the lightweight model: ollama run gemma2:2b.
  5. Run the application: python -m streamlit run app.py.

Designed with pride by Mawaddah Dawood | Computer Systems Engineering Student from Qalqilya, studying at Palestine Technical University - Kadoorie (Tulkarm, Palestine).

Top comments (0)