DEV Community

Anthony
Anthony

Posted on

Briefly: A Local-First, Zero-Data-Egress Filing Assistant for Legal Practice

Hacktoberfest Weekend Challenge: Build for a Friend Submission 🤝

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

What I Built

My wife is a practicing attorney managing an active case load. Like many legal practitioners working against tight court filing deadlines, her day-to-day workflow involves downloading court orders, disclosure bundles, witness statements, client emails, and vendor invoices directly into her local Downloads folder with the intention of filing them later.

Over time, this operational bottleneck created significant administrative drag. Her Downloads directory accumulated hundreds of ambiguously named files—such as WhatsApp Document 2026-09-14 at 11.23.pdf, scan_0042.pdf, and generic court export names. Locating a specific pleading or affidavit before a hearing required manual searching through unstructured files, wasting valuable time and interrupting focused legal work.

I built Briefly to help her work smarter by automating that administrative overhead.

Briefly is a lightweight, local-first document filing assistant. It monitors an inbox directory, extracts text from supported documents directly on her machine, and queries a local open-weight model to identify the corresponding legal matter and document category. High-confidence matches are filed automatically into a standardized matter directory structure. Any ambiguous, scanned, or low-confidence files remain untouched in the inbox for quick manual review.

Handing It Over: Her Real-World Feedback & Immediate Iteration

The challenge prompt specifically notes: "Bonus points if you actually hand it over and tell us what they said."

When I deployed Briefly on my wife's machine with her actual workflow documents, her immediate feedback drove two crucial product insights:

  1. Eliminating Cognitive Friction (Minutes vs. Seconds): When she opened the settings dialog and saw the cooldown timer configured in seconds, she gave immediate, practical feedback: > "I entered 30 mins but is moving stuff in less than a minute. What's going on with this?"

That's when I realized I was having her enter the cooldown as an integer in seconds but she thought she needed to use minutes.

She was completely right. Engineers think in seconds; busy professionals think in minutes. I immediately adjusted the interface to present standardized minute presets (1 min, 5 min, 10 min, 15 min, 30 min, 1 hr) while managing the underlying seconds conversion transparently in the background.

  1. Validating the Data Model (The Matter Search Request): Once Briefly completed the first-pass discovery and organized her scattered documents into clean matter folders, she clicked through the directory tree and asked: > "This is well organized. Can I now search for specific keywords across all the documents inside this matter?"

In product development, a user requesting search across newly organized data is strong confirmation that the foundational organization problem has been solved. She immediately trusted the matter hierarchy and saw it as an operational asset. While cross-document indexing was out of scope for a weekend MVP, it validated the underlying folder architecture and established our roadmap for version 2.

Demo

Here is a walkthrough of Briefly running locally:

Per the challenge announcement's invitation to use ElevenLabs for video narration, the walkthrough audio was generated using **ElevenLabs.

Briefly Local Dashboard

Core Workflow:

  1. First-Pass Collective Discovery: When starting with an unorganized inbox and an empty library, Briefly analyzes the document batch collectively to propose canonical matter folder names (e.g., Marcus Sterling v Gopaul & Colfire, Estate of Joseph Ramdial). It never creates directories or moves files without explicit user confirmation.
  2. Deterministic Classification & Filing: Validated legal files are routed into standardized subdirectories: Pleadings, Correspondence, Disclosures, Orders, and Evidence.
  3. Non-Case Document Routing: Non-legal receipts, tickets, and personal office files can be excluded with a single click, moving them directly to ~/Documents so they do not clutter future reviews.
  4. Safety Defaults: Files are never deleted. Active browser downloads are protected by cooldown timers, and name collisions are resolved automatically using timestamped suffixes.

Code

Briefly is open-source under the MIT license:

Briefly

A local-first filing assistant for legal work. Briefly watches a folder, reads supported documents on your machine, asks an open-weight model to identify the matter and document type, and files only clear matches. Everything uncertain stays where it is for review.

Weekend MVP for DEV's Build for a Friend challenge. Built around a lawyer's real habit: download now, file later.

What works in this MVP

  • Local browser dashboard served by Python; no account, cloud service, or telemetry.
  • Configurable inbox and matter library, cooldown, file type allow/exclude lists, and confidence threshold.
  • PDF, DOCX, TXT, and Markdown text extraction. Scanned PDFs are recognized as unsupported without optional OCR; no silent guess based on an empty extraction.
  • Local Ollama chat API with Gemma (or another locally installed model), constrained JSON output.
  • High-confidence files are moved into an existing matter folder and document-type subfolder. Low-confidence and unassigned documents remain in the inbox and…

Quickstart (Running Locally)

Requirements: Python 3.10+ (Standard Library only — zero pip dependencies) and a local Ollama instance.

# 1. Pull the model locally
ollama pull gemma4:e2b-it-qat

# 2. Run Briefly
python3 briefly.py
Enter fullscreen mode Exit fullscreen mode

Access the dashboard at http://127.0.0.1:8765.

How I Built It

Briefly was architected around two core engineering principles: complete local execution and standard library minimalism.

1. Local Open-Weight Inference with Google Gemma

Inference is executed strictly against local Ollama (http://127.0.0.1:11434) using Google's open-weight Gemma 4 (gemma4:e2b-it-qat).

We specifically evaluated this quantized Gemma 4 model to match her everyday laptop specifications, achieving deterministic structured JSON output in approximately 2 seconds per document without excessive thermal throttling or memory pressure. Briefly enforces a strict JSON schema so classifications map deterministically to verified matters:

{
  "matter": "Marcus Sterling v Gopaul & Colfire",
  "document_type": "Pleadings",
  "confidence": 0.94,
  "reason": "High Court claim form identifying collision dispute between named parties."
}
Enter fullscreen mode Exit fullscreen mode

Briefly is also model-agnostic: practitioners with high-end desktop workstations or dedicated GPUs can specify larger local models (such as Gemma 4 27B or Llama 3.3) directly in the settings panel without modifying code.

2. Trap 1: The "Where Did My File Go?!" Race Condition

If an attorney downloads an affidavit to attach to an outgoing email, and a background watcher immediately moves it 30 seconds later, that automation disrupts active work.

  • Implementation: Background file watching is disabled by default (watch_enabled: false).
  • Cooldown Protection: When automated watching is enabled, files are protected by a configurable cooldown timer (cooldown_seconds). Newly downloaded files remain accessible in the inbox while actively in use.
  • Documents are only evaluated after the cooldown period elapses, or when the user initiates a manual scan.

3. Trap 2: The Parsing Quagmire

Legal files are rarely uniform. They include scanned raster PDFs without text layers, encrypted court bundles, nested .zip archives, and corrupt DOCX files.

  • Implemented a 50 MB file size ceiling to prevent memory exhaustion.
  • Added decompression size validation (info.file_size <= MAX_FILE_SIZE) on DOCX XML streams to protect against zip bombs.
  • Integrated pdftotext with password-protected PDF detection and empty text validation, routing unsupported files to the review queue without failing the pipeline.

4. Trap 3: Hallucinated Directory Trees

If an LLM invents subtle variations of a matter name across runs (Smith v. Jones, Smith_vs_Jones, Smith v Jones (2024)), the local directory structure quickly fragments.

  • Implemented a closed-world schema where the model selects only from existing, validated matter directories.
  • Designed canonical_matter_key normalization that strips punctuation, legal particles (v., vs., versus), and whitespace variations into a single canonical identifier before any filesystem interaction.

5. Crash-Proof Subprocess Isolation

To eliminate GUI driver conflicts (epoll_ctl errors and segmentation faults caused by running Tkinter within a multi-threaded web server), system folder pickers are executed in isolated child processes (osascript on macOS, zenity/kdialog on Linux).

6. Local SQLite Audit Trail

All classification events, confidence scores, and filing destinations are recorded in a local SQLite database (activity table), providing a complete, auditable operational log.

Why Does Open Innovation Matter?

In legal practice, attorney-client privilege and confidentiality are strict professional mandates, not operational preferences.

Transmitting unredacted court pleadings, financial affidavits, or trade secrets across commercial cloud AI APIs risks waiving privilege under professional conduct rules. Cloud AI terms often permit telemetry collection or model retraining, which constitutes an unacceptable disclosure liability for any legal practitioner.

Open innovation provides the only compliant architecture for this domain:

  • Zero Data Egress: Because models like Google Gemma are open-weight, all inference runs in local memory on the practitioner's machine (127.0.0.1). No document text or metadata ever leaves the device.
  • Operational Independence: Open-source systems insulate professional practices from proprietary subscription lock-in, forced cloud migrations, or sudden API deprecation.
  • Verifiable Auditability: Open code paired with local weights allows practitioners to verify exactly how confidential client data is handled.

For high-trust professions—law, medicine, and accounting—open-weight models are not merely an alternative; they represent the only ethical and compliant deployment model.

My Agent Session

This project was developed using a multi-agent pair-programming workflow:

  • Codex CLI: Used for initial project scaffolding, CLI argument handling, and standard library HTTP server structure.
  • Google Antigravity: Used for architectural planning, implementing the safety safeguards (cooldown timers, bounded parsing, canonical matter reconciliation), generating the 21-document synthetic legal test corpus, conducting accessibility design reviews, and performing defensive adversarial security audits.
  • Test-Driven Verification: Antigravity engineered and verified our automated test suite across 40 unit tests (tests/test_briefly.py), covering concurrency, zip bomb mitigation, host header rebinding protection, and subprocess isolation.

(Note: Raw session transcripts were audited locally and kept private to protect local system paths and environment configuration).

Prize Categories

  • Best Use of Gemma (Featured Category)

    Briefly utilizes Google's open-weight Gemma 4 (gemma4:e2b-it-qat) running locally via Ollama as its core intelligence layer. Gemma 4 performs document classification, matter assignment, and collective multi-document reasoning for first-pass discovery, ensuring complete privacy and compliance with legal confidentiality requirements.

  • Best Use of ElevenLabs

    The Hacktoberfest challenge announcement specifically invited competitors to use ElevenLabs for demo video voiceover narration. We used ElevenLabs to produce the voiceover for our technical walkthrough, generating clear, professional audio that effectively communicates the problem space, architectural safeguards, and live filing workflows.


Top comments (0)