DEV Community

Cover image for From Raw Text to Cryptographic Seal: Building a Legal Document Factory in Python
SunverseAI
SunverseAI

Posted on

From Raw Text to Cryptographic Seal: Building a Legal Document Factory in Python

When people think of Artificial Intelligence, they usually think of chat boxes. You type a prompt, text scrolls across the screen, and you copy-paste it.

In the legal world, a chat box isn't enough. A contract on a screen is just a suggestion. A contract in hand—signed, sealed, and cryptographically verified—is a binding asset.

As we build Lawyie (Sunverse AI’s intelligent legal infrastructure for Africa), one of our core mandates was moving beyond the chat interface. We needed a Document Factory.

Here is the engineering breakdown of how we built an in-memory PDF generation pipeline that creates cryptographically-sealed legal documents in Python.


1. The Problem with Standard File Writing
In standard Python web apps, saving a file usually means writing it to the local hard drive and then serving it.

In a cloud environment like Streamlit Cloud, doing this at scale causes concurrency issues (multiple users overwriting the same contract.pdf file) and unnecessary disk read/write latency.

The Solution: Everything must happen in-memory.


2. The In-Memory Buffer (io.BytesIO / Byte-Streams)
Instead of saving a file to the disk, we use Python’s io module to capture the PDF output directly as a byte-stream and feed it straight into the user's browser download button.

Here is how the pipeline works using fpdf2:

from fpdf import FPDF
import io

def generate_legal_pdf(contract_text, signature_id):
    # 1. Initialize the PDF engine
    pdf = FPDF()
    pdf.add_page()
    pdf.set_font("Arial", size=11)

    # 2. Clean text (Handling special characters for Latin-1 encoding)
    clean_text = contract_text.replace("", "NGN").replace("", "-")
    final_content = f"{clean_text}\n\nSECURE HASH ID: {signature_id}"

    # 3. Write to the document
    pdf.multi_cell(0, 10, txt=final_content)

    # 4. Capture the output as bytes (Crucial for fpdf2)
    pdf_output = pdf.output()
    pdf_bytes = bytes(pdf_output) if isinstance(pdf_output, bytearray) else pdf_output

    return pdf_bytes
Enter fullscreen mode Exit fullscreen mode

3. Cryptographic E-Signatures (hashlib)
In emerging markets, document tampering is a major risk. How does a user know the PDF they downloaded wasn't altered?

We solved this by generating a unique SHA-256 Hash ID tied to the user's name and the exact timestamp of generation.

import hashlib
from datetime import datetime

def generate_e_signature(name):
    timestamp = datetime.now().strftime("%Y%m%d%H%M%S")
    # Generate a secure 12-character cryptographic hash
    raw_string = f"{name}{timestamp}"
    sig_hash = hashlib.sha256(raw_string.encode()).hexdigest()[:12].upper()

    return f"SIGNED-BY-{name.upper()}-ID-{sig_hash}"
Enter fullscreen mode Exit fullscreen mode

This hash acts as a digital fingerprint. If even a single comma in the contract changes, the hash changes, proving authenticity.


4. Why This Matters for African Legal-Tech
By combining LLM inference (Groq) with an automated document factory (Python + FPDF2), Lawyie reduces the time it takes to draft, review, and seal a compliant SME contract from 3 days to 5 seconds.

For the 1.4 billion people of Africa, this isn't just about writing cleaner code. It’s about removing the economic barriers that keep millions operating in the "legal shadow."

What’s Next?
We are continuing to scale Lawyie from Abuja, optimizing our Supabase vault, and expanding our multi-language support.

If you're building document automation tools in Python, let's connect in the comments!


Try Lawyie Live: [lawyie.streamlit.app]

python #ai #pdf #opensource #buildinginpublic #africa

Top comments (1)

Collapse
 
sunverseai profile image
SunverseAI

One of the specific 'hidden' challenges I ran into while building this was Character Encoding. Standard PDF libraries like FPDF default to Latin-1, which means the Naira symbol (₦) and even certain smart quotes from mobile keyboards cause immediate crashes.
In the code, I implemented a string-sanitization layer to swap these for compatible characters before the byte-stream is generated. It’s a small detail, but it’s the difference between a functional product and a broken app for my users in Nigeria.
Are any other devs here building document engines? How are you handling non-standard currency symbols in PDFs?