While building my AI culinary platform, Ages & Spices, I was setting up my marketing flows when I stumbled across a terrifying piece of US law: sending a single marketing email without a valid physical address or PO Box can trigger an FTC fine of up to $53,088 per email.
That moment of panic sent me down a massive rabbit hole of digital compliance, from GDPR to "Dark Patterns" (like the Roach Motel, hidden fees, and fake urgency). I realized how incredibly easy it is for solo developers and startups to accidentally break the law just by using standard growth hacks.
From that single realization, my flagship project was born. To solve this for myself and others, I built DarkLens, an automated AI compliance and dark pattern scanner for e-commerce and SaaS companies.
The Architecture
I built the platform using Next.js and Supabase. The core engine uses a custom web crawler that extracts the DOM structure of a target URL and feeds the semantic HTML into an LLM heuristic engine. The AI analyzes the UI against a database of known dark patterns and compliance risks, generating an auditor-ready report.
The hardest engineering challenge so far has been cleaning the DOM and stripping out unnecessary scripts/styles before passing it to the LLM. Doing this properly saves on token limits and significantly reduces AI hallucinations.
Iād love to hear from other developers who are building web-crawling or LLM-based analysis tools. How are you handling DOM parsing before sending it to your AI models?
You can check out the live scanner here: DarkLens
Top comments (0)