Today I ran my first real batch test with 100 URLs.
Result: 61% success. Not 100%.
And I'm going to tell you why I'm still posting this on LinkedIn.
6 months ago, I started building VNML Agent Framework — a multi-agent AI system running 100% local on my GTX 3050 Ti laptop with 4GB VRAM. No fancy GPU. No cloud. No team.
This morning, I ran 100 URLs across 12 different categories through a complete SEO pipeline:
→ Scraping → Analysis → Writing → Reflection → Validation → Save
Results in 3 hours 6 minutes:
✅ 61 URLs succeeded (61%)
✅ 15 articles scored APPROVED (≥90/100)
✅ 45 articles NEEDED_REVISION (60-89/100)
❌ 39 URLs failed — mostly intentional edge cases (404, paywall, bot detection)
Adjusted success rate (excluding 6 intentional 404 test URLs): 65%.
But this isn't a brag post.
I'm posting this because the batch test exposed 3 bugs that 5-10 URL small tests would NEVER reveal:
🔴 Bug #1: One URL took 20 MINUTES at the Analysis step
→ Content was only 790 chars, but LLM got stuck for 1174 seconds
→ Would destroy entire SLA if unfixed
🔴 Bug #2: example.com — the simplest test page in the world — FAILED to crawl
→ Along with api.github.com (public JSON) also failing
→ Content-type handling bug, not network
🔴 Bug #3: Layer L4 (Coverage) fails ~60% of URLs
→ Threshold too high → most need rewrite → wastes 30% of batch time
3 lessons I learned:
1️⃣ Running 5-10 test URLs will NEVER expose hidden bugs.
You MUST run a batch ≥ 50 URLs before release.
2️⃣ Timeouts need a hard cap, not just dynamic.
When the model is "confused" by short content, dynamic timeout doesn't help.
3️⃣ Honest metrics beat cherry-picked demos.
61% success + root cause analysis > 95% success on 20 pretty URLs.
I'm 43. The Vietnamese IT market considers me "too old."
But I still write code every day. Still debug every bug. Still learn from every batch test.
And I believe: transparency with real numbers matters more than showing off pretty stats.
Full source code + detailed metrics report in the first comment 👇
What are you building with local LLMs? Share in the comments.
Top comments (1)
📊 Full metrics report (VI + EN):