DEV Community

Hoang Nguyen
Hoang Nguyen

Posted on

Running Our First 100 URLs — Real Metrics From VNML Agent Framework

Today I ran my first real batch test with 100 URLs.
Result: 61% success. Not 100%.

And I'm going to tell you why I'm still posting this on LinkedIn.

6 months ago, I started building VNML Agent Framework — a multi-agent AI system running 100% local on my GTX 3050 Ti laptop with 4GB VRAM. No fancy GPU. No cloud. No team.
This morning, I ran 100 URLs across 12 different categories through a complete SEO pipeline:
→ Scraping → Analysis → Writing → Reflection → Validation → Save
Results in 3 hours 6 minutes:
✅ 61 URLs succeeded (61%)
✅ 15 articles scored APPROVED (≥90/100)
✅ 45 articles NEEDED_REVISION (60-89/100)
❌ 39 URLs failed — mostly intentional edge cases (404, paywall, bot detection)

Adjusted success rate (excluding 6 intentional 404 test URLs): 65%.

But this isn't a brag post.
I'm posting this because the batch test exposed 3 bugs that 5-10 URL small tests would NEVER reveal:
🔴 Bug #1: One URL took 20 MINUTES at the Analysis step
→ Content was only 790 chars, but LLM got stuck for 1174 seconds
→ Would destroy entire SLA if unfixed
🔴 Bug #2: example.com — the simplest test page in the world — FAILED to crawl
→ Along with api.github.com (public JSON) also failing
→ Content-type handling bug, not network
🔴 Bug #3: Layer L4 (Coverage) fails ~60% of URLs

→ Threshold too high → most need rewrite → wastes 30% of batch time

3 lessons I learned:
1️⃣ Running 5-10 test URLs will NEVER expose hidden bugs.
You MUST run a batch ≥ 50 URLs before release.
2️⃣ Timeouts need a hard cap, not just dynamic.
When the model is "confused" by short content, dynamic timeout doesn't help.
3️⃣ Honest metrics beat cherry-picked demos.

61% success + root cause analysis > 95% success on 20 pretty URLs.

I'm 43. The Vietnamese IT market considers me "too old."
But I still write code every day. Still debug every bug. Still learn from every batch test.

And I believe: transparency with real numbers matters more than showing off pretty stats.

Full source code + detailed metrics report in the first comment 👇
What are you building with local LLMs? Share in the comments.

Top comments (1)

Collapse
 
hoang_nguyen_ad2c9fbdd6db profile image
Hoang Nguyen •

📊 Full metrics report (VI + EN):

  • VI: vnmlstudio.runasp.net/vi/blogs/100-urls-dau-tien-metrics-thuc-te.html
  • EN: vnmlstudio.runasp.net/en/blogs/first-100-urls-real-metrics.html 💻 Source code: github.com/hoangnc/VNMLAgentFramework Curious to hear: what's the biggest bug your first batch test ever revealed? Drop it below 👇