I split crawler hits from my human sessions
On 2026-09-27, crawler hits distorted my web traffic counts. Bot rows were inflating human session numbers in traffic-compounder. The report listed 262 bot events and 206 human sessions in separate rows. The morning brief had reported a paused push. I updated that reporter and the morning brief so they state measured trends and do not make push-pause claims.
On 2026-09-27, several checks stayed red beside the work that shipped. Blog Daily Healer hit 1 timeout during publication. A rescue run recovered the page, but outcome coverage stayed red at 69/73. Four tasks failed in that run, with 0 uncovered. gitleaks scanned 0 of 5 repositories because a volume was not mounted. A green result on 0 repositories is not proof the system is safe. Then I patched traffic-compounder so the report lists bot rows and human sessions apart.
That graphic is the 2026-09-26 note. It also warns that test runs can leak files. It does not show the 262 bot rows or the 206 human sessions from 2026-09-27.
Why did bot traffic distort the session count?
Automated crawlers hit the site and created 262 raw event rows on 2026-09-27. The old code counted those pings as real human visits. Separating the two streams revealed 206 actual human sessions. You must split bot noise at ingest or your analytics will lie.
The 2026-09-27 report listed those two counts in separate rows. I do not know whether they previously shared one table. The fact I can stand on is the split, not a guessed schema. Here is what the counts showed on 2026-09-27:
| Traffic source | Raw count on 2026-09-27 | Action taken |
|---|---|---|
| Bot crawler events | 262 | Listed in a separate row |
| Real human sessions | 206 | Listed in a separate row |
Those 206 sessions represent activity, not product retention. I still do not know if those visitors returned for a 2nd run. You should verify what an agent actually produced before relying on automated summaries. If your data collector bundles 1 bot with 1 user, your growth metric is fiction.
Why rewrite tool docs around user problems?
Internal field names make sense to the developer who wrote the database schema. They tell a reader nothing about what the code fixes. On 2026-09-27, I rewrote the release copy for showwork 0.6.5 to state user problems first. Good documentation explains what breaks when you do not run the tool.
I asked what problem showwork solved. The early draft answered with internal record names. I could not tell what those words meant for an external developer. I rejected that draft and asked for 1 concrete example. At 16:19 on 2026-09-27, I approved the revised copy and release steps. Then I published version 0.6.5 to PyPI.
You can install the updated release with 1 terminal command:
pip install showwork==0.6.5
When building agent pipelines, give an agent a file, not a memory to stop 1 common failure. If you use BMD, see what your agents actually did across 1 run. Clear words in 1 package release protect users from guessing.
How did background workers handle failed checks?
Brain Worker left 2 checks failed on 2026-09-27. The logs kept those limits visible. You must record unverified runs honestly instead of assuming green results.
The nightly sweep on 2026-09-26 merged 2 PRs and opened 7. Brain worker triage found 4 agent tasks and 2 auto-tonight cards. Reports/Weekly/synthesis-2026-39.md shipped as scheduled on 2026-09-27. Reports/Brain/prompt-evolver-2026-09-27.md logged 0 suggested edits and 0 patterns. The outcome-coverage metric remained red at 69/73 with 4 failed tasks. On 2026-09-27, one P0 credential Request was due to expire, and two exposed credentials still needed rotation. The devlog does not say that rotation finished.
The blog needed repair before it reached readers on 2026-09-27. Blog Daily Healer recorded 1 timeout before frontier rescue recovered publication. The final receipt named the article, its publication time, and a matching page response. I can keep that 1 result without turning the whole day green.
That same report says the AgentGuard reporting fix passed 17 focused tests, while the full follow-up suite remained pending. The security scanner still called 0 repository coverage clean. My record has to hold those 2 limits beside the release. Clear words and honest counts are part of the product I am trying to build.
What should you do with this?
You can protect your project by auditing incoming logs and verifying true agent output. Check whether bot rows inflate your session counts before reporting active usage. Rewrite your documentation around 1 real user failure instead of internal database columns. Keep your failed test counts visible so your dashboards stay honest.
- Inspect your access logs for automated bot traffic. Look for user-agent strings that spike your row count by 100 events or more. Separate those crawlers into an isolated table before computing daily active sessions.
- Review your latest tool release notes in
README.md. Replace 3 internal parameter names with the exact problem your user faces. If a user cannot tell what broke, rewrite the explanation. - Audit your automated health checks against 1 real repository. If a scanner reports a clean score on 0 scanned repositories, mark the run red. Never accept a passing status code without proof of work.
Accompanying prompt
What the prompt does: This prompt audits incoming web logs and separates crawler events from real human sessions.
Copy/paste this prompt:
Role:
Data Pipeline Engineer
Context:
Web traffic counters often mix automated crawler pings with real human visits, inflating session numbers.
You need to separate bot events from user sessions to produce accurate activity metrics.
Inputs:
- Log file path: __
- Bot user-agent patterns: __
- Session timeout in minutes: __
- Output summary path: __
Task:
1. Read the raw web access events from the specified Log file path.
2. Match every entry against the provided Bot user-agent patterns.
3. Label matching crawler rows as bot events and route them to an isolated table.
4. Group the remaining non-bot requests by visitor ID using the Session timeout in minutes.
5. Write the verified counts of bot events and human sessions to the Output summary path.
Output:
- A Markdown table showing bot event counts versus human session counts.
- A list of the top 5 detected crawler user-agents.
Constraints:
- Do not count any crawler hit toward human session totals.
- Output clean text without external dependencies or unexplained assumptions.
Copy the block above.
See what your agents did: https://bmdpat.com/bmd
Get the local AI lab notes (benchmark rows, VRAM fit, quant choices, what runs on consumer GPUs), M-F only when there is something worth sending: https://bmdpat.com/newsletter?utm_source=blog_md&utm_medium=aeo&utm_campaign=i-split-crawler-hits-from-my-human-sessions-2026
Originally published on bmdpat.com. I run a one-person AI agent company and write about what actually works.
Want these in your inbox? Subscribe to the newsletter - no spam, unsubscribe anytime.

Top comments (0)