Your Astro blog can serve a lot of readers without running much JavaScript. A crawler can read those same pages without running any at all.
That makes browser analytics an incomplete place to look for automated visitors. A request for an article can reach your hosting provider even if the client never loads your analytics script.
Let's put observation at the edge and check that it works with one request. This walkthrough is for a static Astro site on a custom domain already proxied through Cloudflare. You can keep your existing build and hosting setup.
We build WebDecoy. This guide uses its Cloudflare edge sensor and dashboard, and was prepared with AI assistance.
Start with where your pages are served
Astro prerenders pages by default. Your hosting provider can serve the resulting HTML without starting an Astro server for each visit. Adding code to an Astro component does not make that code run whenever a crawler downloads the built file. Astro's rendering documentation explains the distinction.
For this setup, requests take this path:
Visitor → Cloudflare route → WebDecoy edge sensor → cached page or origin
The sensor observes requests that reach its matching Worker route, including clients that do not execute page scripts. It does not require an Astro adapter or a change to server rendering.
If you use a provider's default hostname rather than a Cloudflare-proxied custom domain, this particular setup will not cover it. Also check for existing Workers on the intended route. Do not replace a Worker that serves your application to make room for a sensor.
Connect the site you want to observe
Create a WebDecoy account, then select the site you want to monitor.
- Open Integrations → Cloudflare.
- Choose Connect with Cloudflare and review the requested permissions.
- Open Edge Sensor and select the zone and route for your site.
- Check which WebDecoy site will receive the detections before deploying.
In Cloudflare's DNS → Records, confirm that your site's hostname is proxied, shown by the orange cloud. A DNS-only record can leave you with an installed Worker that never receives the requests you expect.
Use the hostname people actually visit. A route for example.com/* does not cover www.example.com/*. Cloudflare's route documentation explains how patterns are matched.
The managed installer can exclude asset paths such as /_astro/* to reduce Worker invocations. Check these exclusions if you already have Workers serving assets: a route assigned to no Worker can affect more than the sensor. Worker usage is billed through your Cloudflare account under its current plan.
The current managed setup supports one edge-sensor site per WebDecoy organization. The installation guide covers this and the permissions in detail.
Deploy, then prove a request arrives
Select Deploy edge sensor, then open Sensors in WebDecoy. Look for the Edge Worker entry and confirm reporting is enabled.
Sensing reports activity. Existing deny policies can still be enforced by the same Worker, so review any rules you already have if your goal is observation only.
Now send a request to your own site's real hostname. Replace the example URL with yours:
curl -i -A 'WebDecoy-Test/1.0' https://your-site.example/
This request does not execute JavaScript. It uses a reserved user agent that asks the sensor to report a labeled test detection.
Open Detections for the same site and find the Test record. That is your evidence that the request reached the sensor and the report reached WebDecoy. A successful page response or a successful deployment message alone does not establish both.
The test label is intentional. It is not a claim that an AI crawler visited your site, and you should not include it in a screenshot as real crawler activity.
If the record is missing
Check these before changing your Astro project:
- Is the hostname proxied in Cloudflare DNS?
- Does the Worker route match the hostname and path used by curl?
- Is Reporting on, and is the sensor build current?
- Are you looking at the WebDecoy site selected during deployment?
A request redirected from the bare domain to www may take a different route. Check the response headers and test the final hostname directly. An upstream security rule can also stop traffic before your sensor sees it.
Read the first real detections
After the test, let ordinary traffic arrive. Start with the requested URLs. Are clients reading articles, following pagination, or probing paths such as /.env?
Then look at the crawler classification and the evidence supporting it. A user agent claiming to be GPTBot is a claim made by the client. Keep that separate from any verified identity shown in a record.
Timing helps too. Several article requests spread across a day tell a different story from repeated probes within a few seconds. Neither pattern, on its own, tells you whether your content ended up in model training.
The detections view is not a complete access log. The sensor selects automated-looking traffic for reporting, and sampling or limits can affect real-traffic counts. Keep hosting or CDN logs if you need a total request count.
Once your labeled test appears, you can leave collection running and return to the dashboard to see what actually visits. You do not need to start blocking crawlers to make that useful.
Top comments (0)