DEV Community

Cover image for The 19-year visit that wasn't real
Zenovay
Zenovay

Posted on Originally published at zenovay.com

The 19-year visit that wasn't real

One record in our analytics said a visitor had spent 619,637,320 seconds on a page. That is about 19.6 years. The number came from an automated scanner, not a remarkably patient reader.

That scanner-generated duration was not evidence of a visit, but it could distort an average-time view. It shows why a number can be valid input without being a useful measurement.

A browser timer is a defined measurement, not a direct reading of human attention.

The outlier was data, but it was not evidence of attention

A tracker has to accept information from a visitor's browser. That browser is outside the analytics server's control. A normal page script can report a duration, but a test client can send a number that no real visit could have produced.

The scanner request reported the 19.6-year value in the time_on_page field. Before a browser-reported duration can inform an average, it needs a plausibility boundary and an outlier policy. Otherwise, the average describes the data received, not how long people spent reading.

One impossible duration rises from a quiet trace

The size of the number made this case obvious. Smaller errors are harder to notice. A five-hour "visit" might look plausible in a report even when it came from an abandoned tab, a stale clock, or a malformed client. That is why the boundary matters even if nobody ever sees another value measured in years.

An open tab is not an active session

Elapsed time is simple to calculate: subtract the start timestamp from the current timestamp. It is not the same as time spent looking at a page. Someone can open a tab, switch to another app, have lunch, and return. A wall-clock timer would happily count all of it.

Zenovay's current browser tracker treats that gap differently. Its heartbeat checks whether the document is visible and the window has focus. It does not accrue time while the page is hidden or unfocused. While visible, it also stops adding time after five minutes without input. When a hidden tab becomes visible again, it resets the timing reference so the time away is not added in one jump.

One time trace continues while the active trace pauses

These are implementation choices, not a claim that we can measure attention perfectly. A person might read a long article without touching the mouse. Another might interact with a page without absorbing anything. The tracker can distinguish some browser states; it cannot see intent or comprehension.

There is also a naming trap. In Zenovay today, time_on_page can represent accumulated active session time, not a clean per-page stopwatch. The label alone should not be treated as the definition of the metric.

Put a bound where browser data enters

Zenovay now bounds client-reported durations when they arrive. Values above four hours are capped, with an over-limit signal retained separately for investigation. This prevents a multi-year duration from being stored as such, but the cap alone does not make every average meaningful.

An otherwise valid pageview is not rejected solely because its reported duration is implausible. The timing value still needs careful interpretation: a cap is a guardrail, not evidence of genuine attention.

A sharp outlier settles into a calm bounded trace

Four hours is a defensive ceiling, not our definition of a normal visit. Even a bounded outlier can skew an average, and a new boundary does not revise past reports. Any duration summary still needs an explicit definition and an outlier check.

Ask what the metric is for

Average time can be a useful clue. If a page's value changes sharply, it is worth checking the underlying events, outliers, traffic sources, and the exact definition of the timer. But a higher average is not automatically better. It can mean sustained reading, a confusing workflow, or simply a measurement problem.

For a product page, the better question may be whether visitors reached the next step. For documentation, it may be whether they found an answer. For an analytics setup flow, it may be whether a real event arrived from the intended website. Duration can add context to those outcomes; it cannot replace them.

The practical lesson is to define what the browser records, bound what a client can report, and inspect unusual values before treating an average as a story about people.

Top comments (1)

Collapse
 
omyvnss profile image
Om Yaduvanshi •

the "cap it, don't reject it" choice is the one i'd underline twice. throwing away the whole pageview over a garbage duration field would have cost you a real visit, and cookieless setups can't afford to be that picky with signal. the naming trap bit hit home too, "time on page" meaning three different things across three tools is half the confusion in this space. curious how you landed on four hours, was there anything in your distribution pointing there or just a purely defensive ceiling?