DEV Community

Daniel Pertu
Daniel Pertu

Posted on

Nine analytics defaults we turned off, and the one that has to be decided in the browser

An analytics SDK arrives with its defaults set to "collect everything". That is the right default for the vendor and almost never the right one for the page, because every option is spending one of three budgets: bytes downloaded, requests made, or rows stored. CogniPrep's posthog.init call is mostly a list of things switched off, and one decision that is only correct because it is made in the browser rather than in the vendor dashboard.

First, it is not their domain

The client points at our own origin:

posthog.init(key, {
  // Reverse proxy (see `rewrites` in next.config.mjs) so ad blockers do not
  // silently erase a chunk of the numbers.
  api_host: '/ingest',
  ui_host: host,
  ...
});
Enter fullscreen mode Exit fullscreen mode

api_host is where events go. ui_host stays the real PostHog host, because that is what the SDK uses when it needs to link a human back to the dashboard. Get that pair the wrong way round and your toolbar links point at your own marketing site.

The proxy itself is three rewrites, and the comment on them is load bearing:

async rewrites() {
  return [
    // PostHog reverse proxy - order matters!
    {
      source: '/ingest/static/:path*',
      destination: 'https://us-assets.i.posthog.com/static/:path*',
    },
    {
      source: '/ingest/decide',
      destination: 'https://us.i.posthog.com/decide',
    },
    {
      source: '/ingest/:path*',
      destination: 'https://us.i.posthog.com/:path*',
    },
  ];
}
Enter fullscreen mode Exit fullscreen mode

Next.js matches rewrites in order and stops at the first hit, so the catch-all has to be last. PostHog serves its static assets from a different host to its ingestion API, us-assets.i.posthog.com rather than us.i.posthog.com, so if the generic /ingest/:path* rule sits first it swallows /ingest/static/... and sends the library's own script requests to the ingestion endpoint. You do not get an error for that. You get an analytics client that never finishes loading, on some visits, depending on which files it needed.

This is the whole reason I am wary of "order matters!" comments that do not say which order or why. That one now does.

Sampling replay on the server samples the wrong thing

The one genuinely interesting decision in the file:

const SESSION_RECORDING_SAMPLE_RATE = 0.1;
...
const shouldRecord = Math.random() < SESSION_RECORDING_SAMPLE_RATE;
...
// `disable_session_recording: true` prevents the recorder chunk from being
// fetched at all, which is why the sampling decision is made above.
disable_session_recording: !shouldRecord,
Enter fullscreen mode Exit fullscreen mode

PostHog can sample session replay for you in project settings. That setting decides whether a recording is kept. The rrweb recorder is still downloaded, still initialised, and still runs on every visitor whose recording will be thrown away. It is the single largest analytics cost on the page.

Deciding in the browser, before the library is configured, means nine out of ten sessions never fetch the recorder at all. The recordings stay useful at that rate because of what they are for: watching how a flow actually goes, not auditing one named user. When a specific user's session matters, Sentry's on-error replay already covers it.

The general shape here is worth stealing even if you never touch PostHog. A server-side sampling rate is a storage optimisation. A client-side one is a performance optimisation. They have the same name and they are not the same thing.

The list of noes

// Anonymous visitors are counted but get no person record.
person_profiles: 'identified_only',

// Autocapture fired an event on every click and produced `$autocapture` rows
// keyed by DOM position that nobody ever built an insight on.
autocapture: false,
capture_heatmaps: false,

// Web vitals are the honest measure of the latency users feel, and cheap.
// Network timing records a payload per request for detail we never needed.
capture_performance: { web_vitals: true, network_timing: false },

// Errors belong in Sentry, which has the stacks, releases and grouping.
capture_exceptions: false,

// localStorage only. The default is 'localStorage+cookie'.
persistence: 'localStorage',
Enter fullscreen mode Exit fullscreen mode

Each of those has a sentence attached in the source, because "we turned it off" without a reason is how a setting gets turned back on in eighteen months by someone who assumes it was an accident.

Two are worth expanding.

Autocapture off costs you heatmaps. That is a real trade, not a free win, and it should be written down next to the flag rather than discovered later. The bet is that a small number of named events beats a large number of DOM-position events that nobody queries.

persistence: 'localStorage' is not a performance choice, it is a consistency one. The default is localStorage+cookie, which sets a cookie, and our published cookie policy says analytics identifiers live in local storage. A config default that quietly contradicts your own policy page is the kind of thing that is true for a year before anyone opens DevTools. localStorage alone still recognises a returning visitor, so nothing analytically useful is given up.

There is one default still on that I have not earned yet: the surveys script loads, and we run no surveys. That one is on the list.

None of it starts until the page is idle

The loader is an import() scheduled at the first idle moment, with a Safari fallback, because Safari still has no requestIdleCallback:

useEffect(() => {
  if (typeof window.requestIdleCallback === 'function') {
    const handle = window.requestIdleCallback(() => loadPostHog(), { timeout: 4000 });
    return () => window.cancelIdleCallback?.(handle);
  }
  const timer = setTimeout(loadPostHog, 1500);
  return () => clearTimeout(timer);
}, []);
Enter fullscreen mode Exit fullscreen mode

Nothing is lost by starting late. capture_pageview: 'history_change' records the initial view whenever init happens, and the thin client module queues anything fired before the library lands. The failure path is a .catch that does nothing on purpose: blocked, offline or a failed chunk all mean the queued events are never sent and the page carries on.

The component also renders no context provider. posthog-js/react exists for the usePostHog and feature flag hooks, this app uses neither, so importing it would be bundle weight in exchange for nothing.

See it

  • Open cogniprep.app with DevTools on the Network tab and filter for ingest. You will see config.js, web-vitals-with-attribution.js, surveys.js and an event POST, every one of them on cogniprep.app. Now filter for posthog.com: nothing. That is the reverse proxy working.
  • Note the timings. On my last load the first /ingest request started about 170ms after navigation, well after the page had painted.
  • Application, Local Storage, ph_<project key>_posthog. Then read cogniprep.app/cookies, which describes that same key in the local storage section. The policy and the config are meant to agree, and this is the check that they do.
  • Reload a handful of times and watch for a recorder chunk in the Network tab. Most loads will not have one, which is the client-side sample rate you are watching from the outside.

Top comments (0)