I was monitoring the Vercel dashboard for a Next.js clinical AI platform I built at Synapsis Medical Technologies when the Interaction to Next Paint (INP) metric for our patient portal spiked from 180ms to 450ms. In a HIPAA-compliant environment where clinicians rely on real-time RAG (Retrieval-Augmented Generation) responses, a 300ms delay isn't just a "yellow" warning in a report—it is the difference between an interface feeling responsive and a user double-clicking a button, potentially triggering duplicate API calls to an LLM pipeline.
I ran Lighthouse. The score was a perfect 100. I ran a manual Trace in Chrome DevTools on my M3 Max MacBook; the main thread was clear. Yet, the field data—the real-world experience of patients on three-year-old Android devices and spotty hospital Wi-Fi—showed a catastrophic regression.
We spent two days chasing phantom layout shifts before identifying that the regression wasn't in our code, but in how we were measuring it. If you rely on lab data (Lighthouse/Synthetic) to debug production regressions, you are looking at a curated, sterile version of your app that does not exist for your users.
The Gap Between Lab and Field
Lab data is a snapshot in a controlled environment. Field data (Chrome User Experience Report or CrUX) is the messy reality. At Synapsis, I oversaw the architecture of five production systems, and the most expensive lesson I learned was that a "Performance" score in a CI/CD pipeline is a vanity metric.
Field regressions usually stem from three causes that lab tests cannot simulate:
- Device Heterogeneity: Your CI runner has a fast CPU; your user is on a $200 Motorola.
- Network Jitter: Lab tests use simulated throttling, which is predictable. Real-world 4G has packet loss that wreaks havoc on hydration.
- User Interaction: Lighthouse doesn't click buttons. INP, which replaced FID (First Input Delay), only exists when a user interacts.
Step 1: Isolate the Population, Not the Page
When a regression hits, do not look at the aggregate score. A 200ms jump in Largest Contentful Paint (LCP) across your entire site is often a false signal caused by a change in traffic mix (e.g., a new marketing campaign hitting lower-end devices).
What to do: Open your Real User Monitoring (RUM) tool—whether that is Vercel Speed Insights, Datadog RUM, or a custom web-vitals library implementation. Filter the regression by device type and connection type.
How to confirm: If the LCP regression only appears on "4G" and "Mobile" but "Desktop" remains flat, your issue is likely image encoding or a bloated JavaScript bundle that is choking lower-end CPUs during decompression, not a server-side logic error.
Step 2: Identify the "Attribution" Element
The biggest mistake I see leads make is guessing which element caused a Cumulative Layout Shift (CLS). The documentation tells you what CLS is; it doesn't tell you that the "offending" element in the console is often the victim, not the cause.
What to do: Use the web-vitals JavaScript library to log the attribution object to your analytics provider.
import { onCLS } from 'web-vitals/attribution';
onCLS((metric) => {
console.log(metric.attribution.largestShiftTarget);
// Send this string to your logging service
});
How to confirm: Look for the largestShiftTarget in your logs. In one instance, our "regression" was actually a third-party cookie consent banner that loaded 2 seconds late. Lighthouse missed it because the headless browser didn't trigger the consent flow.
Step 3: Debugging INP with Long Animation Frames (LoAF)
If your INP has spiked, the standard "Performance" tab in DevTools is often too noisy. You need to see what is blocking the main thread at the exact moment of interaction.
What to do: Enable the "Long Animation Frames API" in your testing. If you are on Chrome 123 or higher, you can use the Long Animation Frames (LoAF) entry to see exactly which script contributed to the delay.
- Open DevTools -> Performance.
- Click the "Interactions" track.
- Look for the red bars.
- Identify the "Presentation Delay."
How to confirm: If the "Input Delay" is high, your main thread was busy when the user clicked. If the "Processing Duration" is high, your event listener code is slow. If "Presentation Delay" is high, the browser was struggling to paint the frame (often due to excessive DOM size or complex CSS).
The Cost of Over-Optimisation
At Synapsis, I oversaw a CI/CD overhaul that cut release cycles from 2 days to 4 hours. Part of that was removing "Performance Gatekeeping" in CI that relied on Lighthouse scores.
Why? Because forcing an engineer to fix a 2-point Lighthouse drop in a PR often leads to "hacks" like delaying the load of essential scripts (like analytics or support chats) until after the first interaction. This improves the lab score but destroys the field INP, as the main thread suddenly hits a wall the moment the user tries to use the app.
The trade-off is simple: Lab data is for catching catastrophic failures before they merge; Field data is for making business decisions. Never roll back a release based solely on a Lighthouse score if your RUM data remains stable.
At Your Level
Starting out
Stop using "Fast 3G" throttling in DevTools as a proxy for reality. It is a linear slowdown. Instead, use the "Network" tab to simulate offline and slow speeds, but focus on understanding the "Main Thread" work in the Performance tab. Your goal is to see how your code execution blocks the browser.
Working engineer
Implement the web-vitals library with attribution. You cannot fix what you cannot name. When a ticket comes in for "slowness," your first response should be to pull the largestShiftTarget or the processingDuration for that specific user session.
Senior or staff
Own the RUM strategy. You should be looking at the 75th and 95th percentiles (P75/P95). A regression in P95 often signals a memory leak or a specific edge case (like a very large FHIR resource loading into a state manager), while a regression in P75 signals a general architectural bloat.
Lead or director
Distinguish between "Performance" as a technical metric and "User Experience" as a business one. A 500ms LCP increase is acceptable if it comes from high-resolution product images that increase conversion by 10%. Do not let your team spend 40 hours chasing a green score if the business metric is moving in the right direction.
In the Interview
The Question: "Our LCP has regressed by 1 second in production, but Lighthouse shows no change. How do you find the cause?"
The Weak Answer: "I would check the images, minify the CSS, and maybe try to implement a CDN or use next/image." This is weak because it jumps to solutions without diagnosing why the tools are giving conflicting signals.
The Strong Answer: A strong candidate identifies that Lighthouse is a synthetic environment and the regression is likely environment-dependent. They will talk about Field vs. Lab data. They will mention checking CrUX reports or RUM data to segment by device and geography. They will specifically name Interaction to Next Paint (INP) as a metric that Lighthouse cannot accurately measure.
The Senior Follow-up: "How do you handle a third-party script that is tanking your INP but is required by the marketing team?"
This separates those who have read docs from those who have shipped. A senior engineer knows you can't always delete the script. They will discuss requestIdleCallback, using Partytown to move scripts to a Web Worker, or strategically yielding to the main thread using setTimeout(..., 0) to break up long tasks. They understand the organizational friction between performance and third-party requirements.
Amit Chakraborty is a founding engineer and senior architect — React Native, AI/RAG systems and production architecture. Portfolio: www.amitchakraborty.dev · LinkedIn · GitHub. Open to senior and founding engineering roles, remote worldwide.
Top comments (0)