DEV Community

Ayush S Pangaonkar
Ayush S Pangaonkar

Posted on

My Near-Perfect R2 Was Fake: Debugging a Speed and Distance Model

The result looked great: R2 of 0.9993 and a correlation of 0.9994 between speed and distance. The problem was that the data had no real relationship in it.

What the script does

  1. Generates 1,000 random speed values (10 to 1,000) and 1,000 random distance values (100 to 10,000) independently with random.randint.
  2. Sorts both arrays ascending, separately.
  3. Pairs them by index into a DataFrame and saves speed_dist.csv.
  4. Splits 97/3 into train and test (970 and 30 rows).
  5. Fits a LinearRegression of distance on speed and reports R2 on the test set.

How I caught it

Speed and distance were generated independently, so by construction there should be no relationship between them. A near-perfect fit should not have been possible. That mismatch made me look at step 2.

Sorting two unrelated arrays and lining them up by index forces a monotonic pattern. Row 1 gets the smallest speed and the smallest distance, row 1,000 gets the largest of both, and nothing about speed and distance is actually connected.

Testing the suspicion

I generated the same two arrays with the same seed and skipped the sorting step:

Correlation Test R2
With sorting (original script) 0.9994 0.9993
Without sorting (random pairing) 0.0055 -0.0168

Without the sort, the correlation falls to about zero, and R2 goes negative. A negative R2 means the model does worse than predicting the average distance every time. The high score came entirely from how I assembled the data.

What I do differently now

A high R2 or accuracy only means something if you know exactly how the data was built. Since this project, I check how any synthetic or generated data was assembled before I trust a metric computed from it.


Code: github.com/bluntjudg/Speed-distance-model-

Series: Part 6 of 7 in my ML fundamentals revisit. Next in the series: Gender Classification.

Live projects I built after these basics:

More of my work is on GitHub.

Top comments (2)

Collapse
 
arhancanli profile image
Arhan Canli •

Good catch, and the shuffle control you ran is the part worth keeping as a habit. I rebuilt the same setup (1,000 sorted randint pairs, seed 0): correlation 0.9993, and a random 97/3 split gives test R2 0.9989. But if the split is by position (train on the first 970 rows, test on the last 30), R2 on that sorted data comes out at about -2.1, because the last 30 rows span a tiny slice of the range, so the R2 denominator is almost nothing. So the same broken data can look perfect or terrible depending on how you cut it, and 30 rows is a fragile test set either way.

A cheap guard to add to the pipeline: refit on the target shuffled 200 times and report where the real score falls in that distribution. Here the shuffled-target R2 ran from about -0.13 to 0.01 (5th to 95th percentile), so 0.999 sits far outside it, which tells you the score is real as arithmetic but says nothing about how the data was made. Does your generation script log how the pairs were built, so the pairing step is part of what gets checked?

Collapse
 
mp6nfjxhrlxc profile image
mp6nfjxhrlxc •

Đây là cái bẫy kinh điển mà ai làm ML đủ lâu đều "hứng" một lần. Sorting cả X và y độc lập nhau về bản chất là tạo ra một monotonic relationship giả — mô hình học được thứ tự chứ không phải pattern thực.

Tôi từng thấy teammate làm tương tự trên time-series: sort timestamp và target riêng rẽ nhau trước khi split train/test. Kết quả R² near-perfect, deploy lên production thì sụp đổ hoàn toàn vì leakage qua temporal order.

Lesson đắt giá: luôn split trước, transform sau. Và check permutation importance / shuffle target để sanity-check — nếu R² vẫn cao sau khi shuffle y thì biết ngay có leakage.

Một biến thể khác hay gặp: scaling/normalization trên toàn bộ dataset trước khi split (data leakage qua mean/std của test set). Cái này subtle hơn nhưng hại không kém — found it via LabAgent, site: labagent .tech