DEV Community

henry
henry

Posted on

Machine Learning Interviews and How to Spot Data Leakage

A model achieves an impressive score, but its performance collapses after deployment. In a machine learning interview, one explanation to investigate is data leakage: information reached the training or evaluation process in a way that would not be available for the intended prediction.

Start with the prediction moment. For this exercise, a service predicts on Monday whether an account will cancel during the following 30 days. Every feature must be justified in relation to what the system knows on Monday. A feature can look harmless in a spreadsheet and still reveal the future.Ask when each feature becomes available

Recent activity measured before Monday may be eligible. A final cancellation reason recorded three weeks later is not. A customer status field can also be dangerous if historical training rows were assembled from the latest account table rather than a snapshot from the prediction time.

That problem is not solved by removing a column named cancelled. Information about the outcome can hide in support workflows, billing adjustments, or fields updated after the event. Ask how each feature was produced and whether the historical value matches what a production request would have seen.

Distinguish event time from ingestion time. An event that happened before Monday but only arrived on Wednesday might still be unavailable to Monday's live prediction. The relevant boundary is actual availability under the proposed serving process.

Make the split match the question

If deployment predicts future behavior, a chronological evaluation can help test that setting. If the same person appears in multiple rows, consider whether splitting related rows across training and evaluation creates an easier problem than the one you intend to solve.

There is no universal split that fits every task. Explain the generalization you care about: future periods, unseen people, new organizations, or another population. Then choose a split that approximates it and state what it leaves untested.

Fit preprocessing in the right place

Transformations learned from data can leak information too. A scaler, imputer, or feature-selection step fitted using evaluation data has already learned something from the examples intended to assess the model.

The scikit-learn common-pitfalls guide discusses this issue and the use of pipelines to keep fitted transformations tied to the training process. In cross-validation, learned preprocessing should be fitted within each training fold. A pipeline helps enforce that organization, but it cannot fix features that already contain future information.

Investigate a suspicious score

Ask for a simple baseline, a feature-availability audit, and a comparison using a split that matches deployment. Remove questionable features and observe how performance changes. A score drop is evidence worth examining, but it does not by itself identify which feature was invalid or prove the remaining model is ready.

Check the evaluation metric against the decision too. A churn model may inform a limited outreach budget, so ranking quality at a useful operating point can matter alongside aggregate metrics. Avoid making business claims from one score without defining how predictions will be used.

For general technical rehearsal, explore PhantomCodeAI's interview resources, then validate ML-specific advice against the modeling library and your data construction. A useful practice prompt asks you to audit feature timestamps and defend the split, rather than merely recite a definition of leakage.

Finish with a production question: can every feature be computed with the same meaning and timing when a real request arrives? That question connects offline evaluation to the system you actually intend to build.

Top comments (0)