Your offline GNN looks perfect. The AUC is high. The precision is great.
You ship it to production to catch fraud or optimize a supply chain.
Then the metrics tank.
The model isn't broken. The data is just too old.
It is making real-time decisions using neighborhoods and embeddings from yesterday, or last week. In fraud, a node's connections change in milliseconds. In supply chain, they change in hours.
The model is evaluating a ghost.
GNNs don't just need features. They need fresh ones. They need neighborhoods that reflect the current state of the world.
This means your graph storage and feature-update pipelines cannot be an afterthought. They must be engineered to the exact freshness the use case demands.
A batch job running at midnight is fine for predicting customer churn. It is catastrophic for blocking a fraudulent transaction at checkout.
Data engineering must run on the same clock as the business decision.
Don't trust your architecture diagram. Trust a trace.
Pick one production decision and trace it backward. Start at the served model input.
Look at the embedding it consumed. Look at the neighborhood it aggregated.
Trace those back to the graph storage layer. Trace that back to the feature-update pipeline. Trace that back to the source system change.
If a user changes their email, or a supplier misses a shipment, how long does it take for that reality to reach the model?
If you can't measure the lag, you can't trust the prediction.
Offline evaluation assumes the model has access to the exact state of the graph at the moment of the event. Production rarely offers that luxury.
Top comments (0)