DEV Community

Cover image for Why Your Demand Forecast Is Wrong at the SKU Level
Emmanuel R for CobuildX AI

Posted on Originally published at cobuildx.ai

Why Your Demand Forecast Is Wrong at the SKU Level

ERP demand forecasts look reasonable at the product family level and fall apart at the SKU level — the level where production scheduling, material purchasing, and safety stock decisions actually get made.

Most manufacturing companies have a demand forecast. It lives in their ERP, runs on historical order data, and gives planners a number to work with for the weeks ahead. At the product family or aggregate level, it is usually close enough. At the SKU level — the level that actually drives production scheduling, material purchasing, and safety stock decisions — it is often systematically wrong in ways that quietly accumulate into stockouts, excess inventory, and expedite costs.

ERP forecasting engines — exponential smoothing, moving averages, seasonal index adjustment — were designed for stable, high-volume demand. They smooth over intermittency. They do not incorporate external signals. They treat every SKU the same regardless of how its demand actually behaves. A product that sells 0-0-0-0-50-0-0-0 units over eight weeks is not well-served by a model optimised for products that sell 5-6-5-7-6-5-6-5. Most product catalogues are dominated by the former.

Key insight: Aggregate forecast accuracy is a planning metric. SKU-level accuracy is an operational one. Most companies are tracking the wrong number and wondering why their inventory and service level performance is poor.

"The monthly fill rate report says 94%. The customer service team knows exactly which ten SKUs are causing every complaint."

How ERP Forecasting Works and Where It Breaks

Standard ERP demand forecasting is statistical. The system looks at historical order quantities, applies a smoothing algorithm, adjusts for detected seasonality, and produces a point estimate for future periods. This works when demand is reasonably stable and high-volume — the algorithm has enough signal to learn from and enough volume to average over.

It breaks down in three common situations. First, intermittent demand: products that sell sporadically, in irregular quantities, with long stretches of zeros between orders. Standard smoothing models interpolate across the zeros rather than recognising that zero is meaningful information about demand structure. Second, new products with no meaningful history — the model has nothing to learn from. Third, products with external demand drivers — weather, promotions, competitor stockouts, commodity prices — that do not appear in the historical order data the model trains on.

For most manufacturers, the majority of SKUs fall into at least one of these categories. The forecast is accurate for the products where planning was already easy, and unreliable for the ones where it matters most.

ERP forecasting is accurate where demand is already predictable. Intermittent, externally-driven, and new-product SKUs — where it consistently breaks — are usually a majority of the catalogue

Why SKU-Level Accuracy Is a Different Problem

Aggregate demand forecasting benefits from statistical averaging. The random errors in individual SKU forecasts partially cancel at the product family or business-unit level, producing a reasonable overall picture. This is why the monthly forecast review looks acceptable while the warehouse team is simultaneously managing excess stock and critical shortages on the same day.

The SKU-level errors do not cancel — they compound. An overforecast on SKU A ties up working capital in inventory that is not moving. An underforecast on SKU B creates a stockout, a missed shipment, a customer escalation, and an expedite purchase order at two to three times the standard cost.

The economics are asymmetric. The cost of a stockout on a high-value, fast-moving SKU typically outweighs the cost of the equivalent overstock by a factor of five or more. The aggregate accuracy number systematically understates the problem because it weights both outcomes equally.

Stockouts on critical SKUs cost far more than equivalent overstock on slow-movers — aggregate accuracy hides the asymmetry

What AI Forecasting Actually Adds

ML-based demand forecasting does several things that statistical ERP models cannot.

It handles intermittent demand explicitly. Methods like Croston's, temporal fusion transformers, or gradient boosting with intermittency encodings can learn the distinction between 'this SKU has lumpy demand' and 'this SKU is declining' — a distinction exponential smoothing blurs.

It ingests external signals. Promotional calendars. Pricing history. Weather for weather-sensitive categories. Macroeconomic indicators for industrial products. These signals are available and correlated with demand but standard ERP engines have no mechanism to incorporate them.

It learns across similar SKUs. A new product with no history can be forecasted using the demand patterns of comparable products — same category, same price tier, same channel. Transfer learning approaches move signal from SKUs with history to SKUs without.

Realistic improvement: 15–30% reduction in mean absolute error at the SKU level for organisations with clean historical data and relevant external signals. The upper end of that range requires both. The lower end is achievable with historical data alone.

15–30% mean absolute error reduction at SKU level is realistic — the upper end requires both clean history and external signals

The Data Problem Nobody Scopes Upfront

The forecast model is the straightforward part. The hard part is the data it learns from.

ML demand forecasting requires clean, consistent historical transaction data — typically two to three years of order-level history with reliable SKU identifiers, dates, and quantities. This sounds simple. In practice it rarely is.

Common issues: SKU identifiers that changed during an ERP migration, splitting history across two ID schemas. Order dates that reflect entry date in some records and ship date in others depending on who entered them. Returns and cancellations mixed into order history without flags. Promotional demand mixed with baseline with no promotional calendar to separate them. Stockout periods in historical data that look like zero demand when they were actually constrained supply.

Every one of these degrades the model. A well-built ML model trained on dirty data produces confident wrong answers — which is worse than a noisy statistical model that planners already know to treat with scepticism.

The right first question before scoping an AI forecasting project: can we pull three years of clean, consistent, order-level history for our top 500 SKUs? If the answer is not a clear yes, the first project is data quality, not model selection.

ML forecasting on dirty data produces confident wrong answers. The data audit belongs before the model architecture conversation

Getting It Into the Planner's Workflow

The second place these projects stall is workflow integration. A better forecast that lives in a separate tool — one planners have to log into, compare to the ERP number, and manually transfer back — is not a better forecast. It is additional work with the same decisions at the end.

For AI forecasting to change planning outcomes, the numbers need to land in the system where planning decisions are made. That means writing forecast quantities back into the ERP planning module, or a planning interface that replaces the ERP planning screen rather than sitting alongside it.

The override question matters too. Planners will override the AI forecast, and they should. They have information the model does not — a conversation about an upcoming large order, a product launch the system was not told about, a capacity constraint that affects what can actually be built. A good AI forecasting system surfaces its confidence, shows its key inputs, and makes overrides easy to record and feed back as training signal. A system that presents a number with no context and no clear override mechanism creates friction that kills adoption faster than any accuracy gap.

A better forecast number that lives outside the ERP is additional work for planners, not a capability improvement — workflow integration is not optional


Originally published on the CobuildX blog.

Top comments (0)