I built a small, reproducible workflow for forecasting retail demand when promotions matter. It compares seasonal estimates against per-store, per-product-family promotion models, keeps zero-sales days, and validates every output against the expected IDs.
Repo: https://github.com/officialpm/retail-demand-forecasting
This is a research workflow, not a deployed inventory system. No service, dashboard or production integration is claimed.
The approach
- Predict each store-family series on its own. Zero-sales days stay in.
- Compare three seasonal baselines, a longer-window seasonal control, and two fixed promotion models (56-day and 112-day).
- Fit weekday intercepts and log promotion counts to log sales, with fixed ridge penalties instead of a parameter search.
- Evaluate five nonoverlapping 16-day horizons, using only earlier history for each one.
- Join forecasts to the sample IDs one-to-one. Reject missing, duplicate, nonfinite or negative output.
Results
| Measurement | Result |
|---|---|
| Held-out test RMSLE (lower is better) | 0.40508 |
| Previous held-out test result | 0.41602 |
| Five-window development RMSLE, 56-day promotion model | 0.435361 |
| Same-window seasonal control | 0.460723 |
| Windows won against the seasonal control | 5 of 5 |
Daily sales summed across all stores and product families, zero-sales days included. The scale and variability change over time, so a single average is not a stable demand model. This plot is descriptive, not a causal promotion analysis.
Limits
- The development windows were reused as the method evolved, so those numbers are selection-biased.
- The held-out test score is one measurement on one test period. It says nothing about new stores, years or demand patterns.
- Promotions are associated with sales here. The coefficient is not a causal estimate.
- No holidays, oil prices, transactions or stockout labels are modeled.
- The portable
train.pyis an adaptation of the measured implementation. I ran synthetic tests (output shape, and invariance to future target changes), not a full retrain. - Nothing about production latency or business impact has been measured.
The repo has the code, tests, a full explained notebook and a results ledger with provenance. Feedback on the validation setup is welcome.
Code is Apache-2.0. The dataset is not included and is not covered by that license.
Top comments (0)