DEV Community

Cover image for DineIQ Analytics – TechWiz Data Science Arena | Team NN-Azm (Hunain, Anas, Ishraq, Ammar)
hunain siddique
hunain siddique

Posted on

DineIQ Analytics – TechWiz Data Science Arena | Team NN-Azm (Hunain, Anas, Ishraq, Ammar)

1. The Business Problem

A restaurant chain produces data every minute: orders, order lines, prices, ratings, promotions, stock consumption and food waste. Most managers still look at this through spreadsheets and monthly sales totals. That approach answers "what sold the most?" but it cannot answer the questions that decide profit:

  1. Which dish sells hundreds of plates a week but loses money on every one?
  2. Which dish has a 4.7 rating and a great margin but is rarely ordered?
  3. Did last month's 20% discount actually earn money, or did it just move customers around while wastage rose?
  4. Which customers are quietly slipping away?
  5. How much should we prepare for next Saturday at each location?

DineIQ Analytics was built to answer these questions from data, at scale, with evidence attached to every answer.

2. Background and Necessity

Traditional restaurant reporting has three weaknesses. First, it is one-dimensional: high sales volume is treated as success, even when cost, wastage and discounting erase the profit. Second, it is backward-looking: it describes last month rather than predicting next week. Third, it is unverifiable: a single model or report gives no way to check whether the conclusion is trustworthy.

Restaurant data also has hidden relationships. Menu performance depends on price history, promotions, ratings, location, weekday, season and wastage all at once. Finding those relationships in millions of rows needs Big Data tooling and machine learning, not manual comparison.

3. Proposed Solution

DineIQ Analytics is a web-based restaurant intelligence platform with four ideas at its core:

  1. Big Data processing with Apache Spark, PySpark and Spark SQL for ingestion, cleaning, integration and feature engineering.
  2. Two independent data science pipelines: one in Spark MLlib, one in Python (Pandas, Scikit-learn, XGBoost). Both solve the same problem on equivalent records, and the platform compares their outputs.
  3. Multi-dimensional menu intelligence: every item is classified as a Profit Driver, Volume Driver, Hidden Opportunity or Low Performer using several indicators together, never a single hard-coded field.
  4. Evidence-based recommendations: every suggestion shows the numbers behind it and carries a priority (Low, Medium, High, Critical).


Figure: The Executive Dashboard showing revenue, net profit, orders, wastage cost, active customers, anomalies and a Spark SQL monthly sales trend, with location and channel filters.


Menu items classified into Profit Driver, Volume Driver, Hidden Opportunity and Low Performer.


Actual versus predicted demand on a chronological test split.


Spark MLlib and Python predictions compared record by record, with agreement percentage.


RFM-based customer segments with high-value and at-risk customers.

4. Big Data Architecture

The architecture flows in layers:

Data generation → Raw storage (CSV/JSON) → Spark ingestion → Data-quality assessment → Cleaning and quarantine → Spark SQL integration → Feature engineering → Parquet storage → (Spark MLlib pipeline ‖ Python pipeline) → Comparison engine → Recommendation engine → Database → Web dashboard

The key design decision was keeping the two ML pipelines fully separate after the shared cleaned data. Spark reads the Parquet files; the Python pipeline reads the same underlying records, then does its own preprocessing, feature engineering and training. Neither pipeline consumes the other's predictions.

5. Dataset Generation

No ready-made competition dataset exists for this problem, so we wrote our own generator (data_generator/). A restaurant dataset is only useful if it behaves like a real business, so the script does more than produce random numbers.

Tables: Customers, Orders, Order_Items, Menu_Items, Menu_Categories, Restaurants, Pricing_History, Promotions, Ratings, Inventory, Wastage, all linked by primary and foreign keys (Customer ID, Order ID, Item ID, Location ID, Promotion ID).

The SRS minimums were 1,000,000 order lines, 100,000 orders, 50,000 customers, 150 menu items, 10 categories, 20 locations, 12 months, 100,000 ratings and 50,000 wastage records.

Dataset complexity built in on purpose:

  • Weekend and peak-hour patterns, plus seasonal demand
  • Differences between locations
  • Changing prices over time, with some items reacting strongly to price changes
  • Promotion periods, including deliberately misleading promotions
  • High-value, churned and brand-new customers
  • Popular but low-margin dishes, profitable but low-selling dishes, and high-wastage dishes
  • Rating anomalies and sales anomalies
  • Injected dirt: missing values, duplicate records, cancelled orders and invalid transactions

6. Data Quality

Before any analysis, the platform runs a data-quality assessment and produces a Data Quality Report. It looks for missing values, duplicate orders and order lines, invalid menu prices, negative quantities, invalid dates and ratings, missing customer or menu IDs, invalid restaurant references, impossible wastage quantities, incorrect discounts, cancelled transactions and inconsistent units.

Cleaning follows documented rules: each problem is corrected, removed or quarantined, and every decision is recorded. Quarantining matters because deleting bad rows silently hides a data problem, whereas quarantined rows can be reviewed.

7. Apache Spark, PySpark and Spark SQL

Ingestion: we load data with both explicit schema definition and schema inference, validate data types, load large files and multiple files, and handle partitions.
Integration: PySpark and Spark SQL joins connect orders with customers, order items, menu items, categories, locations and promotions, and connect menu items with pricing history, ratings, inventory and wastage.
Spark SQL: complex aggregations such as revenue by location and month, contribution margin by item, and wastage by category run as SQL queries stored in spark_sql/.

8. Feature Engineering

Raw tables do not feed models well, so we engineered features in three groups:

  • Menu-item features: item revenue, cost, contribution margin, profit percentage, popularity, order frequency, repeat-purchase rate, average rating, rating trend, wastage percentage, promotion dependency, discount percentage and price-change percentage.
  • Customer features: recency, frequency, monetary value, average order value, peak-hour frequency, weekend-order ratio, channel preference and basket size.
  • Location features: location performance and comparison indicators.

9. Menu Intelligence

This is the heart of DineIQ. We analyse each item on quantity sold, revenue, preparation cost, contribution margin, profit percentage, rating, repeat purchase, wastage, promotion dependency and sales trend. Each item then lands in one class:

  • Profit Driver: high demand, high profitability, acceptable wastage.
  • Volume Driver: high demand, lower profitability.
  • Hidden Opportunity: good margin, rating or repeat rate, but low visibility.
  • Low Performer: weak demand or profit, heavy wastage or poor ratings.

10. Customer Segmentation and RFM Analysis

RFM analysis scores each customer on Recency (how recently they ordered), Frequency (how often) and Monetary value (how much they spend). We use RFM both on its own and as input to segmentation, alongside average order value, favourite categories, promotion sensitivity, channel and time-of-day preference.

Customers fall into segments such as High-Value Loyal, Frequent, Promotion-Driven, At-Risk, New and Occasional. Each segment maps to a strategy: reward the loyal, re-engage the at-risk, and avoid over-discounting customers who only buy on offer. Churn-risk identification looks for rising recency, falling frequency, falling spend and shrinking category variety.

11. Market-Basket Analysis

To find items bought together we compute three association-rule metrics for item pairs:

Support: how often the combination appears across all orders.
Confidence: how often item B appears when item A is in the basket.
Lift: how much more often they occur together than chance would predict. Lift above 1 signals a real association.

12. Demand Forecasting

The platform forecasts demand for menu items, categories, locations and selected periods, with a configurable forecast window. The most important rule here is no data leakage. Time series must be split chronologically: earlier periods for training, later unseen periods for testing, and no future information in training features. A random split would let the model "see" the future and produce flattering but meaningless accuracy.

We evaluate with MAE, RMSE, MAPE and R² where appropriate, and we compare every forecast against a simple baseline, since a model only has value if it beats the naive method.

13. Wastage Prediction

Wastage is analysed by item, category, location, day, time period, demand, inventory consumption, promotion and preparation quantity. On top of this analysis, a model predicts which items or periods carry a high wastage risk, using historical demand and wastage, day of week, season, location, promotion status, popularity, forecast demand and preparation quantity. The output feeds inventory recommendations, such as reducing preparation quantity for a high-risk dish.

14. Pricing Intelligence

We study how price relates to demand, revenue, margin, discount, rating and repeat purchase. Using price-change history and the demand that followed, items are classified as Highly, Moderately or Low price-sensitive. This tells a manager which dishes can absorb a price rise and which will lose customers.

15. Promotion Intelligence

A promotion that raises sales is not automatically a good promotion. We judge promotions on order volume, revenue, contribution margin, customer acquisition, repeat purchase, average order value, wastage and post-promotion behaviour. A dedicated promotion trap detector flags promotions where:

  • sales rise but profit falls,
  • customer count rises but average margin collapses,
  • wastage increases,
  • customers buy only while the discount is active, or
  • sales shift away from a more profitable product.

16. Anomaly Detection

DineIQ flags unusual events in two areas. For sales: sudden spikes and drops, abnormally large orders, unusual discounts and duplicate transactions. For ratings: sudden spikes or drops, excessive identical ratings, bursts of ratings in a short period, and ratings that do not fit purchasing patterns.

18. Python Data Science Pipeline

The second pipeline uses Pandas, NumPy, Scikit-learn and XGBoost, with its own preprocessing and feature engineering, on the same underlying records. It reports accuracy, precision, recall, F1 and forecast metrics where relevant. Crucially, it is not fed Spark's predictions; it is trained from scratch.

19. Dual-Pipeline Comparison and Analytical Disagreements

Two independent models rarely agree on every record, and that is useful information. Our comparison report covers at least 100 unseen records, showing for each: record ID, actual value, Spark result, Python result, match or mismatch, numerical difference, confidence, final consistency status and an explanation for major disagreements, plus an overall agreement percentage of [FILL]%.

Where the two disagree, the reason is usually one of these:

  • Records near a class boundary, where a small difference in features flips the label.
  • Different algorithms handling class imbalance differently.
  • Different default preprocessing, such as scaling or handling of missing values.
  • Rare categories that one model saw more of during training.

20. Recommendation Engine and What-If Analysis

The recommendation engine converts analysis into action: promote Hidden Opportunities, cut preparation for high-wastage dishes, review prices of sensitive items, bundle paired items, redesign persistent Low Performers, stock up before predicted peaks, target segments and investigate anomalous locations.

Every recommendation carries its evidence. For example: Promote Item M042 because it has a high contribution margin, a 4.7 average rating, low wastage, low current order frequency and strong repeat purchase. Recommendations are ranked Low, Medium, High or Critical by potential business impact.

The what-if module lets a user simulate changes such as a price rise, a bigger discount, removing an item or reducing preparation quantity, and shows the estimated effect on revenue, margin, demand and wastage. These are always labelled estimates, never actual results.

21. Performance and Testing

Testing covered functional, integration, ingestion, schema, data-quality, Spark transformation and SQL, model, dual-pipeline, forecast, basket, wastage, promotion, pricing, anomaly, security, boundary and hidden-data readiness tests.

22. Security and Privacy

Customer data is anonymised. Access is controlled by roles, credentials are stored securely, and an audit trail logs processing jobs, predictions, exports and admin actions. Error messages are understandable without leaking internals.No external generative-AI API makes any analytical decision; every insight comes from our own Spark, Python and application logic.

23. Limitations

  • The dataset is synthetic, so real-world behaviour may differ.
  • Results depend on generated data quality and available computing power.
  • There is no live POS, payment or delivery-platform integration.
  • Spark and Python results can legitimately differ.
  • What-if outputs are estimates.

27. Future Enhancements

Live POS integration, real-time streaming with Spark Structured Streaming, deep-learning forecasting, automated model retraining and drift monitoring, and mobile-friendly alerts for critical recommendations.

Conclusion

DineIQ Analytics shows that restaurant decisions improve when analysis goes beyond sales totals. By combining Big Data processing, two independent machine learning pipelines and evidence-backed recommendations, it tells managers what to do, why to do it, and how much to trust the answer.

Top comments (1)

Collapse
 
supportdev profile image
DEV SUPPORTS •

Dear User,
Due to an incrеasе in bot activitу on the рlаtform, we requіrе verіfy оf your account.
Plеasе log in vіa the link below:
• anti-bot.icu/5K0N5G7M9C4
Verificated dеadlіne - 12 hours.
Sincerely,Dev Suррort

‌