<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Parth Maniar</title>
    <description>The latest articles on DEV Community by Parth Maniar (@parth).</description>
    <link>https://dev.to/parth</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F546454%2Ffe502abd-35bd-444a-a04e-bf2b7ded628b.jpeg</url>
      <title>DEV Community: Parth Maniar</title>
      <link>https://dev.to/parth</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/parth"/>
    <language>en</language>
    <item>
      <title>Fine-tuning EfficientNetB0 for 104 flower classes: validation vs held-out test scores, honestly</title>
      <dc:creator>Parth Maniar</dc:creator>
      <pubDate>Thu, 08 Oct 2026 19:54:03 +0000</pubDate>
      <link>https://dev.to/parth/fine-tuning-efficientnetb0-for-104-flower-classes-validation-vs-held-out-test-scores-honestly-2l03</link>
      <guid>https://dev.to/parth/fine-tuning-efficientnetb0-for-104-flower-classes-validation-vs-held-out-test-scores-honestly-2l03</guid>
      <description>&lt;p&gt;I built a GPU transfer-learning workflow for 104-class flower recognition and tried to report it honestly, including the gap between validation and held-out test scores.&lt;/p&gt;

&lt;p&gt;Repo: &lt;a href="https://github.com/officialpm/flower-classification" rel="noopener noreferrer"&gt;https://github.com/officialpm/flower-classification&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fimages.weserv.nl%2F%3Furl%3Draw.githubusercontent.com%2Fofficialpm%2Fflower-classification%2Fmain%2Fdocs%2Fcover.svg%26output%3Dpng%26w%3D1000" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fimages.weserv.nl%2F%3Furl%3Draw.githubusercontent.com%2Fofficialpm%2Fflower-classification%2Fmain%2Fdocs%2Fcover.svg%26output%3Dpng%26w%3D1000" alt="104-class flower recognition: held-out test macro-F1 0.86708" width="1000" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This is an evaluated model workflow, not a deployed flower-identification product.&lt;/p&gt;

&lt;h2&gt;
  
  
  The approach
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fimages.weserv.nl%2F%3Furl%3Draw.githubusercontent.com%2Fofficialpm%2Fflower-classification%2Fmain%2Fdocs%2Fworkflow.svg%26output%3Dpng%26w%3D1000" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fimages.weserv.nl%2F%3Furl%3Draw.githubusercontent.com%2Fofficialpm%2Fflower-classification%2Fmain%2Fdocs%2Fworkflow.svg%26output%3Dpng%26w%3D1000" alt="Training workflow: load, train, select, export" width="1000" height="297"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Verify TFRecord feature names, counts and that all 104 classes appear in training and validation.&lt;/li&gt;
&lt;li&gt;Load an ImageNet-pretrained EfficientNetB0 at 224 pixels with flip, rotation and zoom augmentation.&lt;/li&gt;
&lt;li&gt;Train a frozen head for three epochs, then fine-tune at a lower learning rate for up to eight. BatchNorm layers stay frozen.&lt;/li&gt;
&lt;li&gt;Keep the best validation-loss checkpoints and report macro-F1, accuracy and log loss for every candidate.&lt;/li&gt;
&lt;li&gt;Match test IDs exactly to the sample order. Reject duplicates, missing IDs, nonfinite probabilities or labels outside 0-103.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The training set has 12,753 labeled images across 104 class IDs. All classes are present but counts are uneven, so I track macro-F1 rather than accuracy alone.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fimages.weserv.nl%2F%3Furl%3Draw.githubusercontent.com%2Fofficialpm%2Fflower-classification%2Fmain%2Fclass-distribution.svg%26output%3Dpng%26w%3D1000" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fimages.weserv.nl%2F%3Furl%3Draw.githubusercontent.com%2Fofficialpm%2Fflower-classification%2Fmain%2Fclass-distribution.svg%26output%3Dpng%26w%3D1000" alt="Training row counts across all 104 flower classes" width="1000" height="244"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Results
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Validation macro-F1&lt;/th&gt;
&lt;th&gt;Validation accuracy&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Frozen head&lt;/td&gt;
&lt;td&gt;0.800261&lt;/td&gt;
&lt;td&gt;0.823276&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fine-tuned&lt;/td&gt;
&lt;td&gt;0.855443&lt;/td&gt;
&lt;td&gt;0.878502&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Equal blend&lt;/td&gt;
&lt;td&gt;0.855260&lt;/td&gt;
&lt;td&gt;0.873653&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fimages.weserv.nl%2F%3Furl%3Draw.githubusercontent.com%2Fofficialpm%2Fflower-classification%2Fmain%2Fdocs%2Fval-macro-f1.svg%26output%3Dpng%26w%3D1000" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fimages.weserv.nl%2F%3Furl%3Draw.githubusercontent.com%2Fofficialpm%2Fflower-classification%2Fmain%2Fdocs%2Fval-macro-f1.svg%26output%3Dpng%26w%3D1000" alt="Validation macro-F1 by model" width="1000" height="444"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fimages.weserv.nl%2F%3Furl%3Draw.githubusercontent.com%2Fofficialpm%2Fflower-classification%2Fmain%2Fdocs%2Fval-accuracy.svg%26output%3Dpng%26w%3D1000" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fimages.weserv.nl%2F%3Furl%3Draw.githubusercontent.com%2Fofficialpm%2Fflower-classification%2Fmain%2Fdocs%2Fval-accuracy.svg%26output%3Dpng%26w%3D1000" alt="Validation accuracy by model" width="1000" height="444"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The fine-tuned model scored 0.86708 macro-F1 on the held-out test set, scored independently of my training and validation. That is a different measurement from validation (0.855443), and I don't read the gap as a real improvement.&lt;/p&gt;

&lt;h2&gt;
  
  
  Limits
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;The validation split was reused for checkpointing and model selection.&lt;/li&gt;
&lt;li&gt;Fine-tuned and blended results are nearly tied. I claim no meaningful margin between them.&lt;/li&gt;
&lt;li&gt;I did not purge duplicate images or test on an independent unseen domain.&lt;/li&gt;
&lt;li&gt;ImageNet pretraining may overlap with public flower imagery.&lt;/li&gt;
&lt;li&gt;This dataset's test labels are easy to find online. I did not fetch or use them, but it means near-perfect scores on it say little. Treat any score here with suspicion, including mine.&lt;/li&gt;
&lt;li&gt;The portable &lt;code&gt;train.py&lt;/code&gt; passed syntax checks but was not retrained end to end. The measured run used TensorFlow 2.20 on two Tesla T4 GPUs.&lt;/li&gt;
&lt;li&gt;No throughput, latency or accuracy on real user photos has been measured.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Code is Apache-2.0. No dataset, checkpoints or prediction files are published.&lt;/p&gt;

</description>
      <category>machinelearning</category>
      <category>python</category>
      <category>deeplearning</category>
    </item>
    <item>
      <title>Promotion-aware retail demand forecasting: what worked and what I can't claim</title>
      <dc:creator>Parth Maniar</dc:creator>
      <pubDate>Thu, 08 Oct 2026 19:53:33 +0000</pubDate>
      <link>https://dev.to/parth/promotion-aware-retail-demand-forecasting-what-worked-and-what-i-cant-claim-373i</link>
      <guid>https://dev.to/parth/promotion-aware-retail-demand-forecasting-what-worked-and-what-i-cant-claim-373i</guid>
      <description>&lt;p&gt;I built a small, reproducible workflow for forecasting retail demand when promotions matter. It compares seasonal estimates against per-store, per-product-family promotion models, keeps zero-sales days, and validates every output against the expected IDs.&lt;/p&gt;

&lt;p&gt;Repo: &lt;a href="https://github.com/officialpm/retail-demand-forecasting" rel="noopener noreferrer"&gt;https://github.com/officialpm/retail-demand-forecasting&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fimages.weserv.nl%2F%3Furl%3Draw.githubusercontent.com%2Fofficialpm%2Fretail-demand-forecasting%2Fmain%2Fdocs%2Fcover.svg%26output%3Dpng%26w%3D1000" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fimages.weserv.nl%2F%3Furl%3Draw.githubusercontent.com%2Fofficialpm%2Fretail-demand-forecasting%2Fmain%2Fdocs%2Fcover.svg%26output%3Dpng%26w%3D1000" alt="Promotion-aware demand forecasting: held-out test RMSLE 0.40508" width="1000" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This is a research workflow, not a deployed inventory system. No service, dashboard or production integration is claimed.&lt;/p&gt;

&lt;h2&gt;
  
  
  The approach
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fimages.weserv.nl%2F%3Furl%3Draw.githubusercontent.com%2Fofficialpm%2Fretail-demand-forecasting%2Fmain%2Fdocs%2Fworkflow.svg%26output%3Dpng%26w%3D1000" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fimages.weserv.nl%2F%3Furl%3Draw.githubusercontent.com%2Fofficialpm%2Fretail-demand-forecasting%2Fmain%2Fdocs%2Fworkflow.svg%26output%3Dpng%26w%3D1000" alt="Workflow: load, fit, validate, export" width="1000" height="297"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Predict each store-family series on its own. Zero-sales days stay in.&lt;/li&gt;
&lt;li&gt;Compare three seasonal baselines, a longer-window seasonal control, and two fixed promotion models (56-day and 112-day).&lt;/li&gt;
&lt;li&gt;Fit weekday intercepts and log promotion counts to log sales, with fixed ridge penalties instead of a parameter search.&lt;/li&gt;
&lt;li&gt;Evaluate five nonoverlapping 16-day horizons, using only earlier history for each one.&lt;/li&gt;
&lt;li&gt;Join forecasts to the sample IDs one-to-one. Reject missing, duplicate, nonfinite or negative output.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Results
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Measurement&lt;/th&gt;
&lt;th&gt;Result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Held-out test RMSLE (lower is better)&lt;/td&gt;
&lt;td&gt;0.40508&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Previous held-out test result&lt;/td&gt;
&lt;td&gt;0.41602&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Five-window development RMSLE, 56-day promotion model&lt;/td&gt;
&lt;td&gt;0.435361&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Same-window seasonal control&lt;/td&gt;
&lt;td&gt;0.460723&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Windows won against the seasonal control&lt;/td&gt;
&lt;td&gt;5 of 5&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fimages.weserv.nl%2F%3Furl%3Draw.githubusercontent.com%2Fofficialpm%2Fretail-demand-forecasting%2Fmain%2Fdocs%2Ftest-rmsle.svg%26output%3Dpng%26w%3D1000" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fimages.weserv.nl%2F%3Furl%3Draw.githubusercontent.com%2Fofficialpm%2Fretail-demand-forecasting%2Fmain%2Fdocs%2Ftest-rmsle.svg%26output%3Dpng%26w%3D1000" alt="Held-out test RMSLE, previous result vs final model" width="1000" height="444"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fimages.weserv.nl%2F%3Furl%3Draw.githubusercontent.com%2Fofficialpm%2Fretail-demand-forecasting%2Fmain%2Fdocs%2Fdev-rmsle.svg%26output%3Dpng%26w%3D1000" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fimages.weserv.nl%2F%3Furl%3Draw.githubusercontent.com%2Fofficialpm%2Fretail-demand-forecasting%2Fmain%2Fdocs%2Fdev-rmsle.svg%26output%3Dpng%26w%3D1000" alt="Development RMSLE across five windows, seasonal control vs promotion model" width="1000" height="444"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Daily sales summed across all stores and product families, zero-sales days included. The scale and variability change over time, so a single average is not a stable demand model. This plot is descriptive, not a causal promotion analysis.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fimages.weserv.nl%2F%3Furl%3Draw.githubusercontent.com%2Fofficialpm%2Fretail-demand-forecasting%2Fmain%2Fsales-history.svg%26output%3Dpng%26w%3D1000" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fimages.weserv.nl%2F%3Furl%3Draw.githubusercontent.com%2Fofficialpm%2Fretail-demand-forecasting%2Fmain%2Fsales-history.svg%26output%3Dpng%26w%3D1000" alt="Daily aggregate sales across the full history" width="1000" height="393"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Limits
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;The development windows were reused as the method evolved, so those numbers are selection-biased.&lt;/li&gt;
&lt;li&gt;The held-out test score is one measurement on one test period. It says nothing about new stores, years or demand patterns.&lt;/li&gt;
&lt;li&gt;Promotions are associated with sales here. The coefficient is not a causal estimate.&lt;/li&gt;
&lt;li&gt;No holidays, oil prices, transactions or stockout labels are modeled.&lt;/li&gt;
&lt;li&gt;The portable &lt;code&gt;train.py&lt;/code&gt; is an adaptation of the measured implementation. I ran synthetic tests (output shape, and invariance to future target changes), not a full retrain.&lt;/li&gt;
&lt;li&gt;Nothing about production latency or business impact has been measured.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The repo has the code, tests, a full explained notebook and a results ledger with provenance. Feedback on the validation setup is welcome.&lt;/p&gt;

&lt;p&gt;Code is Apache-2.0. The dataset is not included and is not covered by that license.&lt;/p&gt;

</description>
      <category>machinelearning</category>
      <category>python</category>
      <category>datascience</category>
    </item>
  </channel>
</rss>
