<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: kennytradingdev</title>
    <description>The latest articles on DEV Community by kennytradingdev (@kennytradingdev).</description>
    <link>https://dev.to/kennytradingdev</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4150218%2Fc97dcab7-b357-4a99-8bb2-5f0db573d910.png</url>
      <title>DEV Community: kennytradingdev</title>
      <link>https://dev.to/kennytradingdev</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/kennytradingdev"/>
    <language>en</language>
    <item>
      <title>I pre-registered a test of a famous market effect — here's how it failed</title>
      <dc:creator>kennytradingdev</dc:creator>
      <pubDate>Tue, 29 Sep 2026 16:08:39 +0000</pubDate>
      <link>https://dev.to/kennytradingdev/i-pre-registered-a-test-of-a-famous-market-effect-heres-how-it-failed-2f3</link>
      <guid>https://dev.to/kennytradingdev/i-pre-registered-a-test-of-a-famous-market-effect-heres-how-it-failed-2f3</guid>
      <description>&lt;p&gt;The rule cleared a randomly-placed alternative 93.8% of the time. The bar, written down and hashed before I pulled a single bar of data, was the 95th percentile.&lt;/p&gt;

&lt;p&gt;It missed by 0.12 percentage points of annualized return.&lt;/p&gt;

&lt;p&gt;Had I written that threshold &lt;em&gt;after&lt;/em&gt; seeing the number — at the 90th percentile, say, which is not an indefensible choice — this post would have been titled "A calendar anomaly that still works." Instead it's a failure report. That's the whole point, and it's the only reason&lt;br&gt;
  the near-miss is worth anything.&lt;/p&gt;

&lt;p&gt;## The effect&lt;/p&gt;

&lt;p&gt;There's a well-documented "turn-of-month" effect in equity indices: returns are said to concentrate in a short window spanning the month boundary. It appears in the academic literature, it appears in trading blogs, and it has the two properties that make a claim worth testing&lt;br&gt;
  rather than believing — it's specific enough to falsify, and it's old enough that if it were still payable, it probably shouldn't be.&lt;/p&gt;

&lt;p&gt;I tested it on SPY. Free daily data, 1993-01-29 (SPY's inception) through 2026-09-24. 8,470 daily return bars — 33.6 years.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Verdict: NOT TRADEABLE.&lt;/strong&gt; All five pre-registered decision clauses failed.&lt;/p&gt;

&lt;p&gt;## The method: hash the hypothesis first&lt;/p&gt;

&lt;p&gt;The core problem with backtesting isn't bad math. It's that the analyst and the referee are the same person, and the referee gets to write the rules after seeing the game.&lt;/p&gt;

&lt;p&gt;So I wrote the rules first, in a file, and hashed it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;  PREREGISTRATION.md   SHA-256  04396d92f6a45d2ffebff0faff7c4e31cfde983e11d3bdfd5b5bc9f8e7d8cba2
  spy_snapshot.csv     SHA-256  45d57b45db6e07b59a4780b3990b6b0a0bcdd21f9f23abe8cb47419a25004d9a
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The pre-registration fixes six things before any statistic exists:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;The window.&lt;/strong&gt; Last 1 trading day of month M plus the first 3 of M+1. Four days per boundary, counted by &lt;em&gt;bar index&lt;/em&gt;, not calendar date, so holidays can't shift it. 1,616 bars — 19.1% of the sample.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The statistic.&lt;/strong&gt; Δ = mean(daily log return | in-window) − mean(| out-of-window), in basis points per day.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The inference.&lt;/strong&gt; Stationary bootstrap, mean block length 20 days, 10,000 resamples, two-sided 95% CI. Block resampling because daily returns are serially dependent — an i.i.d. bootstrap would understate the interval and flatter the result.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The null.&lt;/strong&gt; Circular rotation: hold the return series fixed, rotate the 0/1 window label by a random offset. This preserves both the autocorrelation &lt;em&gt;and&lt;/em&gt; the exact count and spacing of labelled days, so it tests &lt;strong&gt;calendar alignment specifically&lt;/strong&gt; rather than "some days
are better than others."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The decision bar.&lt;/strong&gt; Five clauses. TRADEABLE only if all five hold.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The stopping rule.&lt;/strong&gt; One run. If the primary fails, the report records the failure; it does not go shopping for a variant that passes.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The hash is what makes this more than a promise. The report's header carries the SHA-256 of the pre-registration as frozen, and the runner writes both hashes into a manifest alongside the results. Anyone can verify that the document specifying the test is byte-identical to the&lt;br&gt;
  one that existed before the test ran. You don't have to trust me; you have to run &lt;code&gt;sha256sum&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;This is a cheap trick and it costs nothing. It is also, as far as I can tell, almost never done.&lt;/p&gt;

&lt;p&gt;## What came out&lt;/p&gt;

&lt;p&gt;The point estimate had the sign the literature advertises, and it was not small:&lt;/p&gt;

&lt;p&gt;| Quantity | Value |&lt;br&gt;
  |---|---|&lt;br&gt;
  | Mean return, in-window days | &lt;strong&gt;+7.28 bps/day&lt;/strong&gt; |&lt;br&gt;
  | Mean return, other days | &lt;strong&gt;+3.33 bps/day&lt;/strong&gt; |&lt;br&gt;
  | &lt;strong&gt;Δ&lt;/strong&gt; | &lt;strong&gt;+3.95 bps/day&lt;/strong&gt; |&lt;br&gt;
  | 95% CI (stationary bootstrap) | &lt;strong&gt;[−2.02, +9.80]&lt;/strong&gt; — spans zero |&lt;br&gt;
  | Rotation null, two-sided | &lt;strong&gt;p = 0.212&lt;/strong&gt; |&lt;/p&gt;

&lt;p&gt;Turn-of-month days averaged roughly &lt;strong&gt;2.2× the return of other days&lt;/strong&gt;. That ratio is the number a promotional writeup leads with, and it's true. It just doesn't survive its own interval.&lt;/p&gt;

&lt;p&gt;Here's the arithmetic that kills it. With 8,470 bars and daily volatility near 1.2%, the sampling error on a 19%/81% split of the sample is around &lt;strong&gt;±6 bps&lt;/strong&gt; — the same order of magnitude as the effect being claimed. The rotation null puts the observed Δ at the 79th percentile&lt;br&gt;
  of what random calendar alignment produces. &lt;strong&gt;One in five arbitrary rotations of the label beats it.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;## The bar, clause by clause&lt;/p&gt;

&lt;p&gt;| ID | Clause | Result | |&lt;br&gt;
  |---|---|---|---|&lt;br&gt;
  | B1 | 95% CI on Δ excludes zero | [−2.02, +9.80] | &lt;strong&gt;FAIL&lt;/strong&gt; |&lt;br&gt;
  | B2 | Rotation null p &amp;lt; 0.05 | p = 0.212 | &lt;strong&gt;FAIL&lt;/strong&gt; |&lt;br&gt;
  | B3 | Net return beats matched-exposure null's 95th pct @ 2 bps | 3.25% vs 3.38% | &lt;strong&gt;FAIL&lt;/strong&gt; |&lt;br&gt;
  | B4 | Same sign both halves &lt;strong&gt;and&lt;/strong&gt; late-half CI excludes zero | signs agree; late CI [−5.64, +9.79] | &lt;strong&gt;FAIL&lt;/strong&gt; |&lt;br&gt;
  | B5 | B3 still holds at 5 bps | 2.88% vs 3.02% | &lt;strong&gt;FAIL&lt;/strong&gt; |&lt;/p&gt;

&lt;p&gt;B4 deserves a careful reading, because it failed on its &lt;em&gt;second&lt;/em&gt; clause, not its first. The sign agreed across halves — the effect did not flip. It decayed and widened: +5.84 bps/day in 1993–2009, +2.03 bps/day in 2010–2026. Neither half was individually significant. So "the&lt;br&gt;
  effect reversed" would be wrong, and "confirmed in both halves" would also be wrong. The honest statement is that &lt;strong&gt;this sample cannot resolve the question in either half.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;## The near-miss, and why the baseline is the finding&lt;/p&gt;

&lt;p&gt;B3 and B5 are where the 93.8 lives, and they're the clauses I'd defend hardest.&lt;/p&gt;

&lt;p&gt;The rule is long-only, in the market about 4 of every 21 trading days: 404 round trips, ~12 per year, 19.1% time in market. Net of 2 bps round-trip costs it returned &lt;strong&gt;3.25% annualized&lt;/strong&gt; — annualized over the full sample calendar, so the 19% exposure is priced honestly rather&lt;br&gt;
  than annualized only over the days it was actually long.&lt;/p&gt;

&lt;p&gt;Compared to cash, 3.25% looks like a result. &lt;strong&gt;Compared to cash, it means nothing.&lt;/strong&gt; A rule that is long 19% of a 33-year bull market earns a positive return for reasons that have nothing whatsoever to do with the calendar.&lt;/p&gt;

&lt;p&gt;So the pre-registered comparator draws 10,000 alternative rules that hold the same 4 contiguous bars per month at a &lt;strong&gt;random offset inside the month&lt;/strong&gt;:&lt;/p&gt;

&lt;p&gt;| Cost | Rule | Null median | Null 95th pct | Rule's percentile |&lt;br&gt;
  |---|---|---|---|---|&lt;br&gt;
  | 2 bps | 3.25% | 1.42% | 3.38% | &lt;strong&gt;93.8th&lt;/strong&gt; |&lt;br&gt;
  | 5 bps | 2.88% | 1.03% | 3.02% | &lt;strong&gt;93.7th&lt;/strong&gt; |&lt;/p&gt;

&lt;p&gt;The rule beats a randomly-placed 4-day hold 94% of the time. That is genuinely toward the upper tail. It is also short of a threshold set before the data was opened, by 0.12pp at 2 bps and 0.13pp at 5 bps.&lt;/p&gt;

&lt;p&gt;A near-miss reported as a miss is the entire deliverable.&lt;/p&gt;

&lt;p&gt;## Two things I found along the way&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;One bar of execution delay costs more than the entire cost sweep.&lt;/strong&gt; I pre-declared two execution variants: enter at the prior close (E1) or at the next open (E2). That single bar of delay removes &lt;strong&gt;14.2% of the gross return&lt;/strong&gt; — 3.50% → 3.01%. That is larger than the whole&lt;br&gt;
  cost sweep from 1 bp to 10 bps. A backtest that fills at the same close its signal was computed from isn't conservative-by-a-little. On a 4-day hold, the entry bar is a quarter of the position's life.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The strongest result in the study is the one I'm least allowed to use.&lt;/strong&gt; Three alternative windows were pre-declared with Holm correction at family-wise α = 0.05. The variant that &lt;em&gt;drops&lt;/em&gt; the last day of the prior month more than doubles Δ to &lt;strong&gt;+8.35 bps/day at p = 0.020&lt;/strong&gt;.&lt;br&gt;
  Presented alone, that's a publishable-looking calendar anomaly.&lt;/p&gt;

&lt;p&gt;It fails Holm by a hair — 0.020 &amp;gt; 0.05/3 = 0.0167. And Holm here is &lt;em&gt;generous&lt;/em&gt;, because the family it corrects over is only the three windows I wrote down in advance. The honest family is much larger: the window has two free integer parameters, and a search over plausible&lt;br&gt;
  values of both offers a dozen or more candidates. Correcting over three when the search space held twelve &lt;strong&gt;understates&lt;/strong&gt; the correction.&lt;/p&gt;

&lt;p&gt;That variant is a lead worth re-registering as its own study, on out-of-sample or other-instrument data. It is not a result.&lt;/p&gt;

&lt;p&gt;## What this study cannot say&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;SPY only.&lt;/strong&gt; No cross-instrument replication exists inside the design. Nothing here generalizes to other indices, sizes, or countries.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Survivorship.&lt;/strong&gt; SPY is a survivor by construction. Accepted, because the claim is about calendar timing within one continuously-listed instrument — no cross-sectional claim is made.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Data vintage.&lt;/strong&gt; The adjusted series is revised retroactively when distributions are restated. The verdict is pinned to the hashed snapshot; &lt;code&gt;--offline&lt;/code&gt; reproduces it exactly.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;This is not evidence the effect is absent.&lt;/strong&gt; It is evidence this sample cannot distinguish it from zero. At SPY's 1.17% daily volatility, resolving an effect of this size would take about &lt;strong&gt;2.2× this sample — roughly 75 years of history.&lt;/strong&gt; For a once-a-month window, that's
data that does not exist. A null here is a statement about power as much as about the market.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;## One disclosure&lt;/p&gt;

&lt;p&gt;The study ran twice. The first run computed from the in-memory data pull while writing a snapshot rounded to six decimals, so the offline mode reproduced figures to about the sixth significant figure rather than exactly. I fixed the I/O layer so the runner computes from the&lt;br&gt;
  snapshot it writes, and re-ran on the &lt;strong&gt;same hashed snapshot&lt;/strong&gt; — no new pull, no change of vintage, no change to the hypothesis, the bar, or any parameter. The verdict and all five clause outcomes were identical.&lt;/p&gt;

&lt;p&gt;I found that flaw by running my own tool on my own study.&lt;/p&gt;

&lt;p&gt;## The tool&lt;/p&gt;

&lt;p&gt;The failures this design was built to avoid are mechanical enough to grep for, so I packaged the check: &lt;strong&gt;&lt;a href="https://github.com/kennytradingdev/backtest-honesty-check" rel="noopener noreferrer"&gt;backtest-honesty-check&lt;/a&gt;&lt;/strong&gt; — zero-dependency Python, MIT, 26 rules. Nineteen read your source as an AST&lt;br&gt;
  (indicators referencing future bars, a signal earning the same bar's return with no lag, centred rolling windows, a cost constant defined and never subtracted, &lt;code&gt;dropna(axis=1)&lt;/code&gt; quietly deleting the names that delisted). Seven screen a reported-results artifact (Sharpe above 3&lt;br&gt;
  on daily data, win rate above 90%, a big annual return against a near-zero drawdown, fewer than 30 trades).&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;  python3 honesty_check.py backtest.py &lt;span class="nt"&gt;--results&lt;/span&gt; metrics.json &lt;span class="nt"&gt;--fail-on&lt;/span&gt; HIGH
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Exit 0 clean, 1 findings, 2 critical. It is a smoke detector, not a fire marshal — the reference doc states plainly what it cannot do.&lt;/p&gt;

&lt;p&gt;## The takeaway&lt;/p&gt;

&lt;p&gt;A well-known calendar anomaly, tested once against a frozen bar, produced a positive point estimate that no clause of the bar could confirm. The rule made money in the sense that being long a rising market 19% of the time makes money, and it landed at the 94th percentile&lt;br&gt;
  against equal time in market. Suggestive. Not payable. Short of a threshold set before the data was opened.&lt;/p&gt;

&lt;p&gt;The deliverable was never the anomaly. It's the audit trail: a hashed hypothesis, one run, a null that preserves the structure it needs to preserve, a baseline matched on exposure rather than cash, costs as a swept dimension, and a near-miss recorded as a miss.&lt;/p&gt;

&lt;p&gt;If you take one thing from this: &lt;strong&gt;write down your threshold and hash the file.&lt;/strong&gt; It takes ninety seconds. It is the difference between a test and a story.&lt;/p&gt;

</description>
      <category>datascience</category>
      <category>python</category>
      <category>statistics</category>
      <category>finance</category>
    </item>
  </channel>
</rss>
