DEV Community

kennytradingdev
kennytradingdev

Posted on AI-assisted

I pre-registered a test of a famous market effect — here's how it failed

The rule cleared a randomly-placed alternative 93.8% of the time. The bar, written down and hashed before I pulled a single bar of data, was the 95th percentile.

It missed by 0.12 percentage points of annualized return.

Had I written that threshold after seeing the number — at the 90th percentile, say, which is not an indefensible choice — this post would have been titled "A calendar anomaly that still works." Instead it's a failure report. That's the whole point, and it's the only reason
the near-miss is worth anything.

## The effect

There's a well-documented "turn-of-month" effect in equity indices: returns are said to concentrate in a short window spanning the month boundary. It appears in the academic literature, it appears in trading blogs, and it has the two properties that make a claim worth testing
rather than believing — it's specific enough to falsify, and it's old enough that if it were still payable, it probably shouldn't be.

I tested it on SPY. Free daily data, 1993-01-29 (SPY's inception) through 2026-09-24. 8,470 daily return bars — 33.6 years.

Verdict: NOT TRADEABLE. All five pre-registered decision clauses failed.

## The method: hash the hypothesis first

The core problem with backtesting isn't bad math. It's that the analyst and the referee are the same person, and the referee gets to write the rules after seeing the game.

So I wrote the rules first, in a file, and hashed it:

  PREREGISTRATION.md   SHA-256  04396d92f6a45d2ffebff0faff7c4e31cfde983e11d3bdfd5b5bc9f8e7d8cba2
  spy_snapshot.csv     SHA-256  45d57b45db6e07b59a4780b3990b6b0a0bcdd21f9f23abe8cb47419a25004d9a
Enter fullscreen mode Exit fullscreen mode

The pre-registration fixes six things before any statistic exists:

  1. The window. Last 1 trading day of month M plus the first 3 of M+1. Four days per boundary, counted by bar index, not calendar date, so holidays can't shift it. 1,616 bars — 19.1% of the sample.
  2. The statistic. Δ = mean(daily log return | in-window) − mean(| out-of-window), in basis points per day.
  3. The inference. Stationary bootstrap, mean block length 20 days, 10,000 resamples, two-sided 95% CI. Block resampling because daily returns are serially dependent — an i.i.d. bootstrap would understate the interval and flatter the result.
  4. The null. Circular rotation: hold the return series fixed, rotate the 0/1 window label by a random offset. This preserves both the autocorrelation and the exact count and spacing of labelled days, so it tests calendar alignment specifically rather than "some days are better than others."
  5. The decision bar. Five clauses. TRADEABLE only if all five hold.
  6. The stopping rule. One run. If the primary fails, the report records the failure; it does not go shopping for a variant that passes.

The hash is what makes this more than a promise. The report's header carries the SHA-256 of the pre-registration as frozen, and the runner writes both hashes into a manifest alongside the results. Anyone can verify that the document specifying the test is byte-identical to the
one that existed before the test ran. You don't have to trust me; you have to run sha256sum.

This is a cheap trick and it costs nothing. It is also, as far as I can tell, almost never done.

## What came out

The point estimate had the sign the literature advertises, and it was not small:

| Quantity | Value |
|---|---|
| Mean return, in-window days | +7.28 bps/day |
| Mean return, other days | +3.33 bps/day |
| Δ | +3.95 bps/day |
| 95% CI (stationary bootstrap) | [−2.02, +9.80] — spans zero |
| Rotation null, two-sided | p = 0.212 |

Turn-of-month days averaged roughly 2.2× the return of other days. That ratio is the number a promotional writeup leads with, and it's true. It just doesn't survive its own interval.

Here's the arithmetic that kills it. With 8,470 bars and daily volatility near 1.2%, the sampling error on a 19%/81% split of the sample is around ±6 bps — the same order of magnitude as the effect being claimed. The rotation null puts the observed Δ at the 79th percentile
of what random calendar alignment produces. One in five arbitrary rotations of the label beats it.

## The bar, clause by clause

| ID | Clause | Result | |
|---|---|---|---|
| B1 | 95% CI on Δ excludes zero | [−2.02, +9.80] | FAIL |
| B2 | Rotation null p < 0.05 | p = 0.212 | FAIL |
| B3 | Net return beats matched-exposure null's 95th pct @ 2 bps | 3.25% vs 3.38% | FAIL |
| B4 | Same sign both halves and late-half CI excludes zero | signs agree; late CI [−5.64, +9.79] | FAIL |
| B5 | B3 still holds at 5 bps | 2.88% vs 3.02% | FAIL |

B4 deserves a careful reading, because it failed on its second clause, not its first. The sign agreed across halves — the effect did not flip. It decayed and widened: +5.84 bps/day in 1993–2009, +2.03 bps/day in 2010–2026. Neither half was individually significant. So "the
effect reversed" would be wrong, and "confirmed in both halves" would also be wrong. The honest statement is that this sample cannot resolve the question in either half.

## The near-miss, and why the baseline is the finding

B3 and B5 are where the 93.8 lives, and they're the clauses I'd defend hardest.

The rule is long-only, in the market about 4 of every 21 trading days: 404 round trips, ~12 per year, 19.1% time in market. Net of 2 bps round-trip costs it returned 3.25% annualized — annualized over the full sample calendar, so the 19% exposure is priced honestly rather
than annualized only over the days it was actually long.

Compared to cash, 3.25% looks like a result. Compared to cash, it means nothing. A rule that is long 19% of a 33-year bull market earns a positive return for reasons that have nothing whatsoever to do with the calendar.

So the pre-registered comparator draws 10,000 alternative rules that hold the same 4 contiguous bars per month at a random offset inside the month:

| Cost | Rule | Null median | Null 95th pct | Rule's percentile |
|---|---|---|---|---|
| 2 bps | 3.25% | 1.42% | 3.38% | 93.8th |
| 5 bps | 2.88% | 1.03% | 3.02% | 93.7th |

The rule beats a randomly-placed 4-day hold 94% of the time. That is genuinely toward the upper tail. It is also short of a threshold set before the data was opened, by 0.12pp at 2 bps and 0.13pp at 5 bps.

A near-miss reported as a miss is the entire deliverable.

## Two things I found along the way

One bar of execution delay costs more than the entire cost sweep. I pre-declared two execution variants: enter at the prior close (E1) or at the next open (E2). That single bar of delay removes 14.2% of the gross return — 3.50% → 3.01%. That is larger than the whole
cost sweep from 1 bp to 10 bps. A backtest that fills at the same close its signal was computed from isn't conservative-by-a-little. On a 4-day hold, the entry bar is a quarter of the position's life.

The strongest result in the study is the one I'm least allowed to use. Three alternative windows were pre-declared with Holm correction at family-wise α = 0.05. The variant that drops the last day of the prior month more than doubles Δ to +8.35 bps/day at p = 0.020.
Presented alone, that's a publishable-looking calendar anomaly.

It fails Holm by a hair — 0.020 > 0.05/3 = 0.0167. And Holm here is generous, because the family it corrects over is only the three windows I wrote down in advance. The honest family is much larger: the window has two free integer parameters, and a search over plausible
values of both offers a dozen or more candidates. Correcting over three when the search space held twelve understates the correction.

That variant is a lead worth re-registering as its own study, on out-of-sample or other-instrument data. It is not a result.

## What this study cannot say

  • SPY only. No cross-instrument replication exists inside the design. Nothing here generalizes to other indices, sizes, or countries.
  • Survivorship. SPY is a survivor by construction. Accepted, because the claim is about calendar timing within one continuously-listed instrument — no cross-sectional claim is made.
  • Data vintage. The adjusted series is revised retroactively when distributions are restated. The verdict is pinned to the hashed snapshot; --offline reproduces it exactly.
  • This is not evidence the effect is absent. It is evidence this sample cannot distinguish it from zero. At SPY's 1.17% daily volatility, resolving an effect of this size would take about 2.2× this sample — roughly 75 years of history. For a once-a-month window, that's data that does not exist. A null here is a statement about power as much as about the market.

## One disclosure

The study ran twice. The first run computed from the in-memory data pull while writing a snapshot rounded to six decimals, so the offline mode reproduced figures to about the sixth significant figure rather than exactly. I fixed the I/O layer so the runner computes from the
snapshot it writes, and re-ran on the same hashed snapshot — no new pull, no change of vintage, no change to the hypothesis, the bar, or any parameter. The verdict and all five clause outcomes were identical.

I found that flaw by running my own tool on my own study.

## The tool

The failures this design was built to avoid are mechanical enough to grep for, so I packaged the check: backtest-honesty-check — zero-dependency Python, MIT, 26 rules. Nineteen read your source as an AST
(indicators referencing future bars, a signal earning the same bar's return with no lag, centred rolling windows, a cost constant defined and never subtracted, dropna(axis=1) quietly deleting the names that delisted). Seven screen a reported-results artifact (Sharpe above 3
on daily data, win rate above 90%, a big annual return against a near-zero drawdown, fewer than 30 trades).

  python3 honesty_check.py backtest.py --results metrics.json --fail-on HIGH
Enter fullscreen mode Exit fullscreen mode

Exit 0 clean, 1 findings, 2 critical. It is a smoke detector, not a fire marshal — the reference doc states plainly what it cannot do.

## The takeaway

A well-known calendar anomaly, tested once against a frozen bar, produced a positive point estimate that no clause of the bar could confirm. The rule made money in the sense that being long a rising market 19% of the time makes money, and it landed at the 94th percentile
against equal time in market. Suggestive. Not payable. Short of a threshold set before the data was opened.

The deliverable was never the anomaly. It's the audit trail: a hashed hypothesis, one run, a null that preserves the structure it needs to preserve, a baseline matched on exposure rather than cash, costs as a swept dimension, and a near-miss recorded as a miss.

If you take one thing from this: write down your threshold and hash the file. It takes ninety seconds. It is the difference between a test and a story.

Top comments (1)

Collapse
 
arhancanli profile image
Arhan Canli •

Hashing the pre-registration and the data snapshot is the part more people should copy, and reporting a 93.8th-percentile near-miss as a miss is exactly what the hash is for.

One number I'd add next to the verdict: the test's power. Your CI on Δ implies a standard error of about 3.0 bps/day, so if the true effect were exactly the 3.95 bps you measured, B1 would pass only about 26% of the time. The smallest effect this sample could detect with 80% power is about 8.4 bps/day, and confirming 4 bps at that power would take roughly 150 years of SPY. So "NOT TRADEABLE" here is partly "the sample can't tell", which your B4 reading already says for each half. Putting the minimum detectable effect in the pre-registration makes that visible before the run: a reader can see which verdicts the test was capable of reaching.

On the dropped-last-day variant, I agree the honest family is bigger than three. Counting every (start, length) pair a reasonable person might have tried as a trial, even roughly, and deflating the best one's p-value by that count would put a number on how much of its +8.35 bps is selection.