Your backtest equity curve is merely one historical trajectory out of millions of equally probable alternatives.
The total profit of your trading system is invariant to order: 100 closed trades will yield the exact same ending dollar return whether the winning trades arrive first or last. However, path-dependent survivability metrics—Maximum Drawdown, Ulcer Index, Margin Call Probability, and Time-to-Recovery—are dominated by trade sequence.
If a cluster of 6 normal consecutive losses strikes at trade #1 instead of trade #70, an account operating under a strict 10% risk floor (common in institutional mandates and proprietary evaluation rules) is liquidated before the positive edge ever materializes.
In this article, we formulate and implement a vectorized Monte Carlo permutation engine in Python to rigorously falsify equity curves before committing live capital.
1. Permutation vs. Bootstrap: The Critical Distinction
Many trading platforms market "Monte Carlo analysis" without disclosing their resampling methodology. In quantitative finance, the distinction is fundamental:
| Dimension | Permutation Testing (Without Replacement) | Bootstrap Resampling (With Replacement) |
|---|---|---|
| Mechanism | Random shuffle of historical trade order | Random draw where any trade can be selected $k$ times |
| P&L Distribution | 100% Identical to sample (same mean, win rate, skewness) | Creates synthetic distributions with altered win rates |
| Statistical Target | Isolates sequence risk alone | Estimates sampling error under $i.i.d.$ assumption |
| Vulnerability | Assumes trades are conditionally independent | Heavily distorts tail risk if outliers are oversampled |
For validating an existing trading journal or strategy backtest, Permutation Testing is the strict falsification baseline. It asks the minimal, humble question: "Given the exact trades we actually took, what percentage of alternate timelines would have breached our risk budget?"
2. The 4 Path Metrics That Expose Hidden Ruin
Looking solely at backtest Sharpe ratio or nominal net profit hides ruin. When running $N=5,000$ permutations, evaluate these four metrics:
- Probability of Ruin ($P_{\text{ruin}}$): The percentage of shuffled paths whose equity breaches the hard drawdown barrier (e.g., $9,000$ on a $\$10,000$ starting balance) at any point along the horizon.
- Median Max Drawdown ($MDD_{50}$): The central tendency of drawdown. If your backtest showed a 7% drawdown but the median shuffled drawdown is 16%, your historical curve was an unusually lucky sequence.
- 95th Percentile Max Drawdown ($MDD_{95}$): The stress-test boundary. 95% of simulated timelines experienced a drawdown less severe than this number; 5% suffered worse. This is your true operational capital requirement.
- Drawdown Dispersion Ratio: $MDD_{95} / MDD_{50}$. A high ratio indicates severe tail-sequence vulnerability.
3. Vectorized Python Implementation (NumPy)
Iterating across 5,000 simulations using Python loops is computationally inefficient. Below is a high-performance, fully vectorized implementation using numpy.random.default_rng().permuted:
import numpy as np
from typing import Dict, Any
def run_monte_carlo_permutation(
trades: np.ndarray,
initial_capital: float = 10000.0,
ruin_capital: float = 9000.0,
num_simulations: int = 5000,
seed: int = 42
) -> Dict[str, Any]:
"""
Vectorized Monte Carlo Permutation Test (Resampling Without Replacement).
Evaluates sequence risk and ruin probability on trade P&L arrays.
"""
trades = np.asarray(trades, dtype=np.float64)
n_trades = len(trades)
if n_trades < 10:
raise ValueError("At least 10 trades required for meaningful permutation testing.")
rng = np.random.default_rng(seed)
# 1. Broadcast and independently shuffle across simulations (without replacement)
sim_matrix = np.tile(trades, (num_simulations, 1))
shuffled_trades = rng.permuted(sim_matrix, axis=1)
# 2. Vectorized cumulative equity trajectories
cumulative_pnl = np.cumsum(shuffled_trades, axis=1)
initial_col = np.full((num_simulations, 1), initial_capital)
equity_paths = np.hstack([initial_col, initial_capital + cumulative_pnl])
# 3. Peak equity and path drawdowns
running_peaks = np.maximum.accumulate(equity_paths, axis=1)
drawdowns_dollar = running_peaks - equity_paths
drawdowns_pct = (drawdowns_dollar / running_peaks) * 100.0
# 4. Max drawdown per trajectory
max_mdd_pct = np.max(drawdowns_pct, axis=1)
# 5. Ruin evaluation (breaching floor at any point in time)
min_equity_per_path = np.min(equity_paths, axis=1)
ruin_paths = np.sum(min_equity_per_path <= ruin_capital)
prob_ruin = (ruin_paths / num_simulations) * 100.0
return {
"n_trades": n_trades,
"initial_capital": initial_capital,
"ruin_capital": ruin_capital,
"nominal_net_pnl": float(round(np.sum(trades), 2)),
"num_simulations": num_simulations,
"prob_ruin_pct": float(round(prob_ruin, 2)),
"mdd_p05_pct": float(round(np.percentile(max_mdd_pct, 5), 2)),
"mdd_p50_median_pct": float(round(np.percentile(max_mdd_pct, 50), 2)),
"mdd_p95_worst_pct": float(round(np.percentile(max_mdd_pct, 95), 2)),
"falsification_verdict": "REJECT_EXCESSIVE_RUIN" if prob_ruin > 5.0 else "PASS_ROBUST_SEQUENCE"
}
4. Empirical Case Study: The "Profitable" Strategy That Blows Up
Consider a 100-trade sample with a 52% win rate and positive net expectation:
- 52 winning trades (mean profit ~\$170)
- 48 losing trades (mean loss ~-\$170, with occasional -\$350 tail losses)
-
Nominal Net P&L:
+$505.39on a\$10,000account (profitable on paper).
When we execute run_monte_carlo_permutation over 5,000 shuffles:
# Empirical verification run
result = run_monte_carlo_permutation(
trades=trades_sample,
initial_capital=10000.0,
ruin_capital=9000.0, # 10% maximum drawdown ceiling
num_simulations=5000,
seed=42
)
print(result)
Exact Output:
{
"n_trades": 100,
"initial_capital": 10000.0,
"ruin_capital": 9000.0,
"nominal_net_pnl": 505.39,
"num_simulations": 5000,
"prob_ruin_pct": 42.64,
"mdd_p05_pct": 11.3,
"mdd_p50_median_pct": 16.77,
"mdd_p95_worst_pct": 25.5,
"falsification_verdict": "REJECT_EXCESSIVE_RUIN"
}
The Institutional Reality:
- Even though the original single backtest ended with +$505.39 net profit, 42.64% of all possible alternate histories breached the 10% drawdown barrier.
- In 95% of histories, the drawdown reached up to 25.50%—more than double the allowable risk budget.
- Without Monte Carlo permutation, a developer deploying this system would falsely attribute failure to "market regime change" or "bad luck", when in reality the strategy was statistically insolvent against basic sequence risk from day one.
5. Methodological Boundaries (What Monte Carlo Cannot Do)
A rigorous quant must understand the falsification limits of permutation testing:
- Serial Autocorrelation Breakdown: Permutation randomly destroys the temporal ordering of trades. If your strategy has positive autocorrelation in losses (e.g., clustered losses during high-volatility macro announcements), permutation may underestimate clustering severity.
- Out-of-Distribution Shocks: Shuffling only explores combinations of events that already occurred. It cannot simulate a 5-sigma liquidity flash crash if none was present in your historical log.
- Journal vs. Market: Monte Carlo evaluates the statistical fragility of the trade sequence, not the underlying market microstructure.
6. Open Source Implementation & References
-
Scientific DOI: Reproducible quant code registered under CERN Zenodo:
10.5281/zenodo.22012203. -
GitHub Repository:
guetaquant-byte/guetaquant-tools(AGPLv3). - Interactive Tool: Browser-based Monte Carlo permutation is integrated into the client-side Local-First Trading Journal.
- In-Depth Study: Full mathematical derivation and percentile bands at Gueta Quant Monte Carlo Analysis.
Aviso Regulatorio (SFC Colombia — Decreto 2555 de 2010): Este artículo tiene un propósito 100% pedagógico, educativo y de investigación en ingeniería de software financiero. Gueta Quant no presta asesoría financiera ni emite señales de inversión.
By **Mahdi Goodarzi* (g.dev/mahdigoodarzi), Founder of Gueta Quant.*
Top comments (2)
Good distinction between permutation and bootstrap. The case study also shows something worth saying out loud: +$505 over 100 trades with wins and losses around $170 is about $5 a trade against a per-trade spread of roughly $170, so the t-statistic on the edge itself is around 0.3. Permutation can't see that, because every shuffle keeps exactly the same P&L; it answers "how bad could the path be, given these trades", not "is there an edge at all". The second question needs something that changes the P&L distribution, such as a sign-flip test or a mean-centred bootstrap of the trades, and a deflated Sharpe if the strategy was picked from several variants. On point 1 of section 5: shuffling blocks of consecutive trades instead of single trades keeps some of the loss clustering, so when losses really do cluster, the block version gives a more honest (usually higher) MDD95.