DEV Community

Gueta Quant
Gueta Quant

Posted on Originally published at guetaquant.com

Monte Carlo Permutation Testing in Python: How Random Trade Order Falsifies Your Equity Curve

Your backtest equity curve is merely one historical trajectory out of millions of equally probable alternatives.

The total profit of your trading system is invariant to order: 100 closed trades will yield the exact same ending dollar return whether the winning trades arrive first or last. However, path-dependent survivability metrics—Maximum Drawdown, Ulcer Index, Margin Call Probability, and Time-to-Recovery—are dominated by trade sequence.

If a cluster of 6 normal consecutive losses strikes at trade #1 instead of trade #70, an account operating under a strict 10% risk floor (common in institutional mandates and proprietary evaluation rules) is liquidated before the positive edge ever materializes.

In this article, we formulate and implement a vectorized Monte Carlo permutation engine in Python to rigorously falsify equity curves before committing live capital.


1. Permutation vs. Bootstrap: The Critical Distinction

Many trading platforms market "Monte Carlo analysis" without disclosing their resampling methodology. In quantitative finance, the distinction is fundamental:

Dimension Permutation Testing (Without Replacement) Bootstrap Resampling (With Replacement)
Mechanism Random shuffle of historical trade order Random draw where any trade can be selected $k$ times
P&L Distribution 100% Identical to sample (same mean, win rate, skewness) Creates synthetic distributions with altered win rates
Statistical Target Isolates sequence risk alone Estimates sampling error under $i.i.d.$ assumption
Vulnerability Assumes trades are conditionally independent Heavily distorts tail risk if outliers are oversampled

For validating an existing trading journal or strategy backtest, Permutation Testing is the strict falsification baseline. It asks the minimal, humble question: "Given the exact trades we actually took, what percentage of alternate timelines would have breached our risk budget?"


2. The 4 Path Metrics That Expose Hidden Ruin

Looking solely at backtest Sharpe ratio or nominal net profit hides ruin. When running $N=5,000$ permutations, evaluate these four metrics:

  1. Probability of Ruin ($P_{\text{ruin}}$): The percentage of shuffled paths whose equity breaches the hard drawdown barrier (e.g., $9,000$ on a $\$10,000$ starting balance) at any point along the horizon.
  2. Median Max Drawdown ($MDD_{50}$): The central tendency of drawdown. If your backtest showed a 7% drawdown but the median shuffled drawdown is 16%, your historical curve was an unusually lucky sequence.
  3. 95th Percentile Max Drawdown ($MDD_{95}$): The stress-test boundary. 95% of simulated timelines experienced a drawdown less severe than this number; 5% suffered worse. This is your true operational capital requirement.
  4. Drawdown Dispersion Ratio: $MDD_{95} / MDD_{50}$. A high ratio indicates severe tail-sequence vulnerability.

3. Vectorized Python Implementation (NumPy)

Iterating across 5,000 simulations using Python loops is computationally inefficient. Below is a high-performance, fully vectorized implementation using numpy.random.default_rng().permuted:

import numpy as np
from typing import Dict, Any

def run_monte_carlo_permutation(
    trades: np.ndarray,
    initial_capital: float = 10000.0,
    ruin_capital: float = 9000.0,
    num_simulations: int = 5000,
    seed: int = 42
) -> Dict[str, Any]:
    """
    Vectorized Monte Carlo Permutation Test (Resampling Without Replacement).
    Evaluates sequence risk and ruin probability on trade P&L arrays.
    """
    trades = np.asarray(trades, dtype=np.float64)
    n_trades = len(trades)
    if n_trades < 10:
        raise ValueError("At least 10 trades required for meaningful permutation testing.")

    rng = np.random.default_rng(seed)

    # 1. Broadcast and independently shuffle across simulations (without replacement)
    sim_matrix = np.tile(trades, (num_simulations, 1))
    shuffled_trades = rng.permuted(sim_matrix, axis=1)

    # 2. Vectorized cumulative equity trajectories
    cumulative_pnl = np.cumsum(shuffled_trades, axis=1)
    initial_col = np.full((num_simulations, 1), initial_capital)
    equity_paths = np.hstack([initial_col, initial_capital + cumulative_pnl])

    # 3. Peak equity and path drawdowns
    running_peaks = np.maximum.accumulate(equity_paths, axis=1)
    drawdowns_dollar = running_peaks - equity_paths
    drawdowns_pct = (drawdowns_dollar / running_peaks) * 100.0

    # 4. Max drawdown per trajectory
    max_mdd_pct = np.max(drawdowns_pct, axis=1)

    # 5. Ruin evaluation (breaching floor at any point in time)
    min_equity_per_path = np.min(equity_paths, axis=1)
    ruin_paths = np.sum(min_equity_per_path <= ruin_capital)
    prob_ruin = (ruin_paths / num_simulations) * 100.0

    return {
        "n_trades": n_trades,
        "initial_capital": initial_capital,
        "ruin_capital": ruin_capital,
        "nominal_net_pnl": float(round(np.sum(trades), 2)),
        "num_simulations": num_simulations,
        "prob_ruin_pct": float(round(prob_ruin, 2)),
        "mdd_p05_pct": float(round(np.percentile(max_mdd_pct, 5), 2)),
        "mdd_p50_median_pct": float(round(np.percentile(max_mdd_pct, 50), 2)),
        "mdd_p95_worst_pct": float(round(np.percentile(max_mdd_pct, 95), 2)),
        "falsification_verdict": "REJECT_EXCESSIVE_RUIN" if prob_ruin > 5.0 else "PASS_ROBUST_SEQUENCE"
    }
Enter fullscreen mode Exit fullscreen mode

4. Empirical Case Study: The "Profitable" Strategy That Blows Up

Consider a 100-trade sample with a 52% win rate and positive net expectation:

  • 52 winning trades (mean profit ~\$170)
  • 48 losing trades (mean loss ~-\$170, with occasional -\$350 tail losses)
  • Nominal Net P&L: +$505.39 on a \$10,000 account (profitable on paper).

When we execute run_monte_carlo_permutation over 5,000 shuffles:

# Empirical verification run
result = run_monte_carlo_permutation(
    trades=trades_sample,
    initial_capital=10000.0,
    ruin_capital=9000.0, # 10% maximum drawdown ceiling
    num_simulations=5000,
    seed=42
)

print(result)
Enter fullscreen mode Exit fullscreen mode

Exact Output:

{
  "n_trades": 100,
  "initial_capital": 10000.0,
  "ruin_capital": 9000.0,
  "nominal_net_pnl": 505.39,
  "num_simulations": 5000,
  "prob_ruin_pct": 42.64,
  "mdd_p05_pct": 11.3,
  "mdd_p50_median_pct": 16.77,
  "mdd_p95_worst_pct": 25.5,
  "falsification_verdict": "REJECT_EXCESSIVE_RUIN"
}
Enter fullscreen mode Exit fullscreen mode

The Institutional Reality:

  • Even though the original single backtest ended with +$505.39 net profit, 42.64% of all possible alternate histories breached the 10% drawdown barrier.
  • In 95% of histories, the drawdown reached up to 25.50%—more than double the allowable risk budget.
  • Without Monte Carlo permutation, a developer deploying this system would falsely attribute failure to "market regime change" or "bad luck", when in reality the strategy was statistically insolvent against basic sequence risk from day one.

5. Methodological Boundaries (What Monte Carlo Cannot Do)

A rigorous quant must understand the falsification limits of permutation testing:

  1. Serial Autocorrelation Breakdown: Permutation randomly destroys the temporal ordering of trades. If your strategy has positive autocorrelation in losses (e.g., clustered losses during high-volatility macro announcements), permutation may underestimate clustering severity.
  2. Out-of-Distribution Shocks: Shuffling only explores combinations of events that already occurred. It cannot simulate a 5-sigma liquidity flash crash if none was present in your historical log.
  3. Journal vs. Market: Monte Carlo evaluates the statistical fragility of the trade sequence, not the underlying market microstructure.

6. Open Source Implementation & References


Aviso Regulatorio (SFC Colombia — Decreto 2555 de 2010): Este artículo tiene un propósito 100% pedagógico, educativo y de investigación en ingeniería de software financiero. Gueta Quant no presta asesoría financiera ni emite señales de inversión.

By **Mahdi Goodarzi* (g.dev/mahdigoodarzi), Founder of Gueta Quant.*

Top comments (2)

Collapse
 
arhancanli profile image
Arhan Canli •

Good distinction between permutation and bootstrap. The case study also shows something worth saying out loud: +$505 over 100 trades with wins and losses around $170 is about $5 a trade against a per-trade spread of roughly $170, so the t-statistic on the edge itself is around 0.3. Permutation can't see that, because every shuffle keeps exactly the same P&L; it answers "how bad could the path be, given these trades", not "is there an edge at all". The second question needs something that changes the P&L distribution, such as a sign-flip test or a mean-centred bootstrap of the trades, and a deflated Sharpe if the strategy was picked from several variants. On point 1 of section 5: shuffling blocks of consecutive trades instead of single trades keeps some of the loss clustering, so when losses really do cluster, the block version gives a more honest (usually higher) MDD95.