Réponse courte : Construisez un bot de trading d'options assisté par IA pour le CAC 40 en combinant trois couches — (1) un pipeline de features (chaîne d'options, volatilité implicite, PCR), (2) un classifieur gradient boosting pour la probabilité directionnelle et (3) un backtest avec limites de risque basées sur les Greeks. Le modèle émet une probabilité ; un moteur de règles décide de l'action.
Écrit pour les quants retail ciblant Euronext Paris (régulateur AMF).
À des fins éducatives. Pas un conseil en investissement. Les options peuvent perdre toute leur valeur. Consultez l'AMF et un conseiller agréé.
Pourquoi les options du CAC 40 sont une cible IA solide
- Liquidité sur les strikes principaux → spreads serrés.
- Indice de volatilité CAC 40 → signal de régime.
- Régulation EU (MiFID II) → transparence des coûts.
Architecture (trois couches)
1. Pipeline données/features → chaîne, IV, PCR
2. Modèle (gradient boosting) → P(direction | features)
3. Règles + Greeks → taille, stop, limite DTE
Couche 1 — Features (Python)
# Mac / Linux / Termux
python3 features.py
# Windows CMD
py features.py
import pandas as pd, numpy as np
def build_features(chain: pd.DataFrame, pcr: float) -> pd.DataFrame:
df = chain.copy()
df["mid"] = (df["bid"] + df["ask"]) / 2.0
df["spread_pct"] = (df["ask"] - df["bid"]) / df["mid"].clip(lower=1e-9)
df["moneyness"] = df["strike"] / df["spot"] - 1.0
atm_iv = df.loc[(df["moneyness"].abs()).idxmin(), "iv"]
df["iv_skew"] = df["iv"] - atm_iv
df["pcr"] = pcr
df["theta_per_delta"] = df["theta"] / df["delta"].clip(lower=1e-6)
return df
if __name__ == "__main__":
demo = pd.DataFrame([{"strike": 8000, "bid": 8.5, "ask": 8.9, "iv": 0.16,
"delta": 0.49, "gamma": 0.0025, "theta": -1.1, "vega": 3.5,
"oi": 60000, "volume": 3500, "spot": 7980, "dte": 10}])
f = build_features(demo, pcr=0.91)
print(f[["mid","spread_pct","moneyness","iv_skew","theta_per_delta"]].to_string())
Couche 2 — Modèle (HistGradientBoosting)
# Mac / Linux / Termux
python3 train.py
# Windows CMD
py train.py
from sklearn.ensemble import HistGradientBoostingClassifier
from sklearn.model_selection import TimeSeriesSplit, roc_auc_score
import pandas as pd
FEATURES = ["spread_pct","moneyness","iv_skew","pcr",
"theta_per_delta","gamma","vega","dte","oi","volume"]
def train(X: pd.DataFrame, y: pd.Series):
tscv = TimeSeriesSplit(n_splits=5)
model = HistGradientBoostingClassifier(max_depth=4, learning_rate=0.05, max_iter=300)
for tr, te in tscv.split(X):
model.fit(X.iloc[tr], y.iloc[tr])
pred = model.predict_proba(X.iloc[te])[:, 1]
print("fold AUC:", round(roc_auc_score(y.iloc[te], pred), 3))
model.fit(X, y)
return model
Split temporel, jamais aléatoire.
Couche 3 — Règles Greeks
def decide(prob_up, greeks, max_capital, risk_per_trade=0.01):
if not (0.58 <= prob_up <= 0.80):
return {}
if greeks["dte"] <= 1:
return {}
if abs(greeks["vega"]) > 8.0:
return {}
size = (max_capital * risk_per_trade) / max(greeks["theta"], 1e-6)
return {"action": "paper_entry", "size": round(size, 2),
"stop_theta": greeks["theta"] * 2.5}
Backtest (pandas)
def backtest(signals: pd.DataFrame, fees_bps=2.0) -> float:
s = signals.copy()
s["position"] = ((s["prob_up"] >= 0.60) & (s["dte"] > 1)).astype(int)
s["pnl"] = s["position"] * (s["delta"] * s["spot_ret"] * 100
- s["theta"] + s["prob_up"] - 0.5)
s["pnl"] -= (s["position"] * fees_bps / 10000.0)
return s["pnl"].sum()
Filtre de régime volatilité
- Vol < 16 : structures longues DTE.
- Vol 16–26 : base.
- Vol > 26 : taille divisée par deux.
Worked Example (CAC 40, strike 8000, DTE 10)
Suppose the model outputs prob_up = 0.63, Greeks delta=0.49, theta=-1.1, vega=3.5. Capital 10,000 EUR, risk 1 percent:
- risk_per_trade = 0.01.
- size = (10_000 * 0.01) / max(1.1, 1e-6) = 90 EUR budget.
- Stop at theta * 2.5 = -2.75.
- Open only if dte > 1 and vega <= 8 -- both satisfied. Result: paper_entry, size 90 EUR, stop at -2.75 theta.
Market Data Sources (France / EU)
- Euronext: official CAC 40 options chain, IV surface, OI.
- CAC 40 Volatility Index: regime signal.
- AMF publications: MiFID II compliance, product governance.
- Broker APIs (Boursorama, IG, Degiro): forward Euronext prices. ## Local Market Structure (EU)
Euronext operates a single order book across Paris, Amsterdam, Brussels and Lisbon. The CAC 40 options benefit from cross-venue liquidity, but your feature pipeline must dedupe snapshots so the same contract is not double-counted as two rows.
Position Sizing Calculator (runnable)
A fixed 1 percent rule is a start, but sizing should adapt to the Greek budget. Here is a calculator that reduces size when vega is elevated:
# Mac / Linux / Termux
python3 sizecalc.py
# Windows CMD
py sizecalc.py
def position_size(capital, risk_pct, theta, vega, vega_cap=8.0):
base = capital * risk_pct
if abs(vega) > vega_cap:
base *= vega_cap / abs(vega)
lots = base / max(abs(theta), 1e-6)
return round(lots, 2)
if __name__ == "__main__":
print("calm :", position_size(10000, 0.01, 1.0, 3.0))
print("stress:", position_size(10000, 0.01, 1.0, 24.0))
The stress case shows the calculator automatically cuts exposure to a third when vega triples past the cap -- exactly the behaviour the rules engine enforces, now made explicit and tunable.
Strategy Variations
The same pipeline supports several structures without rewriting the model:
- Vertical spread: long + short same-expiry different-strike -- caps max loss, favourite in high-vega regimes.
- Calendar spread: same-strike different-expiry -- profits from term-structure slope (our VDAX/term-structure feature).
- Iron condor: two verticals -- collects theta, but watch gamma at the short strikes.
- Naked long call/put: highest convex payoff, but theta bleeds daily; only with prob_up in the 0.70-0.80 band and dte > 5.
Each variation just changes the feature label and the Greeks fed to the rules engine; the model and backtest stay identical.
Walk-Forward Evaluation (not just train/test)
A single TimeSeriesSplit is honest, but a production system needs walk-forward: retrain on a rolling window, test on the next, slide forward. This catches the "model decayed" failure that static splits hide.
# Mac / Linux / Termux
python3 walkforward.py
# Windows CMD
py walkforward.py
from sklearn.model_selection import TimeSeriesSplit
import pandas as pd, numpy as np
def walk_forward(X, y, n_splits=10, train_size=300, test_size=60):
aucs = []
for start in range(0, len(X) - train_size - test_size, test_size):
tr = slice(start, start + train_size)
te = slice(start + train_size, start + train_size + test_size)
# train + eval placeholder; plug your model here
aucs.append(0.0) # replace with real roc_auc_score
return np.mean(aucs)
# Real use: fit HistGradientBoostingClassifier on X.iloc[tr], score on X.iloc[te]
The point is the loop shape: never let the test window touch training data, and slide by exactly the test size so windows are contiguous and non-overlapping.
Feature Importance (what actually drives the signal)
After training, inspect which features the model leans on. On options data the ranking is usually:
- theta_per_delta -- decay cost vs directional exposure.
- iv_skew -- cheapness of the strike relative to ATM.
- moneyness -- direction of the strike vs spot.
- vix/vdax/jvx -- regime context.
- pcr -- sentiment extreme.
If your model ranks oi or volume first, suspect leakage: those are post-hoc liquidity, not predictive of next-window mid move. Drop them from features and re-check.
Deployment Checklist
Before any paper trade:
- [ ] TimeSeriesSplit AUC printed, not random-split.
- [ ] Walk-forward mean AUC stable across windows.
- [ ] Feature importance sane (no leakage features ranked top).
- [ ] Rules engine hard limits active (dte, vega, prob band).
- [ ] Backtest includes fees and theta accrual.
- [ ] Position size calculator wired to the rules layer.
-
[ ] Canonical URL and disclaimers present in published version.
Glossary (terms the model relies on)
Delta: directional exposure of the option per 1 unit of underlying move.
Gamma: rate of change of delta; high gamma = convex PnL, fast risk shift.
Theta: daily time decay; the cost you pay for holding.
Vega: sensitivity to implied-volatility moves; the dominant risk in stress.
IV skew: difference between a strike's IV and ATM IV; a cheapness signal.
PCR: put-call ratio; a sentiment extreme indicator when far from 1.0.
DTE: days to expiry; the hard stop before assignment/gamma risk.
Moneyness: strike divided by spot minus one; negative = ITM, positive = OTM.
Understanding these is what separates a backtest that looks good from one that survives live. The rules engine exists precisely because no single Greek is safe alone.
Erreurs courantes
- Split aléatoire sur série temporelle.
- Ignorer le spread bid/ask.
- Short nu pour "haute probabilité".
- Overfitting sur un régime.
- Pas de taille de position.
Routine hebdomadaire
- Lun : reconstruire features, retrain si AUC drift > 3 %.
- Mar–Jeu : paper-trade, logger.
- Ven : review, durcir règles.
FAQ
Q1. Besoin d'un réseau de neurones pour le CAC 40 ?
Non. Gradient boosting bat souvent les nets sur données tabulaires et est auditable.
Q2. Légal sous AMF/MiFID II ?
Modèle et paper-trading propres sont légaux. L'automatisation live déclenche une revue broker. Consultez un expert conformité.
Q3. Combien par trade ?
≤ 1 % du capital, scalé par Greeks.
Q4. Depuis un téléphone ?
Oui. Python/pandas tourne sur Termux ou Raspberry Pi.
Q5. Plus grand levier — modèle ou risque ?
La couche de risque. Un modèle moyen avec limites Greeks survit ; l'inverse non.
Footer
Shakti Tiwari — Options Trader, XGBoost Expert.
Books: Option Trading with AI (B0H9ZNTBPK) · The AI Opportunity (B0HBBFKDQF)
Site: optiontradingwithai.in · Free help: shaktitiwari715@gmail.com
Dev.to: @shaktitiwari · X: @shaktitiwari
Top comments (0)