π The Quantitative Microstructure & Cross-Market Series by FollowSM
- Part 1: Detecting Orderbook Toxicity & Polymarket Divergence: Building a Microstructure Circuit Breaker
Packages:
-
Python (PyPI):
pip install followsm-sdk -
Node.js (npm):
npm install @followsm/sdk
We built an open-source hedger that delta-hedges Polymarket BTC positions with Binance perpetual futures and scales the hedge with order-flow toxicity. The logic: stay unhedged while the flow is calm, hedge half when VPIN climbs, go fully neutral when it spikes or a liquidity sweep hits.
Then we backtested it on 256 real daily Polymarket markets covering nine months, and it mostly didn't work.
π TL;DR
- Gating on VPIN rebuilt from public 1-minute data added nothing. A placebo, with the same hedges at randomly shifted times, did as well. On P&L spread it did as well or better in 19 of 20 seeds.
- High VPIN did not warn of big moves. The top-5% VPIN regime was followed by a tail-sized BTC move less often than average: 0.54β0.61Γ the base rate.
- Most of the tail is the jump at expiry. A binary option's delta explodes in its final hour, and no perp hedge keeps up.
- What worked was constant delta neutrality plus an exit one hour before resolution. With a wider rebalance band, that cut the 5% expected shortfall roughly in half, from β$742 to β$385 per 1,000-share position. Hedging cost fell to $36.
Everything below is reproducible from public data:
Dynamic Inventory Hedger
This tool delta-hedges Polymarket crypto positions with Binance USDβ-M perpetual futures. How hard it hedges depends on FollowSM microstructure toxicity: VPIN percentile, orderbook imbalance, liquidity sweeps and smart-money flow.
A Polymarket "Bitcoin above $84,000 on Friday" position is really a digital option on BTC. When informed flow hits the market, that position moves before any stop-loss can react. The hedger turns each position into an equivalent perp exposure and then:
- stays unhedged while the flow is normal;
- builds a hedge with post-only limit orders when toxicity rises;
- neutralises its exposure with price-protected IOC orders when toxicity spikes or a liquidity sweep hits;
- unwinds the hedge once the flow normalises.
It runs in paper mode by default.
Read the backtest before relying on toxicity gating. On 256 daily BTC markets (JanβSep 2026), gating on VPIN rebuilt from public 1-minute data did no better than the same hedgesβ¦
1. The idea
A Polymarket position such as "Will the price of Bitcoin be above $84,000 on Friday?" is a digital option on BTC. These markets settle on the Binance BTC/USDT close, the same price a Binance perp tracks. So the position has a delta you can hedge.
Take a strike K, spot S and volatility to expiry Ο_Ο. The YES price is roughly N(d), where
The perp-equivalent exposure of N shares is
At the money with one day left (ΟΟ β 2%), that is about $20 of perp per share. A 1,000-share position, worth at most $1,000 at settlement, behaves like a $20,000 BTC position. Near expiry, ΟΟ β 0 and the delta blows up.
The hedger runs four layers:
-
Signals. FollowSM microstructure (VPIN percentile, 1% book toxicity, liquidity sweeps, smart-money flow) maps to one of three levels:
NORMAL,PRE_HEDGING_ALERT,EMERGENCY_HEDGE_EXECUTION. Hysteresis stops it flip-flopping. - Inventory. Polymarket positions become Ξ$ per underlying.
- Execution. Hedge 0% / 50% / 100% by level. Post-only slices when patient, price-protected IOC when urgent, and a hard inventory limit throughout.
- Risk. Spread and slippage guards, a failure circuit breaker, and a JSONL journal.
The hypothesis was simple: informed flow shows up in the microstructure before the price moves, so hedging only then should get most of the protection at a fraction of the cost.
2. The test
Universe. Every daily "Bitcoin above $K on <date>" event from 2 January to 28 September 2026. At 24 hours before resolution, the book holds 1,000 shares of the strike nearest spot, held to the official resolution. Each market is run twice, long YES and long NO. The two books mirror each other, so directional drift cancels. The sample is 256 markets (512 books); 8 were skipped because no strike was priced between 0.10 and 0.90 at entry.
Replay. Minute by minute, using the hedger's own code: the same evaluator, delta model and decide() function that run live.
- Polymarket marks come from the public CLOB minute price history.
- The hedge trades the Binance USDβ-M BTCUSDT perp at the minute close.
- Maker fee 2 bps (post-only, assumed filled within the minute); taker 5 bps + 1 bp slippage (IOC).
- Real funding payments.
Signals. Rebuilt from Binance's public 1-minute spot klines, which carry the exact taker-buy notional per bar:
-
VPIN with $1M notional buckets and a 50-bucket window.
vpin_percentileis its rank within the previous 7 days. A robustness run uses activity-scaled buckets (trailing 7-day average daily volume / 50). - Liquidity sweep: a 15-minute volume Z-score β₯ 6 together with a 15-minute move β₯ 2 Γ NATR.
- Not tested: 1% order-book toxicity, smart-money sweeps and Polymarket book flow have no public history, so they are off in this backtest.
This is our own reconstruction for research, not FollowSM's production feed. Keep that in mind for everything below.
| Strategy | What it does |
|---|---|
| Unhedged | Never trades the perp |
| Inventory limit only | Hedges only the delta above $25k |
| Toxicity-gated | Repo defaults: 0% / 50% / 100% by level, plus the inventory limit |
| Placebo | The gated level path shifted by a random offset in time (20 seeds): same share of time hedged, same regime lengths, wrong timing |
| Always hedged | 100% delta hedge at every level |
The placebo is the key control. If toxicity timing carries information, the gated hedger must beat the same hedges placed at random times. Beating "unhedged" alone isn't enough, because any hedge reduces variance.
3. Results
USD per 1,000-share position, $1M buckets:
| Strategy | Mean P&L | Std dev | Expected shortfall 5% | Worst | Hedge cost | Perp traded |
|---|---|---|---|---|---|---|
| Unhedged | 0 | 473 | β742 | β885 | 0 | 0 |
| Inventory limit only | β115 | 394 | β829 | β1,177 | 101 | $288k |
| Toxicity-gated | β152 | 383 | β887 | β1,177 | 134 | $408k |
| Placebo (mean, 20 seeds) | β159 | 375 | β887 | β1,211 | 141 | $429k |
| Always 100% hedged | β130 | 242 | β719 | β1,020 | 130 | $649k |
Unhedged mean P&L is 0 because YES and NO mirror each other. Every hedged strategy's mean equals minus its hedging cost; the hedges earn no alpha, they only reshape risk.
Finding 1: the toxicity timing adds nothing
The gated hedger and the placebo are indistinguishable: expected shortfall β887 for both, and standard deviation 383 against 375. Across the 20 placebo seeds, random timing matched or beat the gated strategy:
- on expected shortfall in 50% of seeds;
- on standard deviation in 95%.
The activity-scaled VPIN run gave the same picture: 55% and 95%.
Gating also did worse than hedging all the time. It paid about the same cost, $134 against $130, but cut P&L variance by 34% where always-hedging cut it by 74%. Once costs are counted, it had a worse tail than not hedging at all.
Finding 2: VPIN from 1-minute data did not flag tail moves
This result undercuts the premise itself. We split every minute of the nine months by VPIN percentile, then asked: how often did a top-1% BTC move (by size) follow within 15 minutes?
| VPIN percentile | Share of time | Mean |15m move| | Tail-move rate vs base ($1M buckets) | Activity-scaled buckets |
|---|---|---|---|---|
| < 0.60 | 60% | 0.165% | 1.24Γ | 1.05Γ |
| 0.60β0.85 | 25% | 0.132% | 0.65Γ | 1.03Γ |
| 0.85β0.95 | 10% | 0.121% | 0.64Γ | 1.02Γ |
| β₯ 0.95 | 5% | 0.120% | 0.61Γ | 0.54Γ |
With fixed $1M buckets the relation runs backwards: high VPIN comes before calmer markets. With activity-scaled buckets it is flat, then backwards at the top.
The worst position in the sample shows what this means in practice. It was a NO on "Bitcoin above $64k on August 17" that lost $885 unhedged. With $1M buckets the evaluator stayed NORMAL for all 24 hours, and the gated line moves only because the inventory limit kicked in. With activity-scaled buckets it raised PRE_HEDGING_ALERT during the calm first 13 hours, then went quiet before the collapse:
In our previous article, Detecting Orderbook Toxicity & Polymarket Divergence, we reported a 1.4Γ tail-move lift for slow VPIN over 60 days. We called it "a hint, not a result" because it rested on about five events. With nine months of data and a rolling percentile, the hint does not hold up. The other observation from that article does hold: with fast $1M buckets the relationship ran backwards there too.
Why 1-minute VPIN fails: the blurring problem
VPIN was designed for trade-by-trade data: fill fixed-volume buckets in the order trades happen, and measure how one-sided each bucket is. Rebuilding it from 1-minute bars breaks this in three ways.
- Everything inside a minute becomes one number. A minute that prints $56M is split pro rata into 56 buckets of $1M, all with that minute's single buy/sell ratio. The order inside the minute is lost: a one-sided sweep at second 5 and a rebound at second 40 look identical to steady two-way flow. A crash with its rebound averages to "balanced", and VPIN drops at the exact moment flow is most toxic. The 22 August case study in hft-toxicity-circuit-breaker showed exactly this: fast VPIN collapsed as $43M of two-sided volume poured into the rebound.
- Quiet markets look toxic. In thin hours a $1M bucket spans many minutes of sparse, often one-sided flow, so the imbalance reads high. The busiest, most dangerous minutes fill dozens of buckets with averaged-out flow. That is how you get the inverse relation above.
- VPIN looks backwards. A 50-bucket window reports on flow that has already traded. By the time it peaks, the move it describes has often happened, and markets tend to calm afterwards. Andersen & Bondarenko (2014) reached a similar conclusion about VPIN and the 2010 Flash Crash.
None of this proves microstructure gating is useless. It shows that 1-minute public bars can't carry the signal. In the 22 August sweep, what moved first was one-sided taker flow inside the minute and the bid side thinning within 1% of mid. Only tick-level trade data and L1/L2 book snapshots can see those: per-trade VPIN, ob_toxicity_1pct, depth imbalance and sweep detection.
We couldn't test any of that here because none of it has public history. That is the honest limit of this study, and the next experiment is below.
Finding 3: the tail is the expiry jump, and early exit + delta neutrality fixes most of it
Look at the worst position again: most of the damage happens in the final hour. A digital option's delta, Ο(d)/ΟΟ, explodes as ΟΟ β 0 near the strike. No perp hedge can track a delta that swings by tens of thousands of dollars within minutes. The hedger zeroes delta in the final 5 minutes by design, which leaves the settlement jump fully exposed.
So we ran two changes that don't depend on any signal:
- Exit one hour before resolution. Close the Polymarket position at the market price, 60 minutes early.
- Stay delta-neutral the whole time (hedge ratio 1.0 at every level), with a wider rebalance band so the hedge doesn't churn.
| Strategy | Mean P&L | Std dev | Expected shortfall 5% | Worst | Hedge cost | Perp traded |
|---|---|---|---|---|---|---|
| Unhedged, to expiry | 0 | 473 | β742 | β885 | 0 | 0 |
| Unhedged, exit 1h early | 0 | 402 | β682 | β839 | 0 | 0 |
| Gated, exit 1h early | β145 | 324 | β717 | β870 | 128 | $385k |
| Always hedged, exit 1h early, $150 rebalance band | β127 | 171 | β509 | β776 | 126 | $610k |
| Always hedged, exit 1h early, $5,000 band | β38 | 164 | β385 | β535 | 36 | $158k |
βΆ Full rebalance-band sweep (always hedged, exit 1h early)
Rebalance band
Mean P&L
Std dev
Expected shortfall 5%
Worst
Hedge cost
Perp traded
$150
β127
171
β509
β776
126
$610k
$500
β113
171
β494
β760
112
$539k
$1,000
β96
169
β472
β731
94
$451k
$2,500
β58
164
β421
β638
57
$263k
$5,000
β38
164
β385
β535
36
$158k
Three points stand out:
- Exiting early alone helps a little (β742 β β682). Constant delta neutrality does most of the work once the settlement jump is gone (β β509).
- Churn was the real cost. With a $150 rebalance band, a position worth at most $1,000 turned over $610k of perp in a day. At maker fees, that alone is about $122. A $5,000 band cut turnover by three-quarters, lowered the tail (β385), and brought the cost down to $36.
- Both VPIN definitions gave the same result to within a few dollars. None of these strategies use VPIN.
A caution on the band: the grid was chosen in-sample, and the best value sits at its edge. Treat "wider is better" as a direction, not a tuned parameter.
4. Caveats
- Rebuilt signals only. These are our reconstruction from 1-minute public data, not FollowSM's tick-level feed. Book toxicity, smart-money sweeps and Polymarket book flow were not tested. The conclusion is "1-minute VPIN gating doesn't help", not "microstructure gating can't help".
- Narrow scope. Only BTC daily at-the-money markets were tested, over nine months. Other underlyings, strikes far from the money, and short-dated Up/Down markets may behave differently.
- Noisy tail. Expected shortfall rests on the worst 25 of 512 books.
- Optimistic execution. Post-only fills are assumed to be immediate. The early Polymarket exit is marked at the minute price, with no spread or slippage modelled; a real exit pays the spread.
- Manual exit. The hedger doesn't trade Polymarket, so the early exit is your job.
5. What we'd do with a Polymarket crypto book tomorrow
- Hedge the delta constantly. Don't try to time it with 1-minute order-flow statistics.
- Use a wide rebalance band. Binary deltas swing hard; a tight band pays fees to chase noise.
- Get out before the final hour. Delta hedging can't manage the settlement jump; not holding the position through it can.
- Test microstructure gating on the data it was designed for. That means tick-level trades and L2 books, logged live and replayed through the same harness. The hedger's journal already records every metric it acts on, so a few weeks of live paper trading gives the dataset this backtest lacked. That is our next experiment.
In the hedger that means:
BASELINE_HEDGE_RATIO=1.0
PRE_HEDGE_RATIO=1.0
EMERGENCY_HEDGE_RATIO=1.0
MIN_REBALANCE_USD=5000 # in-sample; see the caveat above
plus your own exit an hour before resolution.
The first run downloads and caches about nine months of public Binance and Polymarket data. Results land in βΆ Reproduce the backtest (no API keys needed)
git clone https://github.com/Follow-SM/dynamic-inventory-hedger.git
cd dynamic-inventory-hedger
pip install -e ".[research]"
python research/backtest.py # $1M buckets
VPIN_BUCKET=adv python research/backtest.py # activity-scaled buckets
research/results_*.json and charts in research/assets/.
Get started
- The hedger: dynamic-inventory-hedger. Paper mode by default, then Binance demo, then live.
-
Python SDK:
pip install followsm-sdkΒ· TypeScript:npm install @followsm/sdk - Related: hft-toxicity-circuit-breaker Β· polymarket-arbitrage-starter-kit
The triggers this backtest couldn't reach (1% book toxicity, depth imbalance, sweep detection and smart-money flow) are what FollowSM streams live.
| Free | DEVELOPER ($199/mo) | ENTERPRISE ($499/mo) | |
|---|---|---|---|
| REST requests/min | 30 (per IP) | 300 | 1,000 |
/ws/v1/confluence stream (the hedger's SIGNAL_SOURCE=stream) |
β | β | β |
π Get your API key at follow-sm.com/pricing
References: Easley, LΓ³pez de Prado & O'Hara (2012), "Flow Toxicity and Liquidity in a High-Frequency World", Review of Financial Studies. Andersen & Bondarenko (2014), "VPIN and the Flash Crash", Journal of Financial Markets.
Not financial advice. Prediction markets and perpetual futures carry substantial risk of loss.




Top comments (0)