DEV Community

Cover image for Detecting Orderbook Toxicity & Polymarket Divergence: Building a Microstructure Circuit Breaker
Andrew Rodn for FollowSM

Posted on

Detecting Orderbook Toxicity & Polymarket Divergence: Building a Microstructure Circuit Breaker

Packages:

On 22 August 2026, between 05:07 and 05:10 UTC, BTCUSDT on Binance fell from 78,528 to 76,742, a 1.8k drop in three minutes. The 05:10 minute alone printed $56.3M of notional. The ten minutes before it averaged under $1M per minute.

On Polymarket, the 15-minute "Bitcoin Up or Down, 05:00–05:15 UTC" market was still quoting Up at 0.635 in its 05:07 price-history point. One point later it was 0.085. By 05:10 it was 0.005.

Anyone resting maker quotes on either venue in that window gave free options to traders who could see the other one.

This article covers:

  1. Theory: VPIN, L1/L2 orderbook imbalance and the robust volume Z-score, with the maths behind each.
  2. The problem: how market makers get adverse-selected during news shocks and prediction-market sweeps.
  3. Implementation: a streaming circuit breaker in Python on followsm-sdk.
  4. Case study: the 22 August sweep, measured honestly. That includes the part where VPIN didn't see it coming.

📊 All numbers below come from public data (Binance klines and Polymarket's CLOB price history). The script that reproduces them is linked at the end.


1. Theory

1.1 VPIN: Volume-Synchronized Probability of Informed Trading

The classic PIN model (Easley, Kiefer, O'Hara & Paperman, 1996) treats order flow as a mix of two groups:

  • uninformed traders, arriving on both sides at rate ε;
  • informed traders, who show up with probability α on "information days" and trade one-sidedly at rate μ.

The probability that a given trade is informed is

PIN=αμαμ+2ε \text{PIN} = \frac{\alpha\mu}{\alpha\mu + 2\varepsilon}

Estimating PIN by maximum likelihood is slow and unstable at intraday frequency. VPIN (Easley, López de Prado & O'Hara, 2012) replaces the estimation with a volume-clock approximation.

  1. Cut the trade tape into buckets of equal traded volume V, not equal time. Volume clocks speed up exactly when information is arriving.
  2. Split each bucket τ into buyer-initiated volume V^B and seller-initiated volume V^S.
  3. Average the absolute imbalance over the last n buckets:
VPIN=1n∑τ=1n∣VτB−VτS∣V \text{VPIN} = \frac{1}{n}\sum_{\tau=1}^{n} \frac{\lvert V^B_\tau - V^S_\tau \rvert}{V}

Why this approximates PIN: the expected imbalance per bucket is about αμ, and the expected bucket volume is αμ + 2ε, so VPIN converges to the same ratio.

On equities, the original paper had to infer trade direction (bulk volume classification). Crypto venues give it to you for free: every Binance trade carries an isBuyerMaker flag, and every kline carries the exact taker-buy notional. There is no classification error.

The parameter that matters most is V. The original paper used V = 1/50 of average daily volume (ADV) and n = 50, so VPIN covers roughly one trading day.

  • V too small relative to the market's trade size: most buckets fill with one or two trades. Each such bucket is almost entirely one-sided, and VPIN drifts toward 1 regardless of information content.
  • V too large: VPIN barely moves inside an event.

Calibrate V per symbol. A fixed dollar bucket across BTC and a small-cap alt measures two different things.

Even calibrated, raw VPIN levels differ between pairs, because trade-size distributions differ. The robust signal is VPIN's rank within its own history, which is how the original paper reports it (as a CDF). FollowSM calibrates buckets per symbol and publishes that rank as vpin_percentile.

1.2 L1 and L2 orderbook imbalance

VPIN looks at executed flow. The book shows resting intent. At the top of the book (L1):

IL1=pbqbpbqb+paqa∈[0,1] I_{L1} = \frac{p_b q_b}{p_b q_b + p_a q_a} \in [0, 1]

where p_b, q_b are the best bid price and size and p_a, q_a the best ask. L1 is noisy and easy to spoof, so we also aggregate depth within a band of ±k around the mid m (the L2 view):

B±kbid=∑pi≥m(1−k)piqi,B±kask=∑pj≤m(1+k)pjqj B^{\pm k}{\text{bid}} = \sum{p_i \ge m(1-k)} p_i q_i, \qquad B^{\pm k}{\text{ask}} = \sum{p_j \le m(1+k)} p_j q_j
ParseError: KaTeX parse error: Expected 'EOF', got '_' at position 116: …\qquad \text{ob_̲toxicity}{1\%} …

An ob_toxicity_1pct above 2 means there is twice as much resting sell notional as buy notional within 1% of mid. The bid side is thin, and a market sell of modest size will walk through it.

What matters here is the change, not the level. Books are persistently skewed on many pairs. A circuit breaker should react to a sudden departure from the recent norm, for example:

∣imbalance1%(t)−EWMAλ(imbalance1%)∣>δ \lvert \text{imbalance}{1\%}(t) - \text{EWMA}\lambda(\text{imbalance}_{1\%}) \rvert > \delta

1.3 Robust volume Z-score

Volume is heavy-tailed, so a mean/standard-deviation Z-score is dominated by the very outliers you're trying to detect. Use the median and the median absolute deviation (MAD) instead:

zV=0.6745⋅Vt−median⁡(Vt−k..t−1)MAD⁡(Vt−k..t−1) z_V = 0.6745 \cdot \frac{V_t - \operatorname{median}(V_{t-k..t-1})}{\operatorname{MAD}(V_{t-k..t-1})}

The 0.6745 factor makes the MAD comparable to a standard deviation under normality. When the current candle is still open, pro-rate its volume by elapsed time before comparing, or the Z-score under-counts early in the window.

1.4 Cross-venue divergence with Polymarket

Crypto prediction markets carry an implied probability p for an outcome that is a direct function of spot. Examples are "Will BTC be above $82k on 26 September?" and "BTC Up or Down in the next 15 minutes".

Give each market a direction d: +1 if a YES is bullish for spot, −1 if bearish, 0 if ambiguous. The two venues diverge when:

sign⁡(ΔS15m)≠sign⁡(d⋅Δp15m)and∣Δp15m∣>0.05 \operatorname{sign}(\Delta S_{15m}) \ne \operatorname{sign}(d \cdot \Delta p_{15m}) \quad\text{and}\quad \lvert \Delta p_{15m} \rvert > 0.05

In words: spot is going one way and the prediction market is pricing the other. Either someone knows something, or one venue is about to be run over. Neither is a good time to be quoting size.


2. The problem: adverse selection in bursts

A market maker earns the spread on uninformed flow and loses on informed flow. Over a normal day, informed flow is a small share of fills and the spread covers it. The losses are not spread evenly, though. They arrive in bursts:

  • News shocks. A headline hits, the fastest participants lift or hit everything on the book, and resting quotes are filled at stale prices.
  • Liquidation cascades. Forced sellers are uninformed individually, but a cascade is predictable once it starts. Anyone who sees it early trades against your stale bids.
  • Prediction-market sweeps. Polymarket's short-dated crypto markets price off spot. When Binance moves, the trader who reacts first sweeps every Polymarket quote still priced off the old spot. On a 15-minute binary, "stale" can mean a 60-cent error on a $1 contract.

💡 The fix is not predicting the move; nobody reliably does that. The fix is not being there when it happens: pull or widen quotes within milliseconds of the flow turning toxic, and put them back once it normalises. That is a circuit breaker.


3. Implementation

pip install followsm-sdk
Enter fullscreen mode Exit fullscreen mode

followsm-sdk exposes FollowSM's Binance microstructure metrics and its Binance × Polymarket confluence snapshots as typed pydantic models. The main ones are SymbolToxicityMetrics, ConfluenceSnapshot and RiskConfig.

3.1 One snapshot (free tier, no key)

from followsm_sdk import FollowSMClient, RateLimitExceededException

client = FollowSMClient()  # no key: free, IP-rate-limited tier

try:
    snap = client.get_confluence_snapshot("BTCUSDT")
except RateLimitExceededException as exc:
    raise SystemExit(f"{exc}\nMore headroom: https://follow-sm.com/pricing")

micro = snap.binance_microstructure
print(f"vpin={micro.vpin:.3f} ob_toxicity_1pct={micro.ob_toxicity_1pct:.2f} "
      f"spot_15m={micro.price_delta_15m_pct:+.3%}")

for event in snap.polymarket_confluence.active_events:
    print(f"{event.market_slug}: p={event.implied_probability:.3f} "
          f"prob_delta_15m={event.prob_delta_15m:+.3f} direction={event.direction}")

print("divergence:", snap.composite_signals.cross_market_divergence_flag)
print("backend action:", snap.composite_signals.recommended_action)
Enter fullscreen mode Exit fullscreen mode

3.2 Your own risk ladder

The backend returns a conservative recommended_action. You can re-derive it locally with your own thresholds:

from followsm_sdk import FollowSMClient, RiskConfig

risk = RiskConfig(
    vpin_percentile_widen_threshold=0.85,  # -> WIDEN_SPREAD_2X
    vpin_percentile_halt_threshold=0.97,   # with divergence -> HALT_MAKER_QUOTES
    ob_toxicity_threshold=2.5,             # a lopsided 1% book also counts as toxic
    min_semantic_confidence=0.70,          # never HALT on an ambiguously mapped market
)
client = FollowSMClient(risk_config=risk)
snap = client.get_confluence_snapshot("BTCUSDT")
print(client.evaluate_risk(snap))  # NONE | WIDEN_SPREAD_1_5X | WIDEN_SPREAD_2X | HALT_MAKER_QUOTES
Enter fullscreen mode Exit fullscreen mode

The ladder in full. "Toxic" means vpin_percentile at or above the threshold, or ob_toxicity_1pct above ob_toxicity_threshold:

Condition Action
vpin_percentile ≥ vpin_percentile_halt_threshold (or a toxic book) and divergence HALT_MAKER_QUOTES (downgraded to 1.5x if direction_confidence < min_semantic_confidence)
vpin_percentile ≥ vpin_percentile_widen_threshold, or a toxic book WIDEN_SPREAD_2X
divergence alone WIDEN_SPREAD_1_5X
otherwise NONE

3.3 A streaming circuit breaker

Polling is fine for a Polymarket bot quoting every few seconds. A market maker needs ticks. The Enterprise stream /ws/v1/toxicity pushes every symbol's SymbolToxicityMetrics from sub-10ms in-memory snapshots:

import asyncio
import os

from followsm_sdk import FollowSMClient

PCTL_TRIP, PCTL_REARM, OB_TOX_TRIP = 0.90, 0.80, 2.0


async def cancel_quotes(symbol: str) -> None:
    print(f"KILL SWITCH {symbol}: cancelling resting maker quotes")  # call your OMS here


async def main() -> None:
    client = FollowSMClient(api_key=os.environ["FOLLOWSM_API_KEY"])  # ENTERPRISE key
    halted = set()
    async for m in client.stream_toxicity():
        pctl = m.vpin_percentile  # VPIN ranked vs this symbol's own history; None while warming up
        toxic = (pctl is not None and pctl > PCTL_TRIP) or m.ob_toxicity_1pct > OB_TOX_TRIP
        if toxic and m.symbol not in halted:
            halted.add(m.symbol)
            await cancel_quotes(m.symbol)
        elif m.symbol in halted and (pctl is None or pctl < PCTL_REARM) and m.ob_toxicity_1pct <= OB_TOX_TRIP:
            halted.discard(m.symbol)
            print(f"RE-ARM {m.symbol}")


asyncio.run(main())
Enter fullscreen mode Exit fullscreen mode

The gap between PCTL_TRIP and PCTL_REARM is hysteresis. Without it, a VPIN percentile oscillating around 0.90 would cancel and re-post your quotes on every tick.

The production version adds the rest of what a real desk needs:

HFT Toxicity Circuit Breaker

PyPI Open In Colab Get an API key License: MIT

A kill switch for market makers, built on FollowSM's Enterprise WebSocket streams (/ws/v1/toxicity and /ws/v1/confluence).

It evaluates every Binance microstructure tick in memory. When order flow turns toxic, it fires cancellation callbacks and signed webhooks for your resting maker quotes, before informed traders pick them off.

12:01:42 WARNING | BTCUSDT NONE -> HALT_MAKER_QUOTES [toxicity] ob_toxicity_1pct 6.69 > 2.00 | vpin=0.186 pctl=- ob_tox_1pct=6.69 eval=32.1us feed_age=561ms
12:01:42 WARNING | KILL SWITCH: cancelling every resting maker quote on BTCUSDT
12:01:42 WARNING | ETHUSDT NONE -> HALT_MAKER_QUOTES [toxicity] ob_toxicity_1pct 4.99 > 2.00 | vpin=0.350 pctl=- ob_tox_1pct=4.99 eval=11.9us feed_age=561ms
11:30:47 WARNING | SOLUSDT WIDEN_SPREAD_2X -> HALT_MAKER_QUOTES [toxicity] ob_toxicity_1pct 2.90 > 2.00; 1% depth imbalance spike 0.26 vs ewma 0.48
Breaker evaluation latency: n=240 p50=7.1us p99=27.3us max=38.0us

Real output from the live Enterprise streams (pctl=-: vpin_percentile was still warming up).


Why

Market makers lose money through adverse selection…

On top of the loop above, it adds:

  • Depth-imbalance spike detection (EWMA deviation of the 1% band, §1.2).
  • A second leg on /ws/v1/confluence, evaluated locally with evaluate_risk_action(snapshot, RiskConfig(...)), so the effective action per symbol is the more severe of the two legs.
  • A cooldown before re-arming.
  • Fail-closed reconnects. A dropped stream halts every symbol it guards until fresh frames arrive.
  • Non-blocking callbacks. Cancellations and HMAC-signed webhooks are scheduled as asyncio tasks, so a slow webhook never delays the next tick.

Evaluation stays in-process and synchronous. On live FollowSM data it measured p50 = 4.7 µs, p99 = 24.9 µs per frame on a laptop. The latency budget belongs to the network, not the decision.

For the Polymarket side, the starter kit runs the same ladder on REST polling. It widens or halts its CLOB (central limit order book) quotes, and skews them toward the spot-implied fair value:

center=mid+clip⁡(k⋅100 ΔS15m⋅d−Δp15m, ±0.05) \text{center} = \text{mid} + \operatorname{clip}\big(k \cdot 100\,\Delta S_{15m} \cdot d - \Delta p_{15m},\ \pm 0.05\big)

Polymarket Arbitrage Starter Kit

PyPI Open In Colab Get an API key License: MIT

An asynchronous Python market-making bot for Polymarket CLOB that uses Binance spot microstructure as its risk signal. It is built on followsm-sdk.

Binance spot usually reprices before crypto prediction markets do. This bot uses that lead in two ways:

  1. Defence. It pulls or widens its Polymarket maker quotes when Binance order flow turns toxic (high VPIN, one-sided 1% depth) or when the two venues disagree.
  2. Offence. It skews its quotes toward the spot-implied fair value before the prediction market catches up.
2026-09-26 10:53:07 INFO bot | BTCUSDT vpin=0.697 ob_tox_1pct=2.61 prob_delta_15m=+0.003 spot_15m=+0.072% -> WIDEN_SPREAD_2X (backend: WIDEN_SPREAD_2X)
2026-09-26 10:53:07 INFO bot | Quoting bitcoin-above-82k-on-september-26-2026: mid=0.994 bid=0.97 ask=0.99 (2.0x spread)
2026-09-26 10:53:07 INFO bot.execution | [PAPER] BUY 10.00 YES @ 0.97 (token 8168161214…)
2026-09-26 10:53:07 INFO bot.execution | [PAPER] SELL 10.00 YES @ 0.99 (token 8168161214…)

How it works

flowchart LR
    A[FollowSM API<br/>get_confluence_snapshot] -->|every POLL_INTERVAL_SECS| B[market_monitor.py<br/>MarketSignal]
    B -->
…

4. Case study: the 22 August 2026 BTC sweep

Setup. 86,400 one-minute BTCUSDT klines from Binance (28 July – 26 September 2026), with the exact taker-buy notional per bar. I computed VPIN two ways, both with n = 50:

  • slow: V = ADV/50 ≈ $23.3M, the original paper's calibration;
  • fast: V = $1M, a desk-style short horizon.

I then took the largest 15-minute drop in the sample after a 7-day warm-up, plus Polymarket's public CLOB price history for the matching 15-minute Up or Down market.

4.1 What happened

UTC Close Notional Taker-buy share VPIN fast (7d pctl) VPIN slow Polymarket "Up"
04:55 78,541 $0.5M 0.72 0.236 (23%) 0.130 –
05:00 78,532 $1.0M 0.69 0.228 (21%) 0.130 0.395
05:03 78,544 $1.2M 0.48 0.211 (16%) 0.130 0.695
05:06 78,528 $0.6M 0.66 0.200 (14%) 0.130 0.545
05:07 78,286 $2.9M 0.17 0.210 (16%) 0.130 0.635
05:08 78,138 $3.1M 0.37 0.207 (16%) 0.130 0.085
05:09 77,802 $10.9M 0.40 0.200 (14%) 0.130 0.035
05:10 76,742 $56.3M 0.34 0.326 (63%) 0.140 0.005
05:11 77,087 $43.4M 0.49 0.062 (0%) 0.142 0.001

Over the full 15 minutes, $84M traded, a robust volume Z-score of 22.8 against the prior five hours. The "Up" market resolves Up if BTC finishes the 05:00–05:15 window above where it started.

4.2 What the data says, and what it doesn't

VPIN did not predict this sweep. At the start of the window, slow VPIN was 0.130, the 4th percentile of its trailing week. Fast VPIN was at the 23rd percentile. Neither moved until the crash minute itself, and fast VPIN then collapsed as $43M of two-sided volume poured in during the rebound.

⚠️ This matches the academic record. VPIN's headline result, a rising reading hours before the 2010 Flash Crash, was contested by Andersen & Bondarenko (2014), who showed that much of the effect depended on the trade-classification scheme. Anyone selling VPIN as a crystal ball is selling something.

The conditional edge is weak. Across 5,087 non-overlapping 15-minute windows:

  • With slow VPIN in its top decile, the chance that the next 15 minutes contained a move in the top 1% of sizes (≥ 0.70%) was 1.37%, against 0.95% otherwise: a 1.4x lift. That rests on roughly five events, so treat it as a hint, not a result.
  • With fast buckets, the relationship inverted (0.29% vs 1.09%, 0.3x). Short-horizon VPIN on 1-minute bars mostly flags quiet, one-sided drift, not the start of a cascade.

What did move first was one-sided flow and a volume burst. At 05:07, 83% of taker notional was selling, on three times the notional of the preceding minutes. That is the fingerprint a tick-level feed is built to catch: per-trade buckets, 1% book toxicity and depth-imbalance spikes. One-minute bars blur it. Klines also can't show the bid side thinning within 1% of mid, which is exactly what ob_toxicity_1pct measures.

Polymarket repriced about one sample later. The "Up" market's 05:07 point still read 0.635, even though the 05:07 minute closed below the window's 05:00 opening price. The 05:08 point read 0.085. At one-minute fidelity I can't resolve the lead more finely than "within about a minute". On a binary contract that lag is enough: anyone still bidding Up near 0.60 at 05:07 sold a contract worth about 0.05 to a seller watching Binance.

4.3 Lessons for the breaker

  1. Don't trip on VPIN alone. Use it as a regime input that widens spreads, and let fast signals (taker imbalance, 1% book toxicity, depth-imbalance spikes, volume bursts) trigger the halt. That is why the breaker above ORs several conditions and the ladder reserves HALT_MAKER_QUOTES for VPIN plus divergence.
  2. Calibrate V per symbol, and set thresholds from each symbol's own distribution. A VPIN of 0.70 means something different on BTC than on a thin alt, and something different again with different bucket sizes.
  3. Latency beats prediction. In this sweep, the useful decision window between the first one-sided minute and the Polymarket collapse was about a minute. For the Binance book itself it was seconds. A breaker that reacts in microseconds to a streamed tick is worth more than a model that claims to forecast.
  4. Watch the other venue. Cross-venue divergence (spot down while a bullish-if-YES market still prices high) is the cheapest adverse-selection signal available to a Polymarket market maker. It doesn't require predicting anything, only noticing that two prices for the same risk disagree.

▶ Reproduce the case study (no API keys needed)
git clone https://github.com/Follow-SM/hft-toxicity-circuit-breaker.git
cd hft-toxicity-circuit-breaker
pip install httpx && python research/vpin_case_study.py
Enter fullscreen mode Exit fullscreen mode

The script uses only public Binance klines and Polymarket's public CLOB price history.


Get started

Free DEVELOPER ($199/mo) ENTERPRISE ($499/mo)
REST requests/min 30 (per IP) 300 1,000
Binance pairs Limited 50+ 50+
Snapshot latency Standard Sub-10ms in-memory Sub-10ms in-memory
/ws/v1/toxicity + /ws/v1/confluence ❌ ❌ ✅

👉 Get your API key at follow-sm.com/pricing


References: Easley, Kiefer, O'Hara & Paperman (1996), "Liquidity, Information, and Infrequently Traded Stocks", Journal of Finance. Easley, López de Prado & O'Hara (2012), "Flow Toxicity and Liquidity in a High-Frequency World", Review of Financial Studies. Andersen & Bondarenko (2014), "VPIN and the Flash Crash", Journal of Financial Markets.

Not financial advice. Trading crypto and prediction markets carries substantial risk of loss.

Top comments (0)