DEV Community

Cover image for Polymarket Trading Bot Incident Recovery: How to Reconstruct What Happened

Polymarket Trading Bot Incident Recovery: How to Reconstruct What Happened

A Polymarket trading bot stops.

Maybe a WebSocket connection dropped.

Maybe an order behaved unexpectedly.

Maybe the position no longer matches local state.

Maybe the process restarted.

The immediate reaction is usually:

restart the bot

But restarting only gets the process running again.

It doesn't answer the more important questions:

What happened before the incident?

What state did the bot believe it was in?

What actually happened to the orders and fills?

Which events were missed?

What is the current position?

Is the system safe to trade again?

That's the part I'm interested in.

For a production-oriented trading system, incident recovery isn't just restarting a process. It's reconstructing enough state to understand what happened and verify that the system can safely continue.


A restart is not a recovery

A simple recovery flow might be:

Process crashed
    ↓
Restart
    ↓
Trading resumes
Enter fullscreen mode Exit fullscreen mode

That assumes the state before the crash was correct.

It also assumes nothing important happened while the process was unavailable.

Neither assumption is safe.

A better flow is:

Incident
    ↓
Pause Trading
    ↓
Recover State
    ↓
Reconcile
    ↓
Verify
    ↓
Check Risk
    ↓
Resume
Enter fullscreen mode Exit fullscreen mode

The restart is only one step.


The first problem is figuring out what happened

Imagine the bot was running normally.

Then something goes wrong.

When the process comes back, you may have:

Last known local state
+
Persisted state
+
Orders
+
Fills
+
Market events
+
Position data
+
Risk state
Enter fullscreen mode Exit fullscreen mode

These don't necessarily line up.

You might discover:

Local position:   +100
Remote position:   +40
Enter fullscreen mode Exit fullscreen mode

Or:

Order:
FILLED

Transaction:
UNKNOWN
Enter fullscreen mode Exit fullscreen mode

Or:

Connection:
RECONNECTED

State:
UNVERIFIED
Enter fullscreen mode Exit fullscreen mode

The incident isn't resolved just because the process is running again.


Start with the last known good state

A useful recovery process needs a reference point.

For example:

Last Known Good State
        ↓
Incident
        ↓
Events around incident
        ↓
Current external state
Enter fullscreen mode Exit fullscreen mode

The system should be able to answer:

When was the last verified state?

What orders were active?

What had already filled?

What positions were expected?

What was the exposure?

What was the risk state?
Enter fullscreen mode Exit fullscreen mode

Without that history, recovery becomes guesswork.


Why event history matters

Suppose the system saw:

Order A
Fill A
Fill B
Enter fullscreen mode Exit fullscreen mode

Then the process crashed.

When it restarts, the local database may only contain part of the state.

The recovery system needs to establish:

What was known before the incident?
What happened during the gap?
What is true now?
Enter fullscreen mode Exit fullscreen mode

That means keeping enough structured information to reconstruct state.

Not every raw event has to be stored forever.

But important state transitions should be traceable.


The trading bot can be alive while its state is wrong

This is similar to the problem I covered in Polymarket WebSocket reconnects.

Imagine:

Process:
RUNNING

WebSocket:
CONNECTED

Strategy:
ACTIVE
Enter fullscreen mode Exit fullscreen mode

Looks healthy.

But:

Position:
UNVERIFIED
Enter fullscreen mode Exit fullscreen mode

Now the system has a much more serious problem.

The process is healthy.

The trading state is not.

This distinction is important during incident recovery because the recovery process needs to establish system state, not just process state.


Reconstructing an incident

A useful incident-recovery pipeline looks like:

Incident Detected
      ↓
Freeze New Risk
      ↓
Capture Current State
      ↓
Load Last Known State
      ↓
Inspect Orders / Fills
      ↓
Inspect Execution State
      ↓
Read Remote Position
      ↓
  Compare
      ↓
Identify Mismatches
      ↓
    Repair
      ↓
   Verify
      ↓
Risk Check
      ↓
   Resume
Enter fullscreen mode Exit fullscreen mode

The exact implementation will vary.

The principle is the important part:

don't resume trading before you understand the state you are resuming from.


Freeze new risk first

The first operational action should normally be to stop adding new exposure.

For example:

Incident
   ↓
TRADING PAUSED
Enter fullscreen mode Exit fullscreen mode

Existing state can still be inspected.

Orders can be reconciled.

Positions can be checked.

But the system shouldn't keep creating additional risk while it is trying to figure out what happened.

This is one of the reasons risk controls and incident recovery belong together.


Reconcile orders first

The system needs to understand the order state around the incident.

For example:

Order 1 → Filled
Order 2 → Open
Order 3 → Cancelled
Order 4 → Unknown
Enter fullscreen mode Exit fullscreen mode

An order that is UNKNOWN needs different handling from one that is clearly cancelled.

Likewise, an open order may still create future exposure.

So recovery needs to establish:

Known orders
+
Known fills
+
Unknown orders
+
Expected remaining quantities
Enter fullscreen mode Exit fullscreen mode

before rebuilding the rest of the state.


Then reconstruct execution state

This connects directly to the Polymarket Execution Verifier.

An incident can leave an execution somewhere in the middle of its lifecycle:

SUBMITTED
    ↓
MATCHED
    ↓
FILLED
    ↓
TX_PENDING
Enter fullscreen mode Exit fullscreen mode

The process stops.

When it comes back, the application shouldn't assume the final state.

It should recover the execution information and determine whether the trade is:

VERIFIED
UNKNOWN
INCONSISTENT
FAILED
Enter fullscreen mode Exit fullscreen mode

The important part is that recovery uses evidence rather than assumptions.


Partial fills make recovery harder

Suppose the bot intended:

100
Enter fullscreen mode Exit fullscreen mode

and before the incident it knew:

40 filled
60 remaining
Enter fullscreen mode Exit fullscreen mode

Then the process goes down.

After restarting, several things could be true.

The remaining 60 may still be open.

They may have filled.

They may have been cancelled.

The application may not know yet.

So:

100 requested
40 previously known
Enter fullscreen mode Exit fullscreen mode

does not automatically mean:

40 final
Enter fullscreen mode Exit fullscreen mode

The recovery process has to reconcile the current order and fill state.

This is one reason partial fills and incident recovery are closely related but distinct problems.


Position reconstruction comes next

After resolving orders and executions, the system needs to establish the current position.

For example:

Expected:
+100

Observed:
+40
Enter fullscreen mode Exit fullscreen mode

That's a mismatch.

The system should not simply overwrite one number with the other and continue.

It needs to know why the difference exists.

Possible causes include:

Missed fill
Partial execution
Missed WebSocket event
Unexpected order state
Stale local state
Restart during execution
Enter fullscreen mode Exit fullscreen mode

The goal is to move from:

UNKNOWN / INCONSISTENT
Enter fullscreen mode Exit fullscreen mode

to:

VERIFIED
Enter fullscreen mode Exit fullscreen mode

Incident recovery is different from reconciliation

These two concepts are connected, but I don't think they should be treated as the same thing.

Reconciliation

Asks:

Does my local state match the external state?

Local
  ↓
Remote
  ↓
Compare
  ↓
Repair
Enter fullscreen mode Exit fullscreen mode

Incident recovery

Asks:

What happened, what state was lost or made uncertain, and what do I need to verify before the system can operate again?

Incident
  ↓
Reconstruct
  ↓
Reconcile
  ↓
Verify
  ↓
Recover
Enter fullscreen mode Exit fullscreen mode

Reconciliation is part of recovery.

Recovery is the larger process.


Recovery needs a timeline

One of the most useful things an incident system can provide is a timeline.

For example:

10:41:08  Order submitted
10:41:09  Fill observed
10:41:10  WebSocket disconnected
10:41:11  Execution state uncertain
10:41:18  Process stopped
10:42:03  Process restarted
10:42:04  Trading paused
10:42:05  Reconciliation started
10:42:06  Position mismatch detected
10:42:07  Remote state fetched
10:42:08  Position verified
10:42:09  Risk checks passed
10:42:10  Trading resumed
Enter fullscreen mode Exit fullscreen mode

Now the operator can understand what happened.

Without this, the incident may look like:

Bot crashed
Bot restarted
Enter fullscreen mode Exit fullscreen mode

That's not enough information.


The control plane should record the recovery state

This is where the Polymarket Trading Control Plane fits into the larger architecture.

The system can move through explicit states:

HEALTHY
    ↓
DEGRADED
    ↓
PAUSED
    ↓
RECOVERING
    ↓
 HEALTHY
Enter fullscreen mode Exit fullscreen mode

The important part is that RECOVERING is an actual system state.

It tells the rest of the system:

The process is running, but normal trading has not been restored yet.

That is much clearer than a generic “online” status.


Recovery should have clear stages

A useful model is:

1. Detect
2. Pause
3. Capture
4. Reconstruct
5. Reconcile
6. Verify
7. Check Risk
8. Resume
Enter fullscreen mode Exit fullscreen mode

Each stage should produce observable state.

For example:

RECOVERY_STARTED
RECONCILIATION_STARTED
STATE_MISMATCH
RECONCILIATION_COMPLETED
RISK_CHECK_COMPLETED
RECOVERY_COMPLETED
TRADING_RESUMED
Enter fullscreen mode Exit fullscreen mode

That turns recovery from a collection of retry loops into an operational process.


What happens when reconstruction fails?

Sometimes the system won't be able to establish a clean state immediately.

For example:

Local:
+100

Remote:
UNKNOWN
Enter fullscreen mode Exit fullscreen mode

Or:

Execution:
UNKNOWN

Position:
UNKNOWN
Enter fullscreen mode Exit fullscreen mode

The system should not invent an answer.

It should remain controlled:

RECOVERING
    ↓
PAUSED
Enter fullscreen mode Exit fullscreen mode

and retry or request further reconciliation.

This is another reason I prefer explicit UNKNOWN states.

Unknown state is uncomfortable.

Pretending unknown state is healthy is worse.


Recovery after application restart

A restart deserves the same treatment.

The process starts:

Application
    ↓
Load persisted state
Enter fullscreen mode Exit fullscreen mode

But that state may be stale.

So:

Load Local State
       ↓
   Connect
       ↓
Read Remote State
       ↓
   Compare
       ↓
    Repair
       ↓
    Verify
       ↓
Risk Check
       ↓
Allow Trading
Enter fullscreen mode Exit fullscreen mode

Starting the application and starting trading are therefore separate operations.


Recovery after a WebSocket gap

The same applies when a connection drops.

Suppose:

Event A
Event B
   ↓
DISCONNECT
   ↓
Events C, D
   ↓
RECONNECT
Enter fullscreen mode Exit fullscreen mode

The process may not have seen C and D.

The recovery path has to establish the current state rather than assuming the missing events can be ignored.

That's why the sequence is:

Disconnect
   ↓
Pause
   ↓
Reconnect
   ↓
Reconcile
   ↓
Verify
   ↓
Risk Check
   ↓
Resume
Enter fullscreen mode Exit fullscreen mode

The connection returning is only one part of the recovery.


Recovery should be idempotent

This is important for automated systems.

Suppose the recovery worker gets triggered twice.

You don't want:

Recovery #1
Recovery #2
Enter fullscreen mode Exit fullscreen mode

to produce conflicting state changes.

The operations should be safe to repeat:

RECONCILE
RECONCILE
RECONCILE
Enter fullscreen mode Exit fullscreen mode

without double-counting fills or creating duplicate records.

Likewise:

PAUSE
PAUSE
Enter fullscreen mode Exit fullscreen mode

should remain simply:

PAUSED
Enter fullscreen mode Exit fullscreen mode

The recovery process should be designed around idempotent operations wherever possible.


Incident history is useful after the system is healthy

Recovery isn't only about getting back online.

The incident record can help answer:

What failed?

How long was trading paused?

Which state became inconsistent?

How was it repaired?

Did risk limits change?

What prevented the system from resuming immediately?
Enter fullscreen mode Exit fullscreen mode

Over time, that becomes useful for identifying recurring problems.

For example:

WebSocket incidents
Execution failures
Position mismatches
Stale-data incidents
Recovery failures
Enter fullscreen mode Exit fullscreen mode

You can start seeing patterns instead of isolated outages.


What I want to see from an incident record

Something like:

Incident ID:
INC-2026-0012

Started:
10:41:10

Trigger:
WEBSOCKET_DISCONNECTED

System State:
PAUSED

Affected Execution:
abc123

Position:
MISMATCH

Recovery:
COMPLETED

Duration:
62 seconds

Final Risk State:
HEALTHY

Trading:
RESUMED
Enter fullscreen mode Exit fullscreen mode

The exact implementation doesn't need to look like this.

But the system should retain enough information to explain the recovery.


Don't hide recovery inside logs

Logs are useful.

But recovery state should also be part of the application's domain model.

Instead of only:

ERROR websocket disconnected
Enter fullscreen mode Exit fullscreen mode

the system should know:

system_state = PAUSED
recovery_state = RECONCILING
trading_permission = BLOCK
Enter fullscreen mode Exit fullscreen mode

Now dashboards, alerts, APIs, and automation can use the same source of truth.

That's one of the main reasons I'm separating the Control Plane from the trading strategy.


What should make trading resume?

This is the final gate.

I don't want:

process restarted
    ↓
resume
Enter fullscreen mode Exit fullscreen mode

I want something closer to:

Incident resolved
+
Orders reconciled
+
Executions verified
+
Positions verified
+
Exposure recalculated
+
Risk healthy
+
Market data fresh
=
Trading allowed
Enter fullscreen mode Exit fullscreen mode

If one of those critical conditions remains unknown:

PAUSED
Enter fullscreen mode Exit fullscreen mode

until the state is established.


The broader architecture

The incident-recovery layer fits into the stack I've been building:

Polymarket Trading Bot
        ↓
    Strategy
        ↓
      Risk
        ↓
    Execution
        ↓
Execution Verifier
        ↓
Position Reconciliation
        ↓
Trading Control Plane
        ↓
Incident Recovery
Enter fullscreen mode Exit fullscreen mode

The responsibilities stay separate.

The strategy generates decisions.

Risk controls those decisions.

Execution carries them out.

The verifier determines what happened.

Reconciliation checks whether state matches reality.

The control plane coordinates health and trading permission.

Incident recovery reconstructs and restores the system after something goes wrong.


A practical recovery checklist

Before allowing a bot to resume after an incident:

[ ] Incident trigger identified
[ ] New trading paused
[ ] Last known state recovered
[ ] Orders reconciled
[ ] Fills reconciled
[ ] Execution state verified
[ ] Position verified
[ ] Exposure recalculated
[ ] Risk checked
[ ] Market data confirmed fresh
[ ] No critical UNKNOWN state remains
[ ] Recovery completed
[ ] Trading permission = ALLOW
Enter fullscreen mode Exit fullscreen mode

The exact policy will depend on the system.

But the important part is having an explicit gate.


Final takeaway

A failed trading bot is not necessarily the hard part.

The harder problem is restarting it without knowing what happened before the restart.

A process can start successfully while:

orders are uncertain
fills are incomplete
positions are wrong
exposure is stale
risk is unknown
Enter fullscreen mode Exit fullscreen mode

That's why I think incident recovery should be treated as part of the trading system itself.

The recovery path I want is:

INCIDENT
   ↓
PAUSE
   ↓
RECONSTRUCT
   ↓
RECONCILE
   ↓
VERIFY
   ↓
CHECK RISK
   ↓
RESUME
Enter fullscreen mode Exit fullscreen mode

The goal isn't simply to get the bot running again.

It's to get the system back to a state where it knows what happened, knows what is true now, and has a reason to trust the next trade.

Top comments (0)