DEV Community

Cover image for Polymarket Trading Bot Execution Analytics: What Should You Measure After the Trade?

Polymarket Trading Bot Execution Analytics: What Should You Measure After the Trade?

A Polymarket trading bot can be running normally and still have problems you won't notice from the P&L.

Orders may be taking longer to execute.

Partial fills may be increasing.

Rejections may be happening more often.

Execution state may be staying unresolved longer than expected.

Position reconciliation may be running too frequently.

A recovery that normally takes a few seconds may suddenly take a minute.

None of that is obvious from:

Bot: RUNNING
P&L: +$...
Enter fullscreen mode Exit fullscreen mode

For an automated trading system, I want to know not only whether it traded, but how the execution layer is behaving.

That is where execution analytics becomes useful.

The idea is simple:

Execution
   ↓
Metrics
   ↓
Health / Anomalies
   ↓
Risk / Control
   ↓
Action
Enter fullscreen mode Exit fullscreen mode

I'm treating this as another layer around the Polymarket trading infrastructure I'm building.


What should you measure?

There isn't one universal list.

The useful metrics depend on the strategy and the execution model.

But for a production-oriented Polymarket bot, I would start with:

order latency
fill latency
fill ratio
partial-fill rate
rejection rate
execution failures
unknown execution states
position mismatches
reconciliation frequency
recovery duration
Enter fullscreen mode Exit fullscreen mode

The important part is not collecting dozens of metrics.

It's being able to connect a metric to an operational decision.

For example:

rejection rate ↑
       ↓
execution health ↓
       ↓
investigate / pause
Enter fullscreen mode Exit fullscreen mode

That's much more useful than a dashboard full of numbers nobody acts on.


Order latency

The first useful measurement is how long an order takes to move through the execution path.

For example:

strategy decision
      ↓
order created
      ↓
order submitted
Enter fullscreen mode Exit fullscreen mode

Measure:

decision → order created
order created → submitted
Enter fullscreen mode Exit fullscreen mode

You can then start looking at the distribution rather than a single average.

For example:

p50:  120 ms
p95:  480 ms
p99:  910 ms
Enter fullscreen mode Exit fullscreen mode

The exact numbers aren't important here.

The point is that averages can hide slow executions.

A bot that normally processes most orders quickly but occasionally takes much longer can have a very different operational profile from one with consistently stable latency.


Fill latency

Order submission is not the same thing as execution.

So the next measurement is:

order submitted
      ↓
fill observed
Enter fullscreen mode Exit fullscreen mode

Track the elapsed time.

For example:

Order A → 140 ms
Order B → 165 ms
Order C → 3.8 s
Enter fullscreen mode Exit fullscreen mode

That third observation deserves attention even if the order eventually completed successfully.

Over time, fill latency can be tracked by:

  • market
  • side
  • order type
  • strategy
  • time window

This makes it possible to find patterns instead of treating every slow execution as an isolated event.


Fill ratio

Suppose the strategy requested 100 units.

The execution produced 75.

That's:

requested = 100
matched   = 75
Enter fullscreen mode Exit fullscreen mode

The fill ratio is:

75 / 100 = 75%
Enter fullscreen mode Exit fullscreen mode

This is more informative than simply recording:

FILLED
Enter fullscreen mode Exit fullscreen mode

because requested quantity and actual execution are different things.

A system that starts seeing lower fill ratios may need a different execution policy, more inventory planning, or simply a closer look at the market conditions.

The metric by itself doesn't tell you what to do.

It tells you where to look.


Partial-fill rate

Fill ratio and partial-fill rate answer different questions.

Fill ratio asks:

How much of the requested quantity was actually executed?

Partial-fill rate asks:

How often are orders ending up partially executed?

For example:

100 orders
30 had partial execution
Enter fullscreen mode Exit fullscreen mode

That gives:

partial-fill rate = 30%
Enter fullscreen mode Exit fullscreen mode

This is useful because partial fills can create downstream work:

partial fill
   ↓
remaining quantity
   ↓
position update
   ↓
exposure update
   ↓
possible reconciliation
Enter fullscreen mode Exit fullscreen mode

This connects directly to my earlier Polymarket partial fills work, where requested quantity and actual execution have to remain separate.

The interesting metric isn't simply “how many orders were partial.”

It's what those partial fills cause elsewhere in the system.


Rejection rate

Another straightforward metric is:

rejected orders
----------------
submitted orders
Enter fullscreen mode Exit fullscreen mode

For example:

1,000 submitted
25 rejected
Enter fullscreen mode Exit fullscreen mode

A 2.5% rejection rate means something very different from a system where almost every order succeeds.

More importantly, track changes over time.

For example:

Monday    0.8%
Tuesday   0.9%
Wednesday 1.1%
Thursday  4.7%
Enter fullscreen mode Exit fullscreen mode

The Thursday number is where I'd start investigating.

Possible causes need to be established from the actual execution data rather than guessed from the metric alone.


Execution failures

Not every execution issue is a rejection.

You can also have failures around:

submission
tracking
confirmation
state updates
reconciliation
recovery
Enter fullscreen mode Exit fullscreen mode

So I would separate execution failures into useful categories rather than using one giant counter.

For example:

ORDER_SUBMISSION_FAILURE
FILL_PROCESSING_FAILURE
TRANSACTION_VERIFICATION_FAILURE
POSITION_RECONCILIATION_FAILURE
RECOVERY_FAILURE
Enter fullscreen mode Exit fullscreen mode

That makes the metric much more actionable.

If recovery failures are increasing while order submission remains healthy, you have a very different problem from a system where order submission itself is failing.


Unknown execution states

This is one of the metrics I care about most.

A trading system will sometimes encounter states it can't verify immediately.

For example:

FILLED
   ↓
TX_PENDING
Enter fullscreen mode Exit fullscreen mode

or:

execution = UNKNOWN
Enter fullscreen mode Exit fullscreen mode

The important metric is not just the count.

Track how long executions remain unresolved.

For example:

UNKNOWN executions: 7

oldest unresolved:
18.4 seconds
Enter fullscreen mode Exit fullscreen mode

Now the control plane has something meaningful to evaluate.

You can define a policy around unresolved execution state rather than treating UNKNOWN as an invisible edge case.

This connects directly to the Polymarket Execution Verifier.


Position mismatches

Execution analytics shouldn't stop at the execution layer.

Suppose the system believes:

position = 100
Enter fullscreen mode Exit fullscreen mode

while the external state is:

position = 40
Enter fullscreen mode Exit fullscreen mode

Now you have:

position mismatch = -60
Enter fullscreen mode Exit fullscreen mode

Track:

mismatch count
mismatch duration
markets affected
largest difference
time to resolution
Enter fullscreen mode Exit fullscreen mode

This is directly related to my earlier Polymarket position reconciliation work.

I haven't put a guessed URL into that link because the exact published page URL should come from the live DEV page rather than being invented.

The important analytics question is:

How often does the system disagree with the account, and how long does it take to become consistent again?


Reconciliation frequency

Reconciliation itself is something worth measuring.

For example:

automatic reconciliations
manual reconciliations
reconciliations after reconnect
reconciliations after restart
reconciliations caused by mismatch
Enter fullscreen mode Exit fullscreen mode

Then:

reconciliation count
+
reconciliation duration
+
reconciliation result
Enter fullscreen mode Exit fullscreen mode

gives you a much better picture of system stability.

If reconciliation starts running constantly, that can be a signal that another part of the system needs attention.

Again, the metric doesn't explain the cause.

It tells you where to investigate.


Recovery duration

Suppose a WebSocket disconnect occurs.

The system goes through:

disconnect
   ↓
pause
   ↓
reconnect
   ↓
reconcile
   ↓
verify
   ↓
risk check
   ↓
resume
Enter fullscreen mode Exit fullscreen mode

Measure the total time.

For example:

Recovery #1 → 2.4 s
Recovery #2 → 3.1 s
Recovery #3 → 18.7 s
Enter fullscreen mode Exit fullscreen mode

That third recovery may deserve investigation.

You can also break the duration into stages:

disconnect → reconnect
reconnect → reconciliation
reconciliation → verification
verification → risk check
risk check → resume
Enter fullscreen mode Exit fullscreen mode

Now you can see where the recovery process is actually spending time.

This builds on the Polymarket WebSocket recovery work from earlier in the cluster.


Metrics need context

A single number is rarely enough.

Suppose you see:

fill ratio = 72%
Enter fullscreen mode Exit fullscreen mode

Is that good?

You can't answer that without context.

Compare:

Market A: 72%
Market B: 41%
Enter fullscreen mode Exit fullscreen mode

or:

normal: 74%
today: 72%
Enter fullscreen mode Exit fullscreen mode

or:

strategy A: 91%
strategy B: 53%
Enter fullscreen mode Exit fullscreen mode

Analytics become useful when they're attached to:

market
strategy
side
time
execution type
system state
Enter fullscreen mode Exit fullscreen mode

This is why I prefer structured event data over a dashboard that only stores aggregate counters.


Metrics should come from state transitions

The easiest way to produce useful execution analytics is to make state transitions observable.

For example:

ORDER_CREATED
      ↓
ORDER_SUBMITTED
      ↓
FILL_OBSERVED
      ↓
TX_PENDING
      ↓
CONFIRMED
      ↓
SETTLED
      ↓
POSITION_VERIFIED
Enter fullscreen mode Exit fullscreen mode

Every transition can have:

timestamp
execution_id
order_id
trade_id
market_id
state
Enter fullscreen mode Exit fullscreen mode

Then metrics can be derived from the events.

For example:

ORDER_SUBMITTED → FILL_OBSERVED
Enter fullscreen mode Exit fullscreen mode

gives fill latency.

And:

DISCONNECT → TRADING_RESUMED
Enter fullscreen mode Exit fullscreen mode

gives recovery duration.

This is a much cleaner design than adding unrelated counters throughout the codebase.


An execution timeline is more useful than a single number

Imagine this execution:

10:41:08.120  order submitted
10:41:08.340  fill observed
10:41:08.342  transaction pending
10:41:08.910  transaction confirmed
10:41:09.020  position updated
10:41:09.040  position verified
Enter fullscreen mode Exit fullscreen mode

Now you can calculate:

order → fill
220 ms

fill → confirmation
570 ms

confirmation → position verified
130 ms

total
920 ms
Enter fullscreen mode Exit fullscreen mode

That lets you see where the time is going.

The same timeline can also help when something goes wrong.

For example:

order submitted
fill observed
WebSocket disconnect
process restart
reconciliation
position mismatch
repair
resume
Enter fullscreen mode Exit fullscreen mode

Now the analytics become part of incident investigation.


Analytics and risk should be connected

The same metrics can feed Polymarket trading bot risk controls when an operational threshold is crossed.

Metrics become much more useful when the control plane can act on them.

For example:

execution latency ↑
rejection rate ↑
unknown state duration ↑
        ↓
Execution Health = DEGRADED
Enter fullscreen mode Exit fullscreen mode

Or:

position mismatch count ↑
        ↓
Reconciliation Health = DEGRADED
        ↓
Trading Permission = BLOCK
Enter fullscreen mode Exit fullscreen mode

Or:

recovery duration ↑
        ↓
System Health = DEGRADED
Enter fullscreen mode Exit fullscreen mode

The exact thresholds should be configurable.

The important part is the architecture:

Metrics
   ↓
Health
   ↓
Risk
   ↓
Control
Enter fullscreen mode Exit fullscreen mode

That turns analytics into an operational input instead of a report that nobody reads.

This is part of the broader Polymarket trading bot architecture I've been building.


Don't optimize for dashboards

It's easy to build a beautiful dashboard with:

Orders
Fills
Latency
Volume
P&L
Enter fullscreen mode Exit fullscreen mode

and still not know whether the system is healthy.

The dashboard should answer operational questions.

For example:

Are executions taking longer than normal?

Are partial fills increasing?

Are unresolved executions accumulating?

Are positions frequently drifting from external state?

Is reconciliation taking longer?

Are recovery incidents becoming more frequent?

Those are questions that can lead to engineering action.


A useful execution-health view

I would want something roughly like:

EXECUTION HEALTH

Orders                 1,248
Fill ratio               86%
Partial fills             94
Rejections                11
Unknown executions         3

p50 fill latency         180 ms
p95 fill latency         620 ms

Position mismatches        2
Active recoveries          0
Enter fullscreen mode Exit fullscreen mode

The actual UI doesn't need to look exactly like this.

The important thing is that the metrics represent the system's current operational state.


Anomaly detection can come later

You don't need machine learning to start.

Simple thresholds are enough.

For example:

p95 latency > threshold
Enter fullscreen mode Exit fullscreen mode

or:

unknown executions > threshold
Enter fullscreen mode Exit fullscreen mode

or:

recovery duration > threshold
Enter fullscreen mode Exit fullscreen mode

or:

position mismatches > threshold
Enter fullscreen mode Exit fullscreen mode

Then:

Normal
   ↓
Threshold crossed
   ↓
DEGRADED
   ↓
Investigate
Enter fullscreen mode Exit fullscreen mode

Start with deterministic rules.

You can add more sophisticated anomaly detection later if there is enough historical data to justify it.


Store the raw events

One useful design decision is to preserve enough raw event information to reconstruct metrics later.

For an execution, you might store:

execution_id
order_id
trade_id
market_id

event_type
event_timestamp
observed_timestamp

requested_quantity
matched_quantity

state
reason
Enter fullscreen mode Exit fullscreen mode

Then metrics can be recalculated.

That's better than only storing:

average_fill_latency = 240ms
Enter fullscreen mode Exit fullscreen mode

because once you throw away the underlying observations, it becomes difficult to investigate why the number changed.


Metrics should support incident analysis

Suppose today's recovery time is much worse than yesterday's.

A good system should let you drill down:

Recovery Duration ↑
      ↓
Reconciliation Duration ↑
      ↓
Position Mismatch Count ↑
      ↓
Market X
      ↓
Specific execution / event sequence
Enter fullscreen mode Exit fullscreen mode

Now analytics and incident recovery are connected.

That is the direction I want the infrastructure to take.

That makes analytics useful for Polymarket trading bot incident recovery as well as day-to-day monitoring.


The broader stack

This fits into the rest of my Polymarket work:

Polymarket Trading Bot
        ↓
Execution
        ↓
Execution Verifier
        ↓
Position Reconciliation
        ↓
Risk Controls
        ↓
Trading Control Plane
        ↓
Execution Analytics
Enter fullscreen mode Exit fullscreen mode

Each layer produces information for the next.

The bot produces execution events.

The verifier establishes execution state.

Reconciliation establishes state consistency.

Risk evaluates the resulting exposure and system conditions.

The control plane decides whether trading continues.

Analytics measures how all of those pieces are behaving over time.


What I would measure first

I wouldn't start with 50 metrics.

I'd start with a small set:

1. Order latency
2. Fill latency
3. Fill ratio
4. Partial-fill rate
5. Rejection rate
6. Unknown execution count
7. Unknown execution duration
8. Position mismatch count
9. Reconciliation duration
10. Recovery duration
Enter fullscreen mode Exit fullscreen mode

Those ten already give you visibility into a large part of the execution lifecycle.

Then add metrics when there is a real operational question they can answer.


Final takeaway

A trading bot being profitable doesn't tell you whether the execution system is healthy.

A bot can be making money while:

latency is increasing
partial fills are increasing
rejections are increasing
execution state is staying unknown
positions are drifting
recovery is getting slower
Enter fullscreen mode Exit fullscreen mode

That's why I'm interested in measuring what happens after the strategy produces a trade.

The model is:

Execution
   ↓
Events
   ↓
Metrics
   ↓
Health
   ↓
Risk / Control
   ↓
Action
Enter fullscreen mode Exit fullscreen mode

The goal isn't to build a dashboard full of numbers.

It's to make the trading system measurable enough that you can answer:

What is happening to execution right now, and does it require the system to change its behavior?

That's the level of observability I want around an automated Polymarket trading system.

Top comments (1)

Collapse
 
arhancanli profile image
Arhan Canli •

Tying each metric to an operational decision is the right filter, and the point about distributions over averages for latency holds for everything else on the list too.

The list is all about whether orders execute, though, and nothing yet about the price they execute at, which is where a bot quietly loses money while every health metric stays green. Two price-quality numbers I'd add first:

  • Slippage against the decision price: fill price minus the mid (or best quote) at the moment the strategy decided, in cents of probability, signed so a positive number is a cost. Split it by market and by order size, because thin books on Polymarket make this vary a lot between markets.
  • Markouts: the mid 30 seconds, 5 minutes and 1 hour after each fill, compared with the fill price. If the price keeps moving against you after you're filled, your fills are adverse selection: the orders that fill are the ones better-informed traders were happy to take the other side of. A rising fill ratio with worsening markouts is a warning sign that a fill-ratio dashboard on its own reads as good news.

Both come straight from the raw events you're already storing, as long as each fill keeps the decision timestamp and the book snapshot from that moment.