DEV Community

Anakin
Anakin

Posted on

Real-time market data is mostly a freshness problem

A competitor changes price at 2:15 PM. Your scraper runs at midnight. The dashboard looks clean the next morning, but the business already lost a few hours of margin or conversion. The bug is not that you lack data. The bug is that your data describes the wrong moment.

Real-time usually means fresh enough

Most teams do not need true real-time market intelligence. They need data that is fresh enough for the decision it drives.

A pricing bot may need competitor changes within 5 minutes. An assortment gap report may be fine at 6 hours. A weekly category review does not need streaming infrastructure at all.

The first useful design question is not how do we make this real-time? It is:

What is the maximum age of this data before acting on it becomes unsafe?

That gives you a freshness SLO. For example:

competitor_price_freshness_seconds < 900 for 95% of tracked SKUs
stock_status_freshness_seconds < 300 for top 500 SKUs
Enter fullscreen mode Exit fullscreen mode

Track freshness as a first-class metric. Job success is not enough. A scrape can succeed every hour while still being useless if the source changed 10 minutes after each run.

Wire fits this specific problem when reliable extraction, freshness tracking, and failure handling matter more than running one-off scrapers.

The shape of the system

A simple near-real-time market monitor has four parts:

  1. Fetch the source
  2. Normalize the result
  3. Compare it with the last known state
  4. Emit a change event

Do not plug dashboards directly to raw scrapes. Raw pages are unstable. Prices move between list price, sale price, coupon price, member price, and regional price. Normalize before anyone consumes the data.

Here is a small Python example that watches one product page and emits an event only when the normalized state changes:

import hashlib
import json
import re
import sys
from datetime import datetime, timezone
from decimal import Decimal
from pathlib import Path

import requests
from bs4 import BeautifulSoup

STATE_FILE = Path('product_state.json')


def parse_price(text):
    cleaned = re.sub(r'[^0-9.]', '', text)
    if not cleaned:
        raise ValueError(f'could not parse price from: {text!r}')
    return str(Decimal(cleaned))


def fetch_product(url):
    res = requests.get(
        url,
        timeout=10,
        headers={'User-Agent': 'market-monitor/1.0'}
    )

    if res.status_code == 429:
        raise RuntimeError('rate limited by source: HTTP 429')

    res.raise_for_status()

    soup = BeautifulSoup(res.text, 'html.parser')
    price_el = soup.select_one('[data-testid=price]')
    title_el = soup.select_one('h1')
    sold_out_el = soup.select_one('[data-testid=sold-out]')

    if price_el is None:
        raise RuntimeError('price selector missing, page layout may have changed')

    if title_el is None:
        raise RuntimeError('title selector missing, page may be blocked or changed')

    return {
        'url': url,
        'title': title_el.get_text(strip=True),
        'price': parse_price(price_el.get_text()),
        'in_stock': sold_out_el is None,
        'observed_at': datetime.now(timezone.utc).isoformat()
    }


def fingerprint(record):
    comparable = {
        'url': record['url'],
        'price': record['price'],
        'in_stock': record['in_stock']
    }
    payload = json.dumps(comparable, sort_keys=True)
    return hashlib.sha256(payload.encode()).hexdigest()


def load_state():
    if not STATE_FILE.exists():
        return {}
    return json.loads(STATE_FILE.read_text())


def save_state(state):
    STATE_FILE.write_text(json.dumps(state, indent=2))


def main(url):
    state = load_state()
    record = fetch_product(url)
    key = record['url']
    new_hash = fingerprint(record)
    previous = state.get(key)

    if previous is None or previous['hash'] != new_hash:
        event = {
            'type': 'product_changed',
            'previous': previous['record'] if previous else None,
            'current': record
        }
        print(json.dumps(event, indent=2))

    state[key] = {'hash': new_hash, 'record': record}
    save_state(state)


if __name__ == '__main__':
    main(sys.argv[1])
Enter fullscreen mode Exit fullscreen mode

Run it on a schedule:

python monitor.py 'https://example.com/products/sku-123'
Enter fullscreen mode Exit fullscreen mode

In production, the printed event would go to Kafka, SQS, Pub/Sub, or a webhook. The important part is the event boundary. Downstream systems should react to product_changed, not scrape HTML themselves.

How these systems fail

The boring failures cause most of the damage.

A CSS selector stops matching and your parser returns null. If you silently write null into the database, the dashboard shows missing prices as real prices. Raise an error instead.

A source returns a bot challenge with HTTP 200. Your parser sees a valid HTML document, but it is not the product page. This is why the example checks for both price and title. In real systems, also check canonical URLs, currency, region, and known page markers.

A source rate limits you with 429. Retrying immediately makes the block worse. Back off, reduce concurrency per domain, and record the gap in freshness.

A product has multiple prices. List price, sale price, coupon price, and cart price can all be different. Pick the one that matches the business decision and name it clearly. competitor_price is vague. competitor_checkout_price_usd is much harder to misuse.

Speed without precision creates bad automation

Fast bad data is worse than slow bad data because it triggers action before a human notices the mistake.

If you connect price changes directly to automated pricing, add guardrails:

reject price decrease > 20% unless confirmed by two consecutive observations
reject currency changes unless source region also changed
ignore out-of-stock changes that flip back within 2 minutes
alert when freshness exceeds SLO instead of reusing stale values
Enter fullscreen mode Exit fullscreen mode

These checks reduce false reactions. They also make incidents easier to debug because the system can explain why it did or did not act.

Build for loops, not reports

A report answers what happened. A decision loop answers what changed and what should happen next.

That loop can still start manually. Send a Slack alert when a top competitor drops price. Create a ticket when a rival goes out of stock. Notify merchandising when a category gap appears.

Once the signal quality is good, automate the low-risk actions first. For example, refresh an internal recommendation, flag a SKU for review, or pause a stale campaign. Leave margin-sensitive pricing changes behind stricter checks.

The practical next step: pick one decision your team already makes from stale external data, define its freshness SLO, and instrument the age of that data before changing the scraper itself.

Top comments (0)