You can get a lot of value from public web data, but the useful part is rarely the first scrape. The real problem starts a week later: selectors stop matching, prices contain currency symbols from three regions, your job silently writes zero rows, and nobody notices until a dashboard looks wrong.
Public data needs boring data engineering
Public web data is attractive because it is already out there: product pages, public directories, reviews, job listings, government portals, documentation sites, release notes, and market pages. Smaller teams can answer questions that used to require expensive data vendors.
But public does not mean ready to use.
If you want to track competitor pricing, the hard part is not fetching HTML once. The hard parts are:
- deciding exactly which pages answer the business question
- collecting data without hammering someone else's site
- detecting when extraction breaks
- normalizing messy fields into stable types
- storing observations over time instead of overwriting history
- keeping enough metadata to debug bad data later
That is the difference between a scraper and a pipeline.
Start with the question, not the scraper
A common mistake is to start with, can we scrape this site? A better first question is, what decision will this data support?
For example, competitor pricing sounds specific, but it usually needs more definition:
- Which competitors matter?
- Which SKUs can be matched reliably?
- Do we care about list price, discounted price, shipping, or availability?
- How often does the data need to change before it affects a decision?
- What failure rate is acceptable?
If the answer only needs weekly trends, a daily scrape with manual review may be enough. If the answer feeds an automated repricing system, you need stricter validation, alerting, and rollback behavior.
This matters because collection cost scales with freshness, coverage, and reliability. Scraping 50 pages once is simple. Scraping 50,000 pages every hour while proving the data is still correct is a different system.
A minimal pipeline shape
Here is a small Python example that shows the pattern. It fetches a public product listing page, parses product cards, validates required fields, and stores each observation in SQLite.
import datetime as dt
import sqlite3
import requests
from bs4 import BeautifulSoup
URL = 'https://example.com/products'
HEADERS = {
'User-Agent': 'market-monitor/1.0 contact: data-team@example.com'
}
def fetch_html(url):
response = requests.get(url, headers=HEADERS, timeout=10)
if response.status_code == 403:
raise RuntimeError('blocked: 403 Forbidden, check site policy or bot protection')
if response.status_code == 429:
raise RuntimeError('rate limited: 429 Too Many Requests, back off before retrying')
response.raise_for_status()
return response.text
def parse_products(html):
soup = BeautifulSoup(html, 'html.parser')
cards = soup.select('.product-card')
if not cards:
raise RuntimeError('selector matched 0 product cards, page structure may have changed')
rows = []
observed_at = dt.datetime.utcnow().isoformat(timespec='seconds') + 'Z'
for card in cards:
product_id = card.get('data-product-id')
name_el = card.select_one('.product-name')
price_el = card.select_one('.price')
if not product_id or not name_el or not price_el:
continue
raw_price = price_el.get_text(strip=True)
price = float(raw_price.replace('$', '').replace(',', ''))
rows.append({
'source': URL,
'product_id': product_id,
'name': name_el.get_text(strip=True),
'price': price,
'observed_at': observed_at
})
return rows
def save(rows):
with sqlite3.connect('market.db') as db:
db.execute('''
create table if not exists price_observations (
source text not null,
product_id text not null,
name text not null,
price real not null,
observed_at text not null,
primary key (source, product_id, observed_at)
)
''')
db.executemany('''
insert into price_observations
(source, product_id, name, price, observed_at)
values (:source, :product_id, :name, :price, :observed_at)
''', rows)
html = fetch_html(URL)
products = parse_products(html)
if len(products) < 5:
raise RuntimeError(f'only parsed {len(products)} products, expected at least 5')
save(products)
print(f'saved {len(products)} product observations')
This is intentionally small, but it includes a few important habits:
- It fails loudly when the page returns 403 or 429.
- It treats zero matched elements as a broken extractor, not an empty dataset.
- It stores observations with timestamps instead of keeping only the latest value.
- It validates minimum row count before writing data.
For workflows where the main pain is keeping extraction stable across changing public pages, Wire can sit in the collection part of the pipeline while your application code handles validation, storage, and analysis.
Store raw and normalized data separately
Do not throw away the messy version too early. If a price parser turns 1.299,00 € into 1.299, you will want the original string when someone asks why the chart dropped by 99 percent.
A practical schema often has two layers:
create table raw_pages (
source_url text not null,
fetched_at text not null,
status_code integer not null,
html text not null,
primary key (source_url, fetched_at)
);
create table normalized_prices (
source_url text not null,
product_id text not null,
price_cents integer not null,
currency text not null,
observed_at text not null
);
Raw storage helps you replay parsers after a bug fix. Normalized storage helps analysts and application code query clean fields without parsing HTML every time.
You do not need to keep raw HTML forever. Set a retention period based on debugging needs, storage cost, and any legal or privacy requirements.
Make failures visible
A scraper that crashes is easier to fix than one that quietly writes bad data. Track operational metrics like:
- pages fetched
- non-200 responses
- rows extracted
- rows rejected during validation
- duplicate rate
- parse duration
- last successful run time
Even a simple cron job can report useful signals:
0 * * * * /usr/bin/python3 /app/collect_prices.py >> /var/log/price-job.log 2>&1
Then alert on symptoms that imply broken data, not just process failure. For example, alert if the job exits successfully but extracts fewer than 80 percent of the usual row count.
If a team needs public web data as a recurring feed rather than one-off HTML parsing, Wire is relevant because recurring extraction only becomes useful when failures, retries, and structured outputs are part of the design.
Be careful with access and ethics
Publicly visible data still has rules around it. Before collecting, check the site's terms, robots.txt, API options, and applicable laws. Avoid authenticated areas unless you have permission. Do not collect personal data just because it appears on a page. Rate limit your requests and identify your crawler with a useful user agent.
Also prefer official APIs when they meet the need. APIs usually provide clearer contracts, better stability, and fewer surprises. Scraping makes sense when no suitable API exists, the data is public, and your collection pattern is reasonable.
A practical next step
Pick one public source you already use manually. Write down the exact decision it supports, then build a small pipeline that fetches it, validates row counts, stores timestamped observations, and alerts when extraction drops to zero. That exercise will teach you more about public web data than a much larger scrape with no validation.
Top comments (0)