DEV Community

Hay Equipos
Hay Equipos

Posted on

How to scrape product prices, stock and GTIN from any online store

You have a list of product pages from competitors, suppliers or resellers, and you need the price, the stock status and the barcode for each one in a spreadsheet. Copying them by hand works for ten products and falls apart at five hundred. Writing a scraper per store is worse, because every store lays out its pages differently.

There is a shortcut. Most online stores already publish their product data in a machine readable form (Schema.org product markup and Open Graph product tags), because search engines and social previews need it. The Ecommerce Product Page Scraper by Hay Equipos reads that data from any product URL you give it, so the same tool works on Shopify, WooCommerce, BigCommerce, Magento, Wix, Squarespace, PrestaShop and custom stores.

What you get back

One row per URL. Here is an example with illustrative values:

Field Example
name Classic Canvas Sneaker
brand Example Brand
price 64
currency USD
listPrice 80
onSale true
availability InStock
inStock true
sku CCS-WHT-09
gtin 0123456789012
rating 4.5
reviewCount 312
variantCount 12
status ok

Rows also carry mpn, lowPrice and highPrice, condition, category, color, seller, images, canonicalUrl, language, dataSources (which markup the data came from), an optional plain text description (up to 5,000 characters) and a variants list with size, color, SKU, GTIN, price and stock per variant.

Every row has a status:

  • ok: product data found (the only status that is charged)
  • no_product_data: the page loaded but had no price and no Schema.org product
  • blocked: the site answered with a bot check or access denied page
  • http_error: 404, 410, 500 and similar
  • skipped_robots_txt: the site's robots.txt closes the page to automated tools
  • error: timeout or network failure

A RUN_SUMMARY record in the run's key value store counts each status.

How it works

The actor fetches each page with a plain HTTP request. It reads JSON-LD product blocks first (including @graph, ProductGroup with variants, AggregateOffer and list price specifications), then fills gaps from microdata, then from Open Graph product: tags. No browser and no AI guessing, so the results are repeatable: the same page gives the same row.

It respects robots.txt by default, never logs in, never solves captchas and does not use proxies to get around blocks. Each site gets one request at a time, with a pause between requests.

Step by step in the Apify Console

  1. Open the actor on the Apify Store (link at the end) and click Try for free. You need a free Apify account.
  2. In Product page URLs, paste your links, one per line. Use single product pages, not category or search pages.
  3. Leave Include description and Include variants on, or switch them off for smaller rows.
  4. Keep Respect robots.txt on. Adjust Parallel requests (default 5, across different sites) and Pause between requests to the same site (default 1,000 ms) if you need to.
  5. Set Maximum URLs (default 1,000) as a safety cap, and optionally a maximum charge for the run in the run options.
  6. Click Start. When the run finishes, open the output table and export it as CSV, Excel or JSON.

For price tracking, save the input as a task and add a daily schedule. Then compare price, listPrice and inStock between runs in your sheet or database.

Calling it from code

With curl, using the synchronous endpoint that returns the dataset items directly (good for small lists):

curl -X POST \
  "https://api.apify.com/v2/acts/pistachio_implementation~product-page-extractor/run-sync-get-dataset-items" \
  -H "Authorization: Bearer $APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "productUrls": [
      "https://www.example-store.com/products/classic-canvas-sneaker",
      "https://shop.example.com/p/desk-lamp-black"
    ],
    "includeDescription": false,
    "includeVariants": true
  }'
Enter fullscreen mode Exit fullscreen mode

With Python and the apify-client package (pip install apify-client):

import os
from apify_client import ApifyClient

client = ApifyClient(os.environ["APIFY_TOKEN"])

run = client.actor("pistachio_implementation/product-page-extractor").call(
    run_input={
        "productUrls": [
            "https://www.example-store.com/products/classic-canvas-sneaker",
            "https://shop.example.com/p/desk-lamp-black",
        ],
        "includeVariants": True,
        "maxItems": 1000,
    }
)

for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    if item["status"] == "ok":
        print(item["name"], item["price"], item["currency"], item["inStock"], item.get("gtin"))
    else:
        print("skipped", item["url"], item["status"])
Enter fullscreen mode Exit fullscreen mode

Pricing

Pay per event, with no subscription and no platform usage charge on top:

Event Price
Product extracted $0.0015 per row with status ok ($1.50 per 1,000 products)

There is no start fee. Blocked pages, pages without product data, HTTP errors, robots.txt skips and timeouts are free. Tracking 500 product pages once a day costs about $0.75 a day if every page returns data.

Limits and what it does not do

  • Sites with strong bot protection often answer datacenter traffic with a block page or let the request time out. Those rows come back as blocked or error and cost nothing. The actor does not try to get around blocks.
  • Only published structured data is returned. If a store leaves GTIN or stock out of its Schema.org data, those fields are empty.
  • No browser. Prices that appear only after JavaScript runs, with no structured data behind them, are not seen.
  • Prices are what a visitor from a US datacenter with no cookies sees. Stores that change price or currency by country may show you different values.
  • One product per URL. Category and search pages are not crawled. Collect product URLs first, for example from the store's sitemap.

Use the data in line with each source site's terms.

Try it on the Apify Store: https://apify.com/pistachio_implementation/product-page-extractor

Top comments (0)