I built this actor; it's a paid tool on Apify with a free trial credit.
Coles is one of Australia's two big supermarkets, and it sits behind Imperva bot protection. That's why most quick scripts people write against coles.com.au die after a few requests. If you want grocery price history, a half-price alert, or unit-price comparisons across a category, you need something that holds up.
I'm an 18-year-old engineering student, and I built Coles Scraper for this. This post covers calling it from Python and Node, what the output looks like, and a small weekly specials tracker.
How it works
coles.com.au is a Next.js site. Every search, category and product page has a matching /_next/data/<buildId>/...json route that returns the page data as JSON. The actor calls those routes directly over plain HTTP (no headless browser), picks up the new build ID whenever Coles deploys, and uses Australian residential proxies. It rotates to a fresh session when Imperva pushes back. In my test runs, 2,100 products (the full half-price list plus the whole bakery department) took 53 seconds with 0 failed requests.
Python example
pip install apify-client
from apify_client import ApifyClient
client = ApifyClient("<YOUR_APIFY_TOKEN>")
run = client.actor("rel8ble/coles-scraper").call(run_input={
"searchQueries": ["tim tam", "milk"],
"maxItems": 50,
})
for p in client.dataset(run["defaultDatasetId"]).iterate_items():
print(p["name"], p["price"], p["unitPriceText"], p.get("specialType"))
Node example
import { ApifyClient } from "apify-client";
const client = new ApifyClient({ token: "<YOUR_APIFY_TOKEN>" });
const run = await client.actor("rel8ble/coles-scraper").call({
specials: "halfprice",
maxItems: 0, // 0 = everything Coles lists
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(`${items.length} half-price products`);
What comes back
A real row from the half-price specials list (no details):
{
"productId": "2993706",
"fullName": "COCA-COLA ZERO SUGAR SOFT DRINK BOTTLE 1.25L",
"price": 2.25,
"wasPrice": 4.5,
"saveAmount": 2.25,
"discountPercent": 50,
"unitPriceText": "$1.80/ 1L",
"specialType": "half_price",
"offerDescription": "save $2.25",
"available": true,
"availableQuantity": 4824,
"promotionalLimit": 12,
"department": "Drinks",
"category": "Soft Drinks",
"aisle": "Soft Drink Bottles",
"source": "specials",
"specialsFilter": "halfprice"
}
With includeDetails: true you also get the barcode (GTIN), ingredients, allergens, country of origin and the full nutrition panel. A trimmed real example from a "chocolate" search:
{
"productId": "2351720",
"name": "Dairy Milk Hazelnut Chocolate Block",
"brand": "Cadbury",
"price": 5.5,
"wasPrice": 8,
"specialType": "price_drop",
"unitPriceText": "$3.06/ 100g",
"barcode": "9300617064923",
"allergens": "Contains Hazelnut, Milk, Soy\nMay Contain Peanut, Wheat",
"countryOfOrigin": "Australia",
"nutrition": {
"servingSize": "25g",
"per100": { "Energy": "2340 kJ", "Protein": "9.0 g", "Sugars - Total": "46.4 g" }
}
}
specialType is one of half_price, price_drop, multibuy, percent_off or special, so filtering for one kind of deal is a one-liner.
Use case: a weekly half-price tracker
Coles specials start on Wednesday. The idea: every Wednesday morning (AEST), pull the half-price list, keep it as a CSV with the date, and print anything on your personal watchlist.
import csv
import datetime
from apify_client import ApifyClient
WATCHLIST = ["tim tam", "coffee", "dishwasher", "olive oil"]
client = ApifyClient("<YOUR_APIFY_TOKEN>")
run = client.actor("rel8ble/coles-scraper").call(run_input={
"specials": "halfprice",
"maxItems": 0,
})
items = list(client.dataset(run["defaultDatasetId"]).iterate_items())
week = datetime.date.today().isoformat()
with open(f"coles-halfprice-{week}.csv", "w", newline="", encoding="utf-8") as f:
w = csv.writer(f)
w.writerow(["productId", "name", "brand", "size", "price", "wasPrice", "unitPriceText"])
for p in items:
w.writerow([p["productId"], p["name"], p.get("brand"), p.get("size"),
p["price"], p.get("wasPrice"), p.get("unitPriceText")])
for p in items:
name = f'{p.get("brand", "")} {p["name"]}'.lower()
if any(word in name for word in WATCHLIST):
print(f'{p["name"]} {p.get("size", "")}: ${p["price"]} (was ${p.get("wasPrice")})')
Schedule it for Wednesday morning with cron or an Apify Schedule. After a few weeks, productId lets you see how often a product cycles back to half price, which is the actual useful part: you stop buying things at full price that go half price every month.
If you care about one store's prices, set storeId (the number at the end of a store page URL, e.g. .../ardeer-7799 is 7799). Otherwise you get the online store's prices.
What it costs
$1.50 per 1,000 results, and one result is one product saved. Product details (barcode, nutrition, etc.) cost nothing extra per result, and failed or blocked requests are never charged.
- A full half-price list in my test was 1,549 products ≈ $2.32 per week
- A 20-item keyword watchlist instead of the full list is a few cents a week
- Apify's free plan gives $5 of monthly credit, about 3,300 products
Limits
- Coles' own totals are approximate. A category that says 550 products can return 549 or 551, because Coles re-ranks items between pages. Duplicates are removed.
-
Online prices by default. In-store prices can differ; set
storeIdfor a specific store. -
includeDetailsdoubles the request count, so leave it off when you only need prices. - No customer reviews in this version.
- Sponsored products are skipped unless you turn on
includeSponsored.
Questions or broken fields: the Issues tab on the actor page comes straight to me.
This article was drafted with AI and published by me.
Top comments (1)