DEV Community

Cover image for Scrape Amazon in Python without proxies: best sellers, niche research and a price tracker
Albin Johansson
Albin Johansson

Posted on Fully Autonomous

Scrape Amazon in Python without proxies: best sellers, niche research and a price tracker

Scraping Amazon yourself means rotating proxies, CAPTCHAs and HTML that changes every few weeks. Amazon's official API (PA-API, replaced by the Creators API in 2026) only works for Amazon Associates with 10 qualifying sales in the last 30 days. And seller tools cost $29/month (Jungle Scout) to $99/month (Helium 10), both billed annually.

Here's the other way: Amazon product data as JSON from Python, paid per result. Below are three scripts I actually use: best sellers in a niche, a quick niche check, and a price tracker.

1. Setup

pip install "apify-client>=3"
export APIFY_TOKEN=...   # free account at console.apify.com
Enter fullscreen mode Exit fullscreen mode

I'm using the Amazon Product Scraper on Apify (disclosure: I built it). It reads a commercial Amazon data API instead of loading Amazon in a browser, covers 18 marketplaces and costs $1.50 per 1,000 search results.

2. What actually sells? Rank by "bought in past month"

Amazon shows "5K+ bought in past month" on many listings. It's the closest thing to public sales data, and it comes back as boughtPastMonth:

import os
from apify_client import ApifyClient

client = ApifyClient(os.environ["APIFY_TOKEN"])
run = client.actor("jesting_grass/amazon-product-scraper").call(run_input={
    "searchQueries": ["yoga mat"],
    "marketplace": "us",
    "maxResultsPerQuery": 100,
})
rows = [r for r in client.dataset(run.default_dataset_id).iterate_items() if "error" not in r]
rows.sort(key=lambda r: r.get("boughtPastMonth") or 0, reverse=True)
for r in rows[:6]:
    badge = "Best Seller" if r["isBestSeller"] else "Amazon's Choice" if r["isAmazonChoice"] else ""
    print(f'{r.get("boughtPastMonth") or 0:>6,}/mo  ${r["price"] or "-"!s:<7} {r["rating"]}★ {r["reviewsCount"]:>7,} reviews  {badge:<15} {r["title"][:40]}')
Enter fullscreen mode Exit fullscreen mode

Real output (amazon.com, September 2026):

10,000/mo  $-       4.6★  70,600 reviews                  Amazon Basics Extra Thick Exercise Yoga
 8,000/mo  $22.49   4.2★   1,100 reviews                  18 Pcs EVA Foam Floor Tiles, Puzzle Exer
 7,000/mo  $15.99   4.6★   3,800 reviews                  CAP Barbell 1/2-Inch High Density Exerci
 7,000/mo  $36.97   4.7★   1,400 reviews  Best Seller     CAP Barbell Folding Exercise Mat – Durab
 6,000/mo  $21.23   4.6★  70,600 reviews                  Amazon Basics Extra Thick Exercise Yoga
 6,000/mo  $28.16   4.4★  10,600 reviews                  Gruper Yoga Mat Non Slip, Eco Friendly E
Enter fullscreen mode Exit fullscreen mode

Two things stand out. The "Best Seller" badge is per category, so the product selling most in this search (10,000 a month) doesn't have it. And a foam floor tile set with only 1,100 reviews sells 8,000 a month in a "yoga mat" search. That's the kind of gap a seller looks for. ($- means the listing only shows a price after picking a variant.)

3. Check several niche ideas at once

The same data answers "is this niche worth entering?". For each idea: total monthly sales of the top 50 results, median price, median review count, and "low-review winners" (1,000+ sales a month with fewer than 500 reviews):

from statistics import median

IDEAS = ["dog water bottle", "cable organizer", "bamboo cutting board"]
run = client.actor("jesting_grass/amazon-product-scraper").call(run_input={
    "searchQueries": IDEAS, "marketplace": "us", "maxResultsPerQuery": 50,
})
by_query = {}
for r in client.dataset(run.default_dataset_id).iterate_items():
    if "error" not in r:
        by_query.setdefault(r["searchQuery"], []).append(r)

for q in IDEAS:
    rows = by_query[q]
    sold = sum(r.get("boughtPastMonth") or 0 for r in rows)
    winners = [r for r in rows if (r.get("boughtPastMonth") or 0) >= 1000 and (r.get("reviewsCount") or 0) < 500]
    print(f'{q:<22}{sold:>9,}/mo  ${median(r["price"] for r in rows if r.get("price")):.2f}'
          f'  {median(r.get("reviewsCount") or 0 for r in rows):>6,.0f} reviews  {len(winners)} low-review winners')
Enter fullscreen mode Exit fullscreen mode

Real output:

dog water bottle         22,608/mo  $14.98   2,200 reviews  3 low-review winners
cable organizer         177,150/mo  $13.68   3,850 reviews  4 low-review winners
bamboo cutting board     70,511/mo  $21.99   1,300 reviews  3 low-review winners
Enter fullscreen mode Exit fullscreen mode

Cable organizers move almost 8x the volume of dog water bottles at a similar price, and bamboo cutting boards have the lowest review barrier. The sales numbers add up Amazon's rounded labels ("1K+"), so read them as a floor, not an exact count. That run returned 150 products for about $0.23.

4. A price tracker in 20 lines

For price monitoring, pass ASINs (or product URLs) and append the result to a CSV. Schedule it daily and the CSV becomes your price history:

import csv
from datetime import date
from pathlib import Path

ASINS = ["B0B41YH9B6", "B0F8MHPVPH", "B0BLCW7MBR"]
COLUMNS = ["date", "asin", "price", "listPrice", "discountPercent", "isAvailable", "rating", "reviewsCount", "title"]

run = client.actor("jesting_grass/amazon-product-scraper").call(run_input={"asins": ASINS, "marketplace": "us"})
out = Path("prices.csv")
new_file = not out.exists()
with out.open("a", newline="", encoding="utf-8") as f:
    w = csv.DictWriter(f, fieldnames=COLUMNS, extrasaction="ignore")
    if new_file:
        w.writeheader()
    for row in client.dataset(run.default_dataset_id).iterate_items():
        if "error" not in row:
            w.writerow({"date": date.today().isoformat(), **row})
Enter fullscreen mode Exit fullscreen mode

First rows of the real CSV:

date,asin,price,listPrice,discountPercent,isAvailable,rating,reviewsCount,title
2026-09-27,B0F8MHPVPH,99.99,112.99,12,True,4.5,4857,"FEZIBO Standing Desk 48 × 24 Inch Electric Height Adjustable, Maple"
2026-09-27,B0B41YH9B6,99.99,119.99,17,True,4.5,12020,"ErGear 48 X 24 Inch Height Adjustable Electric Standing Desk, Black"
Enter fullscreen mode Exit fullscreen mode

Full product pages cost $0.008 each, so tracking 3 products daily is about $0.72 a month. The same call also returns brand, specs, all images, variant ASINs and, when Amazon shows them, top reviews.

5. No Python? One HTTP call

curl -X POST "https://api.apify.com/v2/acts/jesting_grass~amazon-product-scraper/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"searchQueries":["standing desk"],"marketplace":"us","maxResultsPerQuery":20}'
Enter fullscreen mode Exit fullscreen mode

It also runs from n8n, Make and Zapier, and AI agents can call it through the Apify MCP server. Switch marketplace to uk, de, jp or 14 others for other Amazon stores.


All examples (best sellers, niche research, price tracker, ASIN details, Node.js, cURL): github.com/emiohr/amazon-product-scraper-api. Questions or feature requests? Drop a comment.

This article was written by an AI agent under my direction, and all code was tested against the live API. Not affiliated with Amazon, Jungle Scout or Helium 10.

Top comments (0)