Mercari Japan (jp.mercari.com) is the single richest source of real-world resale prices for anything from anime figures to camera gear to JDM parts. If you're a reseller, a market researcher, or building an AI agent that needs to know "what does this actually sell for in Japan," you want this data.
But scraping it the obvious way fails. Here's why, and the technique that works.
The trap: Mercari's search API is not callable directly
Mercari's search page is a fully client-side rendered Next.js app. If you open DevTools, you'll spot the endpoint that returns the JSON:
POST https://api.mercari.jp/v2/entities:search
Your first instinct is to replay it with requests or curl. That returns a 401. The endpoint requires a DPoP (Demonstrating Proof of Possession) security header that the page's own JavaScript generates on the fly. It is not a static token you can copy — it's derived per-request from a key the browser holds. Reverse-engineering it means maintaining a fragile, reverse-engineered crypto flow that breaks whenever Mercari updates their bundle.
There's a better way.
The technique: let the browser do the work, capture the response
Instead of issuing the API call, let the real page issue it and intercept the response. Drive a headless Chromium with Playwright, navigate to the search results, and hook the network layer to grab the JSON the page fetches for you.
from playwright.async_api import async_playwright
async def scrape_mercari(keyword: str, max_items: int = 100):
items = []
async with async_playwright() as p:
browser = await p.chromium.launch(headless=True)
page = await browser.new_page()
# Intercept the search API response the page makes itself
async def on_response(resp):
if "entities:search" in resp.url and resp.status == 200:
try:
data = await resp.json()
for it in data.get("items", []):
items.append({
"id": it.get("id"),
"name": it.get("name"),
"price": it.get("price"), # JPY
"status": it.get("status"),
"condition": it.get("condition"),
"brand": it.get("brandId"),
"itemUrl": f"https://jp.mercari.com/item/{it.get('id')}",
})
except Exception:
pass
page.on("response", on_response)
await page.goto(f"https://jp.mercari.com/search?keyword={keyword}")
# SPA pagination: click "次へ" and wait for React hydration
for _ in range(max_items // 120 + 1):
await page.wait_for_timeout(1500)
nxt = page.locator('button[aria-label="次へ"]')
if await nxt.count() and await nxt.is_enabled():
await nxt.click()
else:
break
await browser.close()
return items[:max_items]
No API keys. No login. No cookies. The DPoP header is generated by Mercari's own JavaScript, running in a real browser, so it's always valid. You just read what comes back.
Why pagination is the hard part
The tricky bit isn't the first page — it's that Mercari is a single-page app. A plain scroll doesn't reliably trigger the next fetch, and the "next" button is disabled until React finishes hydrating. The robust pattern is: click 次へ, wait for the hydration + network round-trip, repeat. Cap it with a total-timeout guard so a hung page can't spin forever.
Output shape
One record per listing, ready to drop into a dataframe:
{
"id": "m92852841305",
"name": "Honda CBR250RR MC51",
"price": "7777",
"status": "ITEM_STATUS_ON_SALE",
"condition": "5",
"brand": "Honda",
"itemUrl": "https://jp.mercari.com/item/m92852841305"
}
Verified run: 359 items across 3 pages, zero duplicates.
What you can do with it
- Resale arbitrage — pull Mercari prices, compare against eBay sold comps, surface margin opportunities. This is exactly what cross-border resellers of Japanese collectibles do manually.
- Price monitoring — track what a specific item actually clears at over time.
- Market intelligence — inventory levels, condition distribution, seller activity per product line.
- AI agent tooling — give an agent a "what's the real market price in Japan" function.
If you'd rather not maintain the scraper
I've packaged this exact approach as a hosted Apify Actor so you don't have to run Playwright, manage proxies, or fix it when Mercari changes their frontend:
👉 https://apify.com/fruitful_quintessence/mercari-japan-search-scraper
- Pay-per-result ($2 per 1,000 items), no subscription
- Runs from any language via REST, or plug it straight into an AI agent through the Apify MCP server
- Source and technique write-up: https://github.com/atushi1841/mercari-japan-search-scraper
The proxy note matters: do not enable Apify's auto-proxy for Mercari — it causes page-load timeouts. Mercari serves fine from datacenter IPs directly. Residential proxies only if you need them.
Wrap up
The general lesson for scraping modern SPAs: when an endpoint is protected by a browser-generated proof (DPoP, signed requests, WAF tokens), stop trying to reproduce the proof and start intercepting the response the real page already fetches. It's more robust and dramatically less maintenance.
If this saved you an afternoon of reverse-engineering, a clap helps others find it. Questions in the comments.
Top comments (0)