DEV Community

Cover image for How I scrape any website in under 5 minutes (Python + Playwright)
Kabir Hossain
Kabir Hossain

Posted on

How I scrape any website in under 5 minutes (Python + Playwright)

python, webdev, tutorial, automation
Every scraping tutorial I've read online ends the same way.

"Here's how to grab a title with BeautifulSoup." Twenty lines of Python. Works on the demo page. Fails on a real website.

That bothered me for a long time. Because real websites are messy. They use JavaScript. They block bots. They hide data behind logins. They paginate. They rename their HTML classes every few months and break every script you've written.

So instead of writing another "20-line scraper" tutorial, I want to show you what actual production scraping looks like. Not the theory — the workflow I use for real client jobs.

What most tutorials skip
When someone asks me to scrape a site, I don't open a fresh Python file and start writing selectors. I follow a process.

Step one is an audit. I check the site first. Is it static HTML or JavaScript-rendered? Does it have clear repeating elements like product cards? Does it paginate cleanly? Does robots.txt allow scraping?

If I skip this step, I waste hours trying to scrape a site that was never going to work. Five minutes of auditing saves five hours of frustration.

Building the audit as a tool
Since I audit every site the same way, I automated the audit itself.

Here's what my tool does when I run it:

bash
tendem-audit
It asks for a URL. Then it:

Fetches the page using the same fetch strategy I use for scraping (static HTTP first, Playwright second)

Checks robots.txt

Auto-detects the repeating card block using heuristics

Fingerprints the CMS (Shopify, WooCommerce, WordPress, Magento)

Prints a verdict: green, yellow, orange, or red

Green means I can quote the job today. Red means I politely decline.

That one step — automating the audit — saved me more time than any scraper I've written.

The actual scraping
Once I know the site is scrapeable, the scraping itself is boring. That's how it should be.

bash
tendem-scrape https://shop.example.com/products --preset shopify --max-pages 50 --format csv,json,sqlite
That one command:

Fetches all 50 pages automatically by following the "Next" link

Uses the WooCommerce preset to find the right fields (title, price, image, link)

Handles the pagination URL changes

Merges items into one dataset

Writes CSV, JSON, and SQLite files

Total time: two minutes. For a 1,000-product catalog.

What production scraping actually means
The scraping part is 20% of the job. The other 80% is what makes a client pay.

Here's what I deliver for every scraping job — no exceptions:

Clean CSV, ready to open in Excel

Not a raw dump. Columns with the right names. No hidden characters. UTF-8 BOM so Excel doesn't mangle it.

Interactive HTML dashboard

A single file the client opens in their browser. Sortable columns. Filterable search box. Works offline. No dependency on my server.

Data quality report

A file that lists exactly what problems I found: missing titles, empty prices, duplicate rows, broken image links. Clients trust the data because they can see what was checked.

Audit trail

What was scraped, when, from what URL, by what method. If the client asks "how did you get this?", the answer is one JSON file away.

The technical stack

For those who want to know how it works under the hood:

Python 3.10 or newer

Playwright for anything JavaScript-rendered

BeautifulSoup with lxml for parsing

Pydantic v2 for validated data models

SQLite or PostgreSQL for storage

GitHub Actions for the CI pipeline

Pytest and Allure for tests and reports

The framework has over 70 automated tests. There's a public test dashboard on GitHub Pages that shows them passing.

What I learned doing this for real clients
Three things stand out.

First, sessions and logins matter more than anything. Half the sites I scrape require a login. If you don't learn how to save cookies and reuse them, half the job disappears.

Second, quality control is what clients actually pay for. A raw CSV is worth ten dollars. A CSV plus a quality report plus a dashboard is worth a hundred. Same data, different packaging.

Third, when a site changes, your tests catch it before the client does. That's not a nice-to-have — it's the difference between a one-off job and a monthly subscription.

Try it yourself
bash
git clone https://github.com/pranromumu/tendem-scraper
cd tendem-scraper
python -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"
playwright install chromium

tendem-scrape https://books.toscrape.com/ --max-pages 3 --open
The last command scrapes three pages and opens the dashboard in your browser.

What's next
I'm building this out as a service for e-commerce sellers and marketing agencies who need competitor pricing, product catalogs, and lead lists on a schedule. If you have a scraping project in mind, or you just want to talk about Python automation, feel free to reach out.

The GitHub repo is open source. Fork it, break it, submit a pull request.

About me: I'm a web scraping and automation engineer based in Malaysia. I build Python tools that extract data from websites and deliver clean, structured files with quality guarantees.

Live demo: https://pranromumu.github.io/tendem-scraper/

GitHub repo: https://github.com/pranromumu/tendem-scraper

Fiverr gig: https://www.fiverr.com/s/3AAA3yL

If you found this post useful, connect with me. I write about Python, scraping, and automation for real client work.

Top comments (7)

Collapse
 
koda2026 profile image
Harun - solo dev •

the "80% is delivery, 20% is scraping" insight is exactly what separates hobby scripts from production engineering.

most tutorials stop at getting the html. but delivering a clean utf-8 csv with a data quality report and an audit trail is what actually builds trust. i see this exact same principle in ai data pipelines: if the ingestion layer doesn't validate and clean the data, the downstream llm just hallucinates on garbage.

your automated audit step (fingerprinting the cms and checking robots.txt before even writing a selector) is brilliant. it saves so much wasted dev time on doomed projects.

curious about one edge case in your workflow: when playwright gets flagged by advanced anti-bot systems (like cloudflare turnstile or datadome), do you rely on stealth plugins, or do you have a proxy rotation strategy baked into the pipeline to handle the blocks gracefully?

massive respect for open-sourcing the framework with 70+ automated tests. that's how you build tools that actually survive in the wild. 🐯

Collapse
 
prantomomo profile image
Kabir Hossain •

Thank you,that means a lot coming from someone working in AI data pipelines.

You're right about the ingestion layer. Any downstream consumer of scraped data is only as good as the input. That's why the QA report and audit trail were the first things I built, not the last.

On your question about Cloudflare Turnstile and DataDome-here's my honest approach:

Static request first. If it works, I never open a browser.

Playwright headless with a stealth script that hides navigator.webdriver.

Playwright visible. Oddly effective-the JS challenge runs and the fingerprint looks human.

Proxy rotation. Simple round-robin right now, not enterprise-grade yet.

What I don't do is use CAPTCHA-solving services. Two reasons. First, it crosses an ethical line I'm not comfortable with. Second, clients who need that usually have a deeper problem than tooling can fix.

When I hit a real wall, I tell the client honestly. Sometimes the answer is "use their public API" or "run this manually twice a week." Turning down the wrong project has saved me more time than any single tool.

Curious about your side do you validate at ingestion, or do you let the model flag garbage and filter downstream?

Collapse
 
koda2026 profile image
Harun - solo dev •

the ethical line on captcha solvers is a massive flex. turning down the wrong project instead of forcing a bad technical solution is the ultimate senior dev move. (and the "visible playwright" trick is brilliantly pragmatic—sometimes the best human fingerprint is just being visible).

to answer your question: i do both, but i heavily lean on ingestion validation to save tokens and latency on the edge.

pre-flight (ingestion): before the llm even sees the request, i run a lightweight constitutional ai check and basic structural validation. if the input is obviously malformed, too large, or violates safety rules, it gets rejected immediately. this prevents wasting compute on garbage.

post-flight (downstream): the llm's generated code goes into a secure, isolated js sandbox (iframe). if the code throws a runtime error, the sandbox catches it and feeds the error trace back to the llm for self-correction, rather than just showing the broken code to the user. so the "downstream filter" is actually an automated, silent retry loop.

it's not perfect, but catching errors at the sandbox level before the user sees them has drastically improved the perceived reliability of the app.

massive respect for the honest breakdown, kabir. this is exactly the kind of real-world engineering talk the community needs. 🐯

Thread Thread
 
prantomomo profile image
Kabir Hossain •

Harun,this is the reply I was hoping for.

The constitutional AI pre-flight check is a pattern I hadn't seen framed that way. Filtering at ingestion to save tokens and latency makes total sense you're treating the LLM as the expensive resource, not the compute. That's the right mental model.

The iframe sandbox with a silent retry loop is genuinely clever. Feeding runtime error traces back to the model instead of showing broken code to the user that's the difference between a demo and something people trust. Most AI code tools I've seen fail exactly at that step. The user sees an error and loses confidence in the whole product, even if the model was 95% right.

I'm curious about one thing. When the sandbox retries, do you cap the number of attempts? I've noticed that with LLM self-correction loops, the model often fixes the first issue and introduces a new one. Curious whether you draw a hard line or let it keep going.

On my side, I don't have that kind of pipeline yet. My extraction validation is all rule-based schema checks, price format parsing, deduplication. It's nowhere near a constitutional check, but it catches the obvious stuff before anything downstream sees it.

Thanks for the detailed reply. This is the kind of conversation that's hard to have on Twitter or Reddit-glad dev.to still has this.
Kabir

Thread Thread
 
koda2026 profile image
Harun - solo dev •

you hit the nail on the head with the oscillation problem. that is exactly why a hard cap is non-negotiable.

i cap the silent retry loop at exactly 2 attempts.

here is the logic: if the llm fixes the syntax error but introduces a logical bug on retry #1, retry #2 is the last chance to self-correct. if it fails a third time, the loop breaks.

when it hits the cap, it doesn't spin forever or burn tokens. it gracefully fails, logs the final error trace, and falls back to presenting the user with the last known working state (or a clear, human-readable error explaining what went wrong). it's better to admit defeat than to confidently serve broken code or rack up latency on a mobile connection.

also, don't undersell your rule-based schema checks! for structured data extraction (like prices and formats), regex and strict schema validation are actually superior to llm checks. they are deterministic, zero-cost, and 100% reliable. llms should be reserved for the unstructured reasoning, while rule-based checks handle the strict boundaries.

glad to have this conversation here too. this is what dev.to is for. 🐯

Thread Thread
 
prantomomo profile image
Kabir Hossain •

Harun,the "gracefully fail and fall back to last known working state" pattern is the piece I hadn't seen articulated before.

Most retry loops I've studied either spin forever or hard-crash. Your middle path-try twice, then revert to a known-good output-is the kind of decision that only comes from actually shipping to users. I'm going to steal that idea for my scheduled scrapes. If a scrape fails on Monday, the previous week's data is still valid and useful; I can ship it with a "stale" flag instead of nothing.

On the rule-based checks point, you're right, and it's the same reason I use SQLite + regex for price validation rather than asking a model. Deterministic beats probabilistic whenever you have a well-defined answer. LLMs are for the cases where you don't.

Glad our paths crossed. Following your work now. 🚀

Kabir

Thread Thread
 
koda2026 profile image
Harun - solo dev •

kabir, that genuinely means a lot.

honestly, your adaptation of it for scheduled scrapes (shipping stale data with a "stale" flag) is even better for your specific use case than it is for code generation. it perfectly balances system availability with user transparency.

deterministic rules for the known, probabilistic reasoning for the unknown, and graceful degradation when things break. that’s the holy trinity of resilient systems.

really glad our paths crossed too. i just followed you back—looking forward to seeing what you build next and learning from your automation journey. keep crushing it! 🐯🚀