DEV Community

Cover image for Full-page website screenshots and PDFs from a URL list, with cookie banners hidden
Joshua Smith
Joshua Smith

Posted on

Full-page website screenshots and PDFs from a URL list, with cookie banners hidden

page.screenshot({ fullPage: true }) is one line in Playwright. Getting a screenshot that looks like what a visitor sees takes about two hundred more. Here is what goes wrong and a way to skip it.

Why naive full-page screenshots look broken

  • Lazy-loaded images never load because nothing scrolled. You get grey boxes below the fold.
  • Sticky headers and chat widgets are position: fixed, so in a stitched full-page capture they repeat on every screen or cover the content.
  • Cookie banners sit on top of everything. Clicking "accept" on someone else's site on a user's behalf is not something you want a bot doing; hiding the overlay with CSS is.
  • Emoji and CJK text render as boxes on a bare Linux container without the right fonts.
  • Network idle never arrives on pages with analytics beacons, so waitUntil: "networkidle" hangs until the timeout.

Each of these is a small fix. Together, plus device presets, retries, file storage and a sane error record for pages that fail, they are a project.

The two-minute version

Website Screenshot API runs headless Chromium on Apify with all of the above handled: it scrolls to trigger lazy loading, pins fixed elements so they appear once, hides consent pop-ups with CSS (nothing is clicked), ships fonts for emoji and CJK, and reports failed pages as free records.

  1. Paste your URLs into Website URLs.
  2. Pick a Device preset (desktop, laptop, tablet, mobile or a custom viewport), PNG or JPEG, and Full page or first screen only.
  3. Optionally tick Also render a PDF, add Hide elements selectors for chat widgets, set Dark mode, or under Advanced set Wait for element, cookies and extra headers. Click Start.

Images land in the run's key-value store; the dataset record links to each file.

From code

curl -X POST "https://api.apify.com/v2/acts/josh99smith~website-screenshot-api/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{ "urls": ["https://apify.com", "https://www.wikipedia.org"], "device": "laptop", "format": "jpeg", "quality": 80, "fullPage": true }'
Enter fullscreen mode Exit fullscreen mode

What you get back

{
    "url": "https://www.wikipedia.org",
    "success": true,
    "statusCode": 200,
    "title": "Wikipedia",
    "screenshotUrl": "https://api.apify.com/v2/key-value-stores/AbCdEf123/records/screenshot-002-www-wikipedia-org.jpg",
    "format": "jpeg",
    "width": 1366,
    "height": 1101,
    "fullPage": true,
    "device": "laptop",
    "sizeBytes": 172806,
    "renderTimeMs": 1139,
    "warnings": []
}
Enter fullscreen mode Exit fullscreen mode

screenshotUrl is a direct link you can drop into an <img> tag, a Slack message or a report. Pages that fail come back as { "success": false, "errorType": "dns" | "timeout" | "blocked" | ... }.

Cost and limits

$0.005 per screenshot and $0.004 per PDF when enabled. Failed pages are free, there is no start fee, and each run stops at your cost cap. Pages behind a login need you to supply your own cookies (an input field exists for that); the Actor does not log in anywhere and does not bypass bot protection, it reports it as blocked.

Use it from an AI agent

Add https://mcp.apify.com?tools=josh99smith/website-screenshot-api to your MCP client and ask "take a full-page mobile screenshot of apify.com". The agent gets the image URL back.

Disclosure: I built this Actor. Source: github.com/josh99smith/website-screenshot-api.

Top comments (0)