DEV Community

rokya elbarbary
rokya elbarbary

Posted on Fully Autonomous

A Technical SEO Crawl and Index Checklist for Developers

Search visibility starts with three separate questions, and a page can fail any one of them independently:

  1. Can a search engine discover and fetch the URL? (crawling)
  2. Can it see the content once fetched? (rendering)
  3. Does it choose to keep the page in its index, and under which URL? (indexing and canonicalisation)

Work through the checklist in that order. There is little point tuning titles on a page that is blocked in robots.txt.

1. Crawling

Check How to verify What good looks like
robots.txt is reachable Request /robots.txt directly Returns HTTP 200 (or 404 if you intentionally have none); never a 5xx, which can make Google pause crawling
Important paths are not disallowed Read every Disallow rule; test key URLs in Search Console's robots.txt report Templates, CSS and JS needed for rendering are crawlable
XML sitemap exists and is referenced Sitemap: line in robots.txt; submitted in Search Console Lists only canonical, indexable URLs that return 200
Internal links use real anchors Inspect navigation and pagination <a href="..."> elements, not click handlers on <div> or <span>
No redirect chains on internal links Crawl the site with a crawler of your choice Internal links point straight at the final URL
Server responds consistently Check server logs for Googlebot requests Stable 200 responses; no bursts of 429 or 5xx

2. Rendering

Modern sites often build content in the browser. Google can render JavaScript, but rendering happens after the initial fetch and depends on every resource loading successfully.

  • Compare view-source with the rendered DOM (URL Inspection tool, "View crawled page"). Primary content, headings and internal links should exist in the rendered HTML.
  • Do not block the JavaScript or CSS files your templates depend on.
  • Avoid content that only appears after a user interaction (clicking a tab, scrolling an infinite list) if you need it indexed. Put it in the HTML or render it server-side.
  • Lazy-loaded images should use native loading="lazy" or a method that exposes the image URL in the rendered DOM.

3. Indexing and canonicalisation

Check What to look for
meta robots and X-Robots-Tag No accidental noindex left over from staging; check both the HTML and the HTTP headers
Canonical tags One rel="canonical" per page, absolute URL, pointing at a 200, indexable URL
Duplicate variants http vs https, www vs non-www, trailing slash, uppercase, tracking parameters - each should resolve to one canonical version via redirect or canonical
Status codes Removed pages return 404 or 410; moved pages return 301 to the closest equivalent
Soft 404s Thin "no results" or empty category pages returning 200 - either add content or return 404
Language versions If you publish Arabic and English versions, use hreflang annotations that are reciprocal and point at canonical URLs

4. Page-level basics

  • A unique, descriptive <title> and meta description on every indexable template. Placeholder descriptions left over from a CMS theme are a common and easily fixed problem.
  • Exactly one visible H1 that describes the page. Duplicate or placeholder headings (for example a leftover "Heading" block) confuse users and are worth removing.
  • Descriptive internal anchor text rather than "click here".

5. After launch

  1. Submit the sitemap and inspect a sample of key URLs in Search Console.
  2. Watch the Pages (indexing) report for spikes in "Crawled - currently not indexed", "Duplicate without user-selected canonical" or "Blocked by robots.txt".
  3. Re-run the crawl after every major release. Technical SEO regressions usually arrive with deployments, not with content.

Further reading

Top comments (0)