Search visibility starts with three separate questions, and a page can fail any one of them independently:
- Can a search engine discover and fetch the URL? (crawling)
- Can it see the content once fetched? (rendering)
- Does it choose to keep the page in its index, and under which URL? (indexing and canonicalisation)
Work through the checklist in that order. There is little point tuning titles on a page that is blocked in robots.txt.
1. Crawling
| Check | How to verify | What good looks like |
|---|---|---|
robots.txt is reachable |
Request /robots.txt directly |
Returns HTTP 200 (or 404 if you intentionally have none); never a 5xx, which can make Google pause crawling |
| Important paths are not disallowed | Read every Disallow rule; test key URLs in Search Console's robots.txt report |
Templates, CSS and JS needed for rendering are crawlable |
| XML sitemap exists and is referenced |
Sitemap: line in robots.txt; submitted in Search Console |
Lists only canonical, indexable URLs that return 200 |
| Internal links use real anchors | Inspect navigation and pagination |
<a href="..."> elements, not click handlers on <div> or <span>
|
| No redirect chains on internal links | Crawl the site with a crawler of your choice | Internal links point straight at the final URL |
| Server responds consistently | Check server logs for Googlebot requests | Stable 200 responses; no bursts of 429 or 5xx |
2. Rendering
Modern sites often build content in the browser. Google can render JavaScript, but rendering happens after the initial fetch and depends on every resource loading successfully.
- Compare view-source with the rendered DOM (URL Inspection tool, "View crawled page"). Primary content, headings and internal links should exist in the rendered HTML.
- Do not block the JavaScript or CSS files your templates depend on.
- Avoid content that only appears after a user interaction (clicking a tab, scrolling an infinite list) if you need it indexed. Put it in the HTML or render it server-side.
- Lazy-loaded images should use native
loading="lazy"or a method that exposes the image URL in the rendered DOM.
3. Indexing and canonicalisation
| Check | What to look for |
|---|---|
meta robots and X-Robots-Tag
|
No accidental noindex left over from staging; check both the HTML and the HTTP headers |
| Canonical tags | One rel="canonical" per page, absolute URL, pointing at a 200, indexable URL |
| Duplicate variants |
http vs https, www vs non-www, trailing slash, uppercase, tracking parameters - each should resolve to one canonical version via redirect or canonical |
| Status codes | Removed pages return 404 or 410; moved pages return 301 to the closest equivalent |
| Soft 404s | Thin "no results" or empty category pages returning 200 - either add content or return 404 |
| Language versions | If you publish Arabic and English versions, use hreflang annotations that are reciprocal and point at canonical URLs |
4. Page-level basics
- A unique, descriptive
<title>and meta description on every indexable template. Placeholder descriptions left over from a CMS theme are a common and easily fixed problem. - Exactly one visible H1 that describes the page. Duplicate or placeholder headings (for example a leftover "Heading" block) confuse users and are worth removing.
- Descriptive internal anchor text rather than "click here".
5. After launch
- Submit the sitemap and inspect a sample of key URLs in Search Console.
- Watch the Pages (indexing) report for spikes in "Crawled - currently not indexed", "Duplicate without user-selected canonical" or "Blocked by robots.txt".
- Re-run the crawl after every major release. Technical SEO regressions usually arrive with deployments, not with content.
Further reading
- Google Search Central: How Google Search works
- Google Search Central: JavaScript SEO basics
- Google Search Central: Consolidate duplicate URLs
- For a longer discussion of why this work is ongoing infrastructure rather than a one-off fix, see Technical SEO in 2026: Building the Infrastructure for Organic Revenue.
Top comments (0)