`
Contents
- Why keyword research is a revenue problem
- What search intent actually means
- A repeatable keyword research methodology
- The opportunity scoring model
- Competitor gap analysis
- Mapping keywords to URLs and topic clusters
- Technical validation: can the page actually rank?
- Measuring ROI: a conservative forecast model
- Limitations of this framework
- Keyword research FAQ
- References
Why keyword research is a revenue problem
The most common failure in content programmes is ranking for the wrong thing. A page reaches the first page for a high-volume query, traffic rises, and nothing downstream changes: no demo requests, no add-to-carts, no pipeline. The query was informational, the audience was students or job seekers, and the business had no next step to offer them.
Volume is therefore a starting filter, not a decision variable. In audit work at RankJockey, the queries that produced qualified leads were usually long, specific and low-volume: implementation timelines, integration constraints, pricing structures, "X vs Y for [use case]". That is a practitioner observation rather than a controlled study, but the pattern is consistent enough to plan around: the distribution of queries that drive traffic and the distribution that drive revenue overlap far less than most keyword lists assume.
Three tests keep a keyword set honest:
- Revenue fit. Does the query sit on a measurable path to a conversion event, and is the customer value behind it meaningful?
- Rank fit. Is the difficulty realistic given your domain's current authority, link profile and publishing capacity?
- Experience fit. Can you serve the page type the searcher expects (a comparison, a calculator, a product listing), or would you be forcing a blog post into a transactional slot?
Google's helpful-content guidance frames the same idea from the other side: content should be built for a specific existing audience with a specific need, not for a query string. A keyword target that passes all three tests is, almost by definition, people-first.
What search intent actually means
Most SEO tools label keywords as informational, navigational, commercial or transactional. Those labels are useful shorthand, but the primary source is Google's Search Quality Evaluator Guidelines, which classify queries by what the user wants to accomplish: Know (and the narrower Know Simple, where a short fact answers the query), Do (complete a task, including buying), Website (reach a specific site) and Visit-in-person. Raters judge results against those needs, and ranking systems are tuned to agree with raters.
Two practical consequences follow. First, intent is a property of the results page, not the keyword. "Enterprise CRM" and "best enterprise CRM" differ by one word and return different result types: vendor and category pages for the first, comparison articles and review sites for the second. Second, a query can have a dominant intent and a secondary one; you build for the dominant type and satisfy the secondary in a section.
The fastest way to classify intent at scale is to sample the top ten results for each candidate query and record what kinds of pages rank. If seven of ten are comparison listicles, the query is commercial regardless of what your tool says.
| Intent | Example query | Dominant SERP pattern | Page type to build | Primary KPI |
|---|---|---|---|---|
| Informational (Know) | how to reduce ecommerce return rates | Guides, documentation, video carousel | Guide, calculator, documentation | Qualified engagement, assisted conversions |
| Commercial (Know leading to Do) | best headless commerce platform | Comparison articles, review sites | Comparison page, buyer's guide | Consultation or trial requests |
| Transactional (Do) | shopify plus migration service pricing | Service and pricing pages | Pricing or service page | Leads, orders |
| Navigational (Website) | [brand] api docs | Brand-owned pages | Make sure the right URL ranks | Task completion |
| Retention (existing customers) | [platform] webhook retry limit | Help centres, forums | Documentation, changelog | Ticket deflection, expansion |
A repeatable keyword research methodology
Single-tool research produces single-tool blind spots. Third-party keyword databases estimate volume from sampled clickstream data and their own models; they are directionally useful and individually unreliable. Triangulate them against sources you own:
- Search Console: impressions, clicks, average position and CTR per query and URL. The only first-party view of demand you have.
- Web analytics and site search logs: what visitors look for once they arrive, in their own words.
- Paid search query reports: real conversion rates per query, a far better proxy for intent than any volume estimate.
- CRM notes and sales call transcripts: the vocabulary buyers use before they learn your product's terminology.
- Competitor indexes: which queries other domains rank for and with what page types.
- Crawl data: which of your URLs are indexable, linked and fast enough to carry a target.
The five stages
- Seed. Extract candidate terms from product catalogues, navigation labels, existing rankings, site search, CRM language and paid search reports.
- Expand. Generate variants: modifiers (best, pricing, vs, for [industry]), questions, locations, comparisons, integrations, error messages.
- Classify. Sample the live SERP for each query and record intent, dominant page type, SERP features (Shopping, video, AI summaries, People Also Ask) and whether the top results are brand-owned.
- Score. Apply the opportunity model in the next section.
- Assign. Map every surviving target to exactly one URL, new or existing, with an owner and a KPI.
Start with striking-distance queries
Before hunting for new terms, mine what already half-works. Queries where a URL sits in positions 5 to 20 with meaningful impressions have demonstrated relevance; a content refresh, better internal links or a corrected title tag can move them faster than any new page. The filter is a few lines of pandas on a Search Console export:
import pandas as pd
# Columns: query, page, clicks, impressions, ctr, position
gsc = pd.read_csv("gsc_queries_last_90_days.csv")
striking = gsc[
gsc["position"].between(5, 20) & (gsc["impressions"] >= 200)
].sort_values("impressions", ascending=False)
print(striking.head(25))
The opportunity scoring model
Keyword lists fail in leadership meetings because they present a thousand equal rows. A scoring model turns the argument from "which keywords" into "which weights", which is a conversation executives can actually have.
The naive version of the model is a ratio:
Opportunity = (Intent × Conversion value × Search demand) ÷ (Ranking difficulty × Effort)
It has one flaw worth fixing: difficulty enters linearly, so a huge head term with near-impossible difficulty still outscores a winnable long-tail term. Difficulty is better modelled as a probability of ranking that collapses towards zero beyond what your domain has demonstrated it can do. A logistic curve centred on your domain's "ceiling" (the difficulty level at which you currently reach the top ten about half the time, measurable from your own rankings) does this cleanly:
import math
from dataclasses import dataclass
@dataclass
class Keyword:
term: str
volume: int # monthly searches, or your own GSC impressions
intent: float # 0.2 informational ... 1.0 transactional (from SERP review)
value: float # expected revenue per conversion
difficulty: float # 1-100 from your SEO tool, sanity-checked against the SERP
effort: float # 1-5 internal estimate: content + design + engineering
def p_rank(difficulty: float, ceiling: float = 45, slope: float = 8) -> float:
"""Probability of a top-10 ranking within ~12 months.
ceiling = difficulty at which your domain currently ranks top-10 ~50% of the time."""
return 1 / (1 + math.exp((difficulty - ceiling) / slope))
def opportunity(k: Keyword, ceiling: float = 45) -> float:
expected_value = k.volume * p_rank(k.difficulty, ceiling) * k.intent * k.value
return round(expected_value / max(k.effort, 1))
candidates = [
Keyword("enterprise crm", 40_000, 0.4, 15_000, 85, 4),
Keyword("crm implementation timeline", 900, 0.7, 15_000, 35, 2),
Keyword("crm data migration checklist", 600, 0.6, 15_000, 28, 2),
]
for k in sorted(candidates, key=opportunity, reverse=True):
print(f"{k.term:32} {opportunity(k):>12,}")
# crm implementation timeline 3,672,742
# crm data migration checklist 2,411,935
# enterprise crm 401,571
With a ceiling of 45, the 40,000-search head term drops to third place because a mid-authority domain is unlikely to reach the top ten for it inside a year. Raise the ceiling to 80 for a site that already owns the category and the order flips back. That is the point: the weights encode a claim about your site, and the claim is what leadership should debate.
Each input is a number your team can defend:
- Intent comes from the SERP classification, not the keyword string.
- Conversion value is the average deal value or average order value of whatever the query maps to.
- Demand is tool volume, or better, your own impressions for terms you already appear for.
- Difficulty is your tool's score, checked against the referring-domain counts of the pages actually ranking.
- Effort is an honest estimate that includes engineering time, which content-only models always omit.
Competitor gap analysis
The competitor list from your sales team is the wrong input. In search, your competitors are whoever occupies the results for the queries you want: trade publishers, marketplaces, review aggregators, open-source documentation, forum threads. Some of them will never sell against you and will still take every click.
Build the SERP competitor set empirically. For your scored queries, tally which domains appear in the top ten and how often. The ten to fifteen domains with the highest frequency are your real competitors for this campaign, and the set will differ by cluster.
With ranking data for those domains, gap analysis is set arithmetic:
ours = {row.query for row in our_rankings if row.position <= 20}
theirs = {row.query for row in competitor_rankings if row.position <= 10}
gap = theirs - ours # they rank, we do not appear at all
shared = theirs & ours # head-to-head battles
unique = ours - theirs # defend these
Not every gap is worth attacking. Classify each gap query by why the competitor holds it:
| Gap type | Signal | Action |
|---|---|---|
| Thin winner | Ranking page is short, outdated or generic | Attack now with a deeper, better-structured page |
| Format mismatch | They rank with a tool, template or dataset; you only have prose | Build the format the SERP rewards |
| Authority gap | Ranking pages sit on domains with far more referring domains | Publish linkable assets first; revisit in two quarters |
| Irrelevant demand | Query converts poorly in your paid search data | Deprioritise regardless of volume |
A large keyword footprint is not evidence of a healthy programme; it can hide low conversion rates and high content costs. Judge competitors by the pages that matter to buyers, not by the size of their index. RankJockey's competitor overview runs this classification across a full domain if you would rather not build the pipeline yourself.
Mapping keywords to URLs and topic clusters
Research that ends as a spreadsheet has no value. The deliverable is a URL-level roadmap: for every target, a page (existing or new), an intent, one primary action, an owner, a set of internal links and a KPI.
The most useful rule is one intent, one URL. When two pages on the same domain target the same intent they compete for the same query (keyword cannibalisation) and search engines alternate between them, suppressing both. Before writing anything, check whether an existing URL already ranks for the cluster; refreshing, consolidating with a 301 redirect, or re-titling that page is often faster than a new one, especially when the URL already has backlinks.
Cluster structure
Group targets around the purchase journey rather than around topics in the abstract. A hub-and-spoke cluster typically looks like this:
| Role | Example | Links to |
|---|---|---|
| Hub (commercial) | Headless commerce platforms: an evaluation guide | Every spoke; product or service pages |
| Informational spokes | What is headless commerce; headless vs monolithic | Hub |
| Comparison spokes | Platform A vs Platform B for B2B | Hub; relevant product pages |
| Proof | Case study: checkout latency after migration | Hub; product pages |
| Transactional | Product, pricing or service page | Hub, for context |
Recommended step: before publishing, write the target query, intended intent and primary call to action into the page brief or front matter. If a reviewer cannot state in one sentence what the page is for, the mapping is not done.
For ecommerce, the same logic applies to category and facet architecture: buyers must be able to reach products through indexable category paths without crawlers wasting budget on infinite filter combinations. Sites with large libraries usually get more from a prioritised content refresh programme than from net-new publishing.
Technical validation: can the page actually rank?
Technical checks are usually done after content ships. That order is backwards: if a target URL cannot be crawled, rendered and indexed, the research behind it was wasted. Treat validation as a gate in the pipeline.
| Check | Why it matters | How to verify |
|---|---|---|
| Indexability | A noindex tag, a canonical pointing elsewhere or a robots rule removes the page from contention | Search Console URL Inspection; crawler |
| Crawl depth and internal links | Pages several clicks from the home page are crawled and ranked less reliably | Crawler depth report; inbound internal link counts |
| Duplication | Parameters, faceted navigation and slash variants split signals across URLs | Crawler duplicate report; server logs |
| Rendering | Content injected by client-side JavaScript may be indexed late or not at all | URL Inspection rendered HTML compared with source |
| Redirect chains | Each hop delays crawling and weakens the signal that reaches the final URL | Crawler; curl -I
|
| Performance | Core Web Vitals are a modest ranking signal and a large conversion factor | CrUX, Lighthouse, field RUM |
Performance thresholds worth knowing
Google's Core Web Vitals define "good" as Largest Contentful Paint of 2.5 s or less, Interaction to Next Paint of 200 ms or less and Cumulative Layout Shift of 0.1 or less, measured at the 75th percentile of real users. INP replaced First Input Delay as the responsiveness metric in March 2024. Time to First Byte is not a Core Web Vital, but web.dev recommends 800 ms or less, and a much tighter internal budget (many engineering teams aim for 100 to 200 ms on cached landing pages) gives LCP the headroom it needs. Treat that as an engineering target, not a ranking requirement.
Structured data: what it does and does not do
Schema markup does not raise rankings by itself. It helps search engines disambiguate entities and can qualify pages for rich results. The landscape narrowed in August 2023, when Google limited FAQ rich results to authoritative government and health sites and retired HowTo rich results; FAQPage markup on a commercial page still validates but rarely produces anything visible. Spend the effort on Article, Product, Organization and BreadcrumbList markup, and on entity consistency across the site, rather than on site-wide FAQ blocks.
A note for dev.to readers: the platform strips <script> tags from posts, so any JSON-LD pasted into an article is discarded. Page-level structured data on a publishing platform belongs to the platform, not the author.
Large sites also need crawl budget management, and Google's own guide on the subject is the reference. This is where an enterprise technical SEO review earns its keep: the failure modes at a million URLs are different from those at a thousand.
Measuring ROI: a conservative forecast model
Rankings are a leading indicator, not an outcome. Fix the measurement plan before the first brief is written and track what paid media is judged on: qualified sessions, conversion events, pipeline or revenue, and cost.
A forecast for a cluster is a chain of multiplications, each with an explicit assumption:
Expected clicks = Demand × CTR at target position × P(rank)
Conversions = Expected clicks × conversion rate for that page type
Expected value = Conversions × average deal or order value
Return = Expected value − (content + engineering + link acquisition cost)
The weakest link is CTR. Published click-through curves vary enormously with SERP features; a query with Shopping units, a video carousel and an AI summary above the organic results has a different curve from a plain ten-link page. Derive the curve from your own data instead:
import pandas as pd
gsc = pd.read_csv("gsc_queries_last_90_days.csv")
gsc["pos_bucket"] = gsc["position"].round().clip(1, 20).astype(int)
ctr_curve = (
gsc.groupby("pos_bucket")[["clicks", "impressions"]]
.sum()
.assign(ctr=lambda d: d["clicks"] / d["impressions"])
)
print(ctr_curve["ctr"].round(3))
Segment the curve by branded versus non-branded queries and by SERP type where you can. Then forecast at the 25th percentile rather than the mean, and print the assumptions next to the number. A forecast that survives a sceptical CFO is one that shows its inputs.
| Stage | Leading indicators | Lagging indicators |
|---|---|---|
| 0 to 6 weeks | Indexed URLs, impressions, average position | None yet |
| 6 to 16 weeks | CTR, clicks, striking-distance movement | Qualified and engaged sessions |
| 4 to 12 months | Share of voice across the cluster | Conversions, pipeline, revenue, acquisition cost |
Limitations of this framework
- Volume figures from third-party tools are modelled estimates built on sampled clickstream data; treat them as ordinal, not cardinal.
- Zero-click results, AI-generated summaries and expanding SERP features reduce organic CTR unevenly across query types, and the effect is still shifting.
- Intent weights, effort scores and the ranking-probability curve are judgement calls; the model makes them explicit, it does not make them objective.
- The scoring model ignores brand effects, seasonality and the compounding value of topical authority, all of which matter.
- The link between rankings and revenue is confounded by everything else marketing does; use holdout clusters or time-series comparisons where you can.
Keyword research FAQ
What is keyword research?
Keyword research is the process of identifying the queries people type into search engines, estimating their demand and intent, and deciding which of them a site should target with which pages. Done well it is a prioritisation exercise tied to business outcomes, not a list-building one.
How is keyword research different for B2B and ecommerce?
B2B research is dominated by low-volume, high-value queries about implementation, integration and evaluation, and by long sales cycles in which informational content assists rather than closes. Ecommerce research is dominated by category and product modifiers, seasonality and facet architecture, and by the need to keep crawlers focused on indexable paths.
How many keywords should one page target?
One primary intent, expressed by a cluster of closely related queries. A single page can and should rank for dozens of variants of the same need; it should not try to serve two different needs.
How long does keyword research take to show results?
The research itself takes days to a few weeks. Ranking impact depends on what you do with it: refreshing striking-distance pages can move within weeks, while new pages on competitive terms typically need several months and supporting links. Phase launches and measure leading indicators early.
Does keyword research still matter for AI search and answer engines?
Yes, with a shift in emphasis. Conversational queries are longer and more specific, and answer engines cite pages that state facts clearly, cover entities consistently and are technically accessible. The same intent-first, well-structured content that ranks in classic search is what gets cited; the research adds question-form and entity-level targets to the list.
Which free tools can a developer start with?
Google Search Console for demand and performance, Google Trends for seasonality, the Keyword Planner inside Google Ads for volume ranges, Bing Webmaster Tools for a second first-party view, Lighthouse and the Chrome UX Report for performance, and a crawler such as Screaming Frog's free tier. Add Python with pandas and you have the entire pipeline in this article.
References
- Google Search Central: Creating helpful, reliable, people-first content
- Google: Search Quality Evaluator Guidelines (PDF)
- Google Search Central: A guide to Google Search ranking systems
- web.dev: Web Vitals and Time to First Byte
- Google Search Central: FAQPage structured data
- Google Search Central: Large site owner's guide to managing your crawl budget
- Google Search Console Help: Performance report (Search results)
Disclosure: the author works at RankJockey, an SEO agency. This article is an expanded, research-oriented version of a guide first published on rankjockey.com. If you would rather see the framework applied than build it yourself, the keyword research service page describes the deliverables.
`
Top comments (0)