DEV Community

Олександр Дуднев
Олександр Дуднев

Posted on AI-assisted

Google keyword volumes are full of spikes. Here's how I clean them in Python

I'm building a small keyword research tool, and the first thing that broke was my trend logic. Not the code. The data.

Here is what Google's keyword data says about "how to start a blog" in the US:

For a year it sits around 12,000 searches a month. Then June 2025 reports 301,000, July 1,500,000 and August 165,000. In September it is back to 8,000.

Now compute the usual metrics on that series:

  • Year-over-year (this August vs. last August): −96%. The data provider's own yearly trend field says −96 too.
  • Typical month, this year vs. last year: 13,450 → 7,350, or −45%.

Both numbers describe the same keyword. Interest really is falling, but −96% would tell you the topic is dead, and it isn't.

This is not a one-off. In a sample of 21 common US keywords, 4 had a recent month at least 4× their typical month. "web hosting" is the worst:

"web hosting" (US)
Google's 12-month average 110,000
Typical month (median) 30,100
Raw YoY (Aug 2026 vs. Aug 2025) +509%
YoY of typical months +11%

I don't know what causes these spikes. It could be bot traffic, a data artifact or a real one-off event. For keyword research it doesn't matter much: a single month shouldn't be allowed to drive your volume, trend or seasonality numbers. Here are the four fixes I ended up with. Each one is a few lines of standard-library Python.

1. Typical month instead of average, plus a spike flag

Google's search_volume is an average of the last 12 months, so one spike inflates it. The median ignores it.

from statistics import median

def has_spike(last_12, ratio=4):
    """True when one recent month is at least `ratio` times the typical month."""
    return max(last_12) >= ratio * median(last_12)

typical = median(last_12)
Enter fullscreen mode Exit fullscreen mode

I keep both numbers: the average is what every SEO tool shows, the median is what you should plan with, and the flag tells you when they disagree.

2. Trend with a Theil–Sen slope

A least-squares trend line gets dragged by a spike. Theil–Sen takes the slope between every pair of months and uses the median slope. With 12 months there are 66 pairs, and a single month takes part in only 11 of them, so it can't move the median much.

import math
from statistics import median

def trend_pct_per_month(last_12):
    """Typical month-over-month change in %, from the Theil-Sen slope of log volume."""
    logs = [math.log1p(v) for v in last_12]
    slopes = [(logs[j] - logs[i]) / (j - i)
              for i in range(len(logs)) for j in range(i + 1, len(logs))]
    return math.expm1(median(slopes)) * 100
Enter fullscreen mode Exit fullscreen mode

Working in log space makes the slope a percentage, so a keyword with 500 searches and one with 500,000 are judged on the same scale. I call anything above +3% a month "rising" and below −3% "falling".

3. Year-over-year on medians

Comparing one month to the same month a year ago is exactly what spikes break. Compare the typical month of the last 12 months with the 12 before:

from statistics import median

def yoy_pct(last_24):
    """Typical month this year vs. the year before, in %. last_24 is oldest first."""
    before, now = median(last_24[:12]), median(last_24[12:])
    return (now - before) / before * 100
Enter fullscreen mode Exit fullscreen mode

It also fixes the opposite problem. "air fryer recipes chicken" had a quiet summer in 2025 (2,900 a month), so a raw August-to-August comparison says +524%. The typical month actually went from 11,250 to 7,350: −35%.

4. Seasonality as recurrence

The naive rule, "the peak month is 1.5× the average", marks every spiky keyword as seasonal. Real seasonality repeats: the peak lands in the same month every year with roughly the same strength. So the test is:

  • in each of the last three years, find the peak month and its strength (peak ÷ that year's median);
  • ignore years whose peak is below 1.5×;
  • call it seasonal if two or more years peak in the same month (±1) with strengths within 4× of each other.
from statistics import median

def seasonal_peak(last_36):
    """last_36: (month_number, volume) pairs, oldest first. Returns the peak month or None."""
    peaks = []
    for start in (24, 12, 0):                     # most recent year first
        year = last_36[start:start + 12]
        month, top = max(year, key=lambda mv: mv[1])
        strength = top / median(v for _, v in year)
        if strength >= 1.5:
            peaks.append((month, strength))
    for month, strength in peaks:
        recurring = [s for m, s in peaks
                     if min(abs(m - month), 12 - abs(m - month)) <= 1   # same month, +-1
                     and max(s, strength) / min(s, strength) <= 4]      # similar strength
        if len(recurring) >= 2:
            return month
    return None
Enter fullscreen mode Exit fullscreen mode

The strength check is what separates "web hosting" from a real season. It has a 1.8× bump in June 2025 and an 18× spike in May 2026. Same month ±1, but one is ten times stronger than the other, so it's an anomaly, not a season. Keywords that do come out seasonal look right: "coffee maker" peaks in November, "noise cancelling headphones" in December, and the German "steuererklärung online" (filing a tax return online) in January.

Where to get the data in bulk

You need monthly volumes for at least 24–36 months. Three options:

Google Ads API (Keyword Planner). It's free, but it's meant for advertisers. Google's Required Minimum Functionality policy says that a tool offering KeywordPlanIdeaService to third parties must also implement campaign creation, management and reporting. Fine for internal scripts, not for a public keyword tool.

DataForSEO. Pay-as-you-go. Their Labs keyword_overview endpoint costs $0.012 per request plus $0.00012 per keyword (up to 700 keywords per request) and returns volume, CPC, competition, up to several years of monthly history, keyword difficulty and intent. The minimum deposit is $50. Their database doesn't know every long-tail phrase: in my 80-keyword test it found 95% of short and mid-length keywords and 85% of questions, but only a quarter of 6–10-word phrases. Google's live Keyword Planner had no data for those either.

The Apify actor I built (disclosure: it's mine). It wraps DataForSEO and computes everything above: median_monthly_searches, has_spike, trend_direction, yoy_change_pct, is_seasonal and peak_month, next to the usual volume, CPC, competition, difficulty and intent. It costs $6 per 1,000 keywords that come back with data, keywords without data are free, and there's no deposit. Apify's free plan includes $5 of monthly credit, which covers roughly 800 keywords.

pip install apify-client
Enter fullscreen mode Exit fullscreen mode
from apify_client import ApifyClient

client = ApifyClient("YOUR_APIFY_TOKEN")  # Apify Console > Settings > API & Integrations
run = client.actor("keywordlab/keyword-search-volume").call(
    run_input={"keywords": ["web hosting", "how to start a blog", "coffee maker"], "country": "US"}
)
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item["keyword"], item["search_volume"], item["median_monthly_searches"],
          item["has_spike"], item["trend_direction"], item["yoy_change_pct"], item["peak_month"])
Enter fullscreen mode Exit fullscreen mode

Output:

coffee maker 201000 165000 False falling 0 November
how to start a blog 8100 7350 False falling -45.4 None
web hosting 110000 30100 True rising 11.1 None
Enter fullscreen mode Exit fullscreen mode

It works for 28 countries and up to 10,000 keywords per run: apify.com/keywordlab/keyword-search-volume.

Takeaway

Never compute a keyword metric from a single month. Use the median for volume, a Theil–Sen slope for trend, medians again for year-over-year, and require a season to repeat before you call it one. It's a few dozen lines of code, and it turns "+509%" back into "+11%".

If you work with keyword data and have a theory about where these spikes come from, I'd like to hear it in the comments.

Top comments (0)