DEV Community

Daniel Pertu
Daniel Pertu

Posted on

A campaign that has searched every keyword gets eight new ones, and the worst eight switch off

Nakodo finds creators and local businesses for a brand by searching platforms for keywords. The keywords come from a research pass over the brand's website: a brief, a set of target markets, and a list of search phrases grouped by angle.

That list is finite. Every plan caps it, and the caps are published on the pricing page: 8 search keywords per campaign on Free, 25 on Pro, 60 on Business.

So a campaign that runs for long enough reaches a state where every keyword it has has been searched, on every platform, in every market. The searching stops. From the outside that is indistinguishable from the product having broken, and the brand still needs people to email.

Running the same searches again does not help much either: platform search results for a phrase are stable over days. What a searched-out campaign needs is different phrases.

When a round is allowed to happen

The entry point is addAngles(campaignId), and most of it is refusals:

const campaign = await db.query.campaigns.findFirst({ where: eq(campaigns.id, campaignId) });
if (!campaign?.brief) return [];

const { limits } = await planForUser(campaign.userId);
const platforms = searchedPlatforms(campaign.platforms, limits.platforms).filter((p) => p !== "web");
if (platforms.length === 0) return [];

const on = all.filter((k) => k.enabled && platforms.includes(k.platform));
if (on.length === 0 || on.some((k) => !k.lastSearchedAt)) return [];

const [waiting] = await db.select({ id: jobs.id }).from(jobs).where(
  and(eq(jobs.campaignId, campaignId), inArray(jobs.type, ["search", "social_search"]), inArray(jobs.status, ["pending", "running"])),
).limit(1);
if (waiting) return [];
Enter fullscreen mode Exit fullscreen mode

Two of those are the interesting ones.

on.some((k) => !k.lastSearchedAt) means a single never-searched keyword cancels the round. The campaign is not out of ideas, it is behind on work, and adding more inputs to a backlog is the wrong move.

The queue check is the same idea one level up: if a search job for this campaign is pending or running, its results have not landed, so the counts we are about to feed the model are stale. We would be proposing new angles while the evidence for the old ones is still in flight.

Both are cheap, and both convert "ask for more keywords" from something that could be triggered repeatedly into something that can only fire from a quiet, finished state. The caller asks once a day and whenever a campaign runs low on leads; the guards are what make asking that often harmless.

Counting what each keyword actually found

To choose new angles you need to know which old ones worked. Every creator or business a campaign has found carries the keywords that matched it, in an array column. The counts come out in one query:

const counts = await db.execute<Tried>(sql`
  select lower(k) as keyword, count(*)::int as found,
         (count(*) filter (where ${sql.raw(fitCondition("cc"))}))::int as fits
  from ${campaignChannels} cc, unnest(cc.matched_keywords) k
  where cc.campaign_id = ${campaignId}
  group by 1`);
Enter fullscreen mode Exit fullscreen mode

Three bits of Postgres worth stealing.

from campaign_channels cc, unnest(cc.matched_keywords) k is an implicit lateral join. The set-returning function in the FROM clause can refer to a column of the table before it, so each row expands into one row per array element, and group by 1 then groups by the element rather than by the array. No jsonb, no normalised join table, no application-side counting.

count(*) filter (where ...) is the standard-SQL aggregate filter, and it is both shorter and clearer than sum(case when ... then 1 else 0 end). Two aggregates over one scan.

fitCondition("cc") is a function that returns the SQL fragment for "this one is a good fit", taking the table alias as its argument. The same rule is applied in TypeScript elsewhere, and keeping one generator for the SQL form means the place that counts fits and the place that shows fits cannot drift apart. (What is in that fragment is the part of the product we do not publish.)

The result is a map from keyword to { found, fits }, which is then attached to every keyword the campaign has ever had, sorted best first, and capped:

const tried = [...new Set(all.map((k) => k.keyword.trim().toLowerCase()))]
  .map((keyword) => ({ keyword, ...statsOf(keyword) }))
  .sort((a, b) => b.fits - a.fits || b.found - a.found);
Enter fullscreen mode Exit fullscreen mode

Note all, not on: retired and disabled keywords are included. A keyword that found nothing is evidence, and dropping it from the list is how you get the model to propose it again.

What the model is asked for, and what is done with the answer

The model gets the brief, the target markets, the platforms, the things the brand does not want, and the tried list with its two numbers per keyword, capped at the best 200. It returns groups:

const anglesSchema = z.object({
  keywordGroups: z.array(z.object({
    theme: z.string().describe("The angle, in 2 to 4 words"),
    languageCode: z.string().describe("ISO 639-1 language the keywords are written in"),
    keywords: z.array(z.string()),
  })),
});
Enter fullscreen mode Exit fullscreen mode

A structured-output schema is not a validator. It constrains shape, not content, so everything that comes back is washed:

export function cleanGroups(groups, languages: string[]) {
  const tidy = (s: string) => s.replace(/["“”#]/g, "").replace(/\s+/g, " ").trim();
  return groups
    .map((g) => {
      const code = g.languageCode.trim().toLowerCase();
      return {
        theme: tidy(g.theme).slice(0, 60) || "New angle",
        languageCode: languages.includes(code) ? code : languages[0],
        keywords: g.keywords.map(tidy).filter((k) => k.length > 1 && k.length <= 80),
      };
    })
    .filter((g) => g.keywords.length > 0);
}
Enter fullscreen mode Exit fullscreen mode

Quotes and hash signs come off because a model that has seen a million search tips likes to emit "exact phrase" and #hashtags, and both change what a platform search means. A language code outside the campaign's own target markets is replaced by the campaign's first one rather than trusted: we have a closed set, so there is no reason to accept a value outside it. A group whose keywords all fail the length check disappears entirely instead of becoming an empty theme.

Then the round is deduplicated against everything the campaign has ever searched, by a key that defines what "the same search" means:

export const searchKey = (k: { keyword: string; regionCode: string | null; relevanceLanguage: string | null }) =>
  `${k.keyword.trim().toLowerCase()}|${k.regionCode ?? ""}|${k.relevanceLanguage ?? ""}`;

const seen = new Set(all.map(searchKey));
for (const row of keywordRows(campaignId, groups, brief.targetMarkets, places)) {
  if (fresh.length >= size || seen.has(searchKey(row))) continue;
  seen.add(searchKey(row));
  fresh.push(row);
}
Enter fullscreen mode Exit fullscreen mode

The same phrase in two markets is two searches. The same phrase in the same market is one, whatever the casing. seen.add inside the loop matters because a model asked for eight keywords will sometimes give you the same keyword twice in two different groups.

The nine lines that make room

A round is Math.min(8, limits.keywordsPerCampaign) keywords. Adding them would take the campaign over its plan's cap, so something has to switch off, and this is the whole rule:

// The searches to turn off so `adding` more fit within `limit`: the ones that
// found the fewest good fits, then the fewest at all; the oldest of equals.
export function toRetire<T extends { fits: number; found: number }>(searches: T[], adding: number, limit: number): T[] {
  const over = searches.length + adding - limit;
  if (over <= 0) return [];
  return [...searches].sort((a, b) => a.fits - b.fits || a.found - b.found).slice(0, over);
}
Enter fullscreen mode Exit fullscreen mode

Things I like about this function, which is the most-tested thing in the feature:

It is pure and generic over { fits, found }, so its tests are a table of numbers with no database and no model.

over <= 0 returns an empty array rather than falling through to a sort and a slice(0). Campaigns under their cap are the common case, and "nothing is retired" is a stated outcome, not an accident of arithmetic.

The tie-break chain ends in the input order, which is createdAt ascending from the caller's query. Array#sort has been required to be stable since ES2019, so the oldest of several equally useless keywords is the one that goes. That is a load-bearing use of sort stability, which is worth a comment in the code and a line in the tests, because if it is ever broken you get a different keyword retired and no error.

Retiring is enabled: false, never a delete:

const retired = toRetire([...searches.values()], fresh.length, limits.keywordsPerCampaign).flatMap((s) => s.ids);
if (retired.length > 0) await db.update(keywords).set({ enabled: false }).where(inArray(keywords.id, retired));
Enter fullscreen mode Exit fullscreen mode

The row stays, its results stay attributed to it, and it keeps appearing in the tried list with its real numbers. A deleted keyword would make the next round's evidence a lie.

One consequence we accepted rather than engineered around: on the Free plan, the cap is 8 and a round is 8, so a round is a complete turnover of the campaign's live searches. That is the correct reading of the cap. The plan buys you eight searches at a time, not eight searches ever, and the eight that get replaced are by definition the eight that found the least.

Finally, the keywords are written once per platform, and only the ids that were actually inserted get queued:

const added = await db
  .insert(keywords)
  .values(fresh.flatMap((k) => platforms.map((platform) => ({ ...k, platform, source: "angle" as const }))))
  .onConflictDoNothing()
  .returning({ id: keywords.id });
Enter fullscreen mode Exit fullscreen mode

onConflictDoNothing().returning() is the pattern that makes a job-queueing insert safe to run twice: whatever a unique index rejected is simply absent from the returned rows, so nothing gets a duplicate search job. source: "angle" marks where they came from, which is how the app can say that a campaign is searching angles it found rather than the ones the brand's site suggested.

The honest trade-offs

Angle keywords drift from the brief. That is the point, and it is also the risk: the fifth round of angles for a kettle brand is further from kettles than the first. Three things hold it in: the brief itself is never rewritten and is sent with every request, the brand's own negative keywords go along too, and the retirement rule is pure selection pressure. An angle that produces nothing is gone within a round or two.

The cost is bounded by construction. A campaign's live search surface is always exactly the plan's cap, so a campaign that has run for a year is not searching more than one that started last week. It is searching differently. The expensive resources in this system are platform API quota and search credits, and both of them care about how many searches are live, not how many have ever existed.

And it is allowed to do nothing. Half a dozen checks in this function return an empty array, which means "not now" and leaves the campaign exactly as it was. In a pipeline that runs unattended on a schedule, a function whose most common answer is "nothing to do" is a feature.

If you want to see the inputs this all starts from, the guide on how to find YouTube influencers is the human version of what the first round of keywords is trying to do, and the pricing page has the per-plan caps the retirement rule enforces.

Top comments (0)