DEV Community

Karma Domains
Karma Domains

Posted on

Your Link-Building Database Is Decaying: How I Audit Donor Sites with MCP

By Lucy Kim, founder of Karma.Domains

A spreadsheet full of publishers looks reassuring. It has domains, prices, contacts, DR, traffic, and neat status labels. The problem is that most of those fields describe the day somebody checked the site, not the site as it exists now.

A donor that passed review six months ago may have lost traffic, changed owners, filled its archive with an unrelated topic, or started publishing every category from crypto to pet insurance. The placement itself may have disappeared while the domain still looks respectable in a third-party metric.

This is why I do not treat a donor database as a catalog. I treat it as a set of decisions that need to be revalidated.

MCP makes that process much less tedious. It lets an AI assistant run tools, compare fresh results with stored values, and return only the records that need attention. It does not replace the link builder. It removes the repetitive part between “we should audit this list” and “here are the 23 domains a person needs to inspect.”

A donor audit has three separate layers

The first mistake is trying to answer every question with one SEO score. Domain Rating, Domain Authority, Authority Score, and similar metrics can be useful inputs. None of them tells you whether your link is still on the page, whether the publisher changed its business model, or whether your team already rejected the site for a client-specific reason.

I split the audit into three layers.

The placement layer covers the actual page and link: HTTP status, final URL after redirects, canonical, indexability, anchor, target URL, and rel attributes. Google recommends qualifying paid placements with rel="sponsored"; nofollow is still acceptable for that purpose. That detail belongs to the placement record, not the domain score (Google Search Central).

The domain layer covers the site around the placement: authority signals, estimated traffic, backlink and referring-domain counts, anchor distribution, spam signals, archive age, page footprint, DNS, and WHOIS data.

The relationship layer covers what the tools cannot infer from a domain: who contacted the publisher, the quoted price, placement terms, previous refusals, client exclusions, and the result of the last negotiation.

The three layers of a donor audit: placement, domain, relationship

Those layers can disagree. A healthy domain can contain a deleted placement. A live followed link can sit on a site that has drifted into an unsuitable niche. A promising publisher can still be the wrong prospect because another teammate contacted the same person last week.

Why I use Karma.Domains MCP for the domain layer

Karma.Domains is best known as an expired domains and auction intelligence platform. Its database contains well over 9 million domains, covers more than 40 sources and many gTLDs and ccTLDs, and exposes more than 90 filters. The platform brings together signals from Ahrefs, Semrush, Moz, Majestic, Similarweb, the Web Archive, and other sources for expired and auction-domain research.

That scale matters, but it is not the same workflow as checking an existing publisher site. For arbitrary live domains, I use the live-check tools available through the Karma.Domains MCP server: authority, traffic, spam score, backlinks, anchor text, archive age, page count, DNS, and WHOIS. I do not query the auction reports endpoint and pretend it is a universal domain checker.

ChatGPT and Claude can connect through OAuth, so the setup is closer to signing into a service than building an integration. Once connected, I can work in normal language while the assistant calls the appropriate live tools.

The practical advantage is orchestration. Instead of copying a domain through eight checker pages, I can give the assistant a batch, specify the review policy, and ask for a change report. The final judgment remains mine.

Step 1: Make the old decision comparable

An audit needs a baseline. A row with DR 42 but no check date is barely historical data. It might be yesterday's value or a number imported three years ago.

At minimum, I keep these fields:

  • root domain;
  • placement URL and target URL;
  • client or project;
  • workflow status;
  • first checked and last checked dates;
  • previous metrics and their sources;
  • contact and owner details;
  • price and placement terms;
  • approval or rejection reason;
  • exclusions, notes, and next action.

The rejection reason is especially important. “Rejected” is not a reusable decision if nobody knows whether the cause was price, topical mismatch, a suspicious link profile, or simply no response.

My first prompt is deliberately administrative:

I’ll upload a CSV of donor domains. Normalize the root domains, flag duplicates, and show which records do not have a last_checked date or a previous metric baseline. Do not score the domains yet. Preserve all client exclusions and notes.

This catches duplicate www and protocol variants, missing timestamps, and records that cannot be compared responsibly. It also prevents the AI from rushing into a verdict before the data model is usable.

Step 2: Check the placement before judging the domain

If the database contains existing placements, I start with the URL. A crawler or backlink-monitoring tool should establish whether the page returns 200, redirects, is gone, has become noindex, still contains the link, and still points to the intended destination.

Ahrefs distinguishes several reasons for a lost link, including a removed link, a missing page, a redirect, noindex, and crawl errors. It also warns that some crawl failures are temporary (Ahrefs Help Center). Semrush similarly separates new, lost, and broken referring domains in its Lost and Found report (Semrush).

That distinction matters. A timeout is not proof that a placement vanished. I route temporary crawl failures to a retry state rather than marking the donor as dead.

Google Search Console is useful as another input, but not as the master ledger. Its Links report is a sample, may include links that have since been removed, and does not show whether a link is nofollow (Google Search Console Help). I use it for discovery and corroboration, not for a complete placement audit.

The output of this pass should be factual: live, removed, redirected, noindexed, rel changed, or retry crawl. No domain-quality conclusion yet.

Step 3: Run live domain checks through MCP

I do not spend the same amount of checking capacity on every row. I start with domains tied to active or expensive placements, publishers the team plans to contact again, records that have not been checked recently, and domains where the placement pass found a change.

Then I give the assistant an explicit task:

Run live checks for authority, traffic, spam score, backlink profile, anchor text, archive age, page count, DNS, and WHOIS for these domains. Compare current values with the CSV baseline. Treat missing data as unknown, not zero. Return the old value, new value, absolute change, percentage change where meaningful, and source for each field.

The wording matters. “Find bad donors” invites a black-box judgment. “Show the changes” produces an auditable result.

I look for combinations rather than a single red flag:

  • a material traffic decline alongside a shrinking page footprint;
  • a new concentration of casino, payday, pharma, or unrelated commercial anchors;
  • an unusual change in backlinks or referring domains;
  • a WHOIS or DNS change combined with a topical shift;
  • a very small site after what appears to be a large content purge;
  • an archive footprint that no longer matches the publisher presented in the CRM.

These are investigation triggers, not automatic convictions. Estimated traffic can fluctuate. Providers update on different schedules. A missing result means “we do not have a value,” not “the value is zero.”

There is another boundary worth keeping clear: the live age check can describe the Web Archive footprint, such as first and last snapshots and snapshot count. It does not turn an auction-domain report workflow into a complete historical content audit for every arbitrary live site. If topical history is decisive, I open the archive and review it.

Step 4: Classify the change, not the publisher

I use three operational states:

Keep means the placement and domain show no material change under the team's policy.

Re-review means something changed, the data is incomplete, or the signals conflict. A specialist needs to inspect the record.

Freeze means the team should pause new outreach or purchases until somebody resolves a specific problem. It is not a permanent blacklist.

Here is a prompt I can reuse with different clients:

Apply the policy below and return Keep, Re-review, or Freeze. For every non-Keep result, list the exact rule that fired, the previous value, the current value, and the source. Never infer zero from missing data. Do not recommend disavow. Do not change CRM exclusions or contact anyone.

The policy itself belongs to the team. A news publisher, a local legal client, and an iGaming affiliate will not have identical thresholds or compliance rules. The assistant can apply a policy consistently; it should not invent one silently.

Step 5: Spend deep-research credits on the short list

Once the change report is ready, the expensive tools become more useful. I can open Ahrefs, Semrush, or DataForSEO for the small set of high-value or changed domains rather than running a full investigation on every row.

That second pass is where I examine link velocity, lost-link reasons, ranking distribution, country and keyword shifts, competing pages, or a suspicious cluster in detail. It is also where a human reads the site.

The person reviewing the shortlist should still decide:

  • whether the publication is genuinely relevant to the client's audience;
  • whether its editorial standards are acceptable;
  • whether the price and terms make sense;
  • whether a change is temporary or structural;
  • whether outreach should resume;
  • whether any compliance or search-policy issue needs escalation.

Google's spam policies explicitly cover links created primarily to manipulate rankings, including paid links that pass ranking credit (Google Search Central). An automated “good donor” label is not a compliance review.

Use a review policy, not a magical calendar

There is no universal “audit every 30 days” rule. The sensible frequency depends on value, activity, and risk.

I usually think in tiers:

  • Tier A: active prospects, expensive placements, and important live links. Check more frequently and before a new transaction.
  • Tier B: proven publishers that are not in an active campaign. Review periodically.
  • Tier C: archived, refused, or inactive records. Recheck when somebody proposes reusing them.

Events can override the schedule. A lost-link alert, a new campaign, a new owner, a sharp traffic change, or a planned repeat order can trigger a fresh review immediately.

The donor database audit loop

At the end of every completed review, I save the new values, timestamp, sources, status, and reason. That becomes the next baseline. Without that write-back step, the team repeats the same detective work during the next audit.

Guardrails I would not remove

Automation is most helpful when its boundaries are explicit.

I do not let it reduce a site to DR, DA, or AS. I do not let it convert missing data into a failing score. I do not mix up a live placement with a healthy domain. I do not send domains to a disavow file automatically. I do not launch outreach without checking previous contacts and exclusions. And I do not allow a model-generated status to masquerade as the final judgment of an SEO specialist.

The purpose of this workflow is narrower and more valuable: keep a large operational database honest.

A static donor list tells me what the team believed at some point in the past. An MCP-assisted audit tells me what changed, why a record was flagged, and which decisions now need a person. That is the part of link-building operations worth automating.

Top comments (0)