DEV Community

vyixor
vyixor

Posted on

Two weeks of trying to teach Google my site isn't a movie streaming platform

The domain is soroflix.xyz.

I bought it about six months ago for a movie rating side project. Something like a personal TMDB with watch links and ratings. Built a bit of it, lost interest, moved on.

Then I started working on something completely different. A set of browser based file tools. PDF editor, EXIF remover, image editor, that kind of thing. Everything runs locally in the browser, nothing gets uploaded. When the tools were ready to ship, the movie idea was dead. So I renamed the project SORO File Tools and pointed it at the domain I already owned.

Which is how I ended up with a file tools site on a domain called soroflix.

Three weeks after launch, I opened Search Console for the first time. Around 22 impressions, 4 clicks. Fine for a brand new site. Then I looked at the Queries tab.

One query. "soro films."

Every impression. Every country. Every day. People searching for movie streaming, landing on a page about stripping GPS coordinates from photographs.

Google read "flix" and stopped there.

What I did about it
I spent the next two weeks trying to override the domain signal with content signals. Here is the actual list, in the order I tried things.

Rewrote every title
Before:

<title>SORO — Free file tools that never see your files</title>
Enter fullscreen mode Exit fullscreen mode

After:

<title>SORO File Tools — Free browser-based PDF, image, text utilities</title>
Enter fullscreen mode Exit fullscreen mode

The hypothesis was simple. Google builds an entity from repeated mentions of a name across pages. If every title says the same two word brand, the classifier has more material to work with. Ten pages, ten titles, all leading with the exact string.

Rewrote every meta description to include the full brand
Before:

<meta name="description" content="EXIF remover, PDF editor, image editor...">
Enter fullscreen mode Exit fullscreen mode

After:

<meta name="description" content="SORO File Tools is a free suite of ten browser-based utilities...">
Enter fullscreen mode Exit fullscreen mode

Same idea. Every meta description now contains the exact phrase. Combined with the titles, that is 20 mentions of the full brand name across the site before Google even parses the body content.

Changed the nav and footer logos
Every page used to say:

<a href="/" class="nav-logo">SORO</a>
Enter fullscreen mode Exit fullscreen mode

Now it says:

<a href="/" class="nav-logo">SORO File Tools</a>
Enter fullscreen mode Exit fullscreen mode

That sounds cosmetic. It is not. The nav logo appears on every single page, in the same position, at the top of the DOM. It is the single most repeated string on the entire site. Making it say the full brand name means every page reinforces the same entity ten times per page load.

Added alternateName to the WebSite schema
The original schema was minimal:

{
  "@context": "https://schema.org",
  "@type": "WebSite",
  "name": "SORO",
  "url": "https://soroflix.xyz"
}
Enter fullscreen mode Exit fullscreen mode

Rewritten to include every variant:

{
  "@context": "https://schema.org",
  "@type": "WebSite",
  "name": "SORO File Tools",
  "alternateName": ["SORO", "Soroflix", "SORO File Utilities", "SORO Tools"],
  "url": "https://soroflix.xyz/",
  "publisher": {
    "@type": "Organization",
    "name": "SORO File Tools",
    "url": "https://soroflix.xyz/"
  }
}
Enter fullscreen mode Exit fullscreen mode

The alternateName array tells Google that all four names point to the same entity. This is the piece I am most confident in, because it directly addresses the classifier problem. Google's entity model already knows what Soroflix is. The alternateName block is how you tell it what Soroflix actually means.

Added an ItemList schema to the homepage
This one is direct. It tells Google exactly what the site contains:

{
  "@context": "https://schema.org",
  "@type": "ItemList",
  "name": "SORO File Tools Catalog",
  "numberOfItems": 10,
  "itemListElement": [
    { "@type": "ListItem", "position": 1, "name": "EXIF Remover", "url": "..." },
    { "@type": "ListItem", "position": 2, "name": "EXIF Viewer", "url": "..." },
    // ... 8 more
  ]
}
Enter fullscreen mode Exit fullscreen mode

Each tool gets its own ListItem with a name and a URL. Google can now read the site as "a catalog of file utilities" from the schema alone, without needing to parse the HTML.

Added isPartOf to every tool page schema
Each tool already had a WebApplication schema. I added a back reference:

{
  "@type": "WebApplication",
  "name": "PDF Editor",
  "url": "https://soroflix.xyz/pdf-editor/",
  "isPartOf": {
    "@type": "WebSite",
    "name": "SORO File Tools",
    "url": "https://soroflix.xyz/"
  },
  "publisher": {
    "@type": "Organization",
    "name": "SORO File Tools",
    "url": "https://soroflix.xyz/"
  }
}
Enter fullscreen mode Exit fullscreen mode

Now every tool page explicitly tells Google "I am part of SORO File Tools." Ten pages, ten back references. That is a lot of signal all pointing at the same parent entity.

Added a definition paragraph to the homepage
The single most direct thing on the site:

<p>
  SORO File Tools is a free collection of browser-based utilities 
  for PDFs, images, text files, and links. Every tool runs locally 
  on your device with no uploads and no accounts.
</p>
Enter fullscreen mode Exit fullscreen mode

Plain text, right under the hero. Not hidden, not in the footer. Google reads page content, and this is the sentence that tells it, in human language, what the site is.

What actually happened
Two weeks later, Search Console still shows "soro films" as the only query.

Impressions have ticked up slightly. Around 40 now. Still zero organic clicks from outside Benin, where I live and where my friends are. The title changes have not shown up in the SERP yet either, which the Search Console documentation says can take one to three weeks for a fresh crawl.

Some things are moving. Ranked number two for "soro" on SaaSHub. AlternativeTo approved the listing. A few Dev.to posts are pulling visitors. But Google, which is the one that matters, has not updated its classification.

The honest answer is I do not know if this is going to work. The theory is sound. The execution is thorough. But Google's entity model is a black box and I have no way to know if I am on a two week timeline or a six month one.

What I learned about the entity model
A few things worth knowing if you ever end up in this situation.

Domain names carry more weight than I expected. The classifier runs before the content parser. If the domain has a strong enough prior, the content never gets a chance to override it in the initial pass.

Schema is a signal, not a command. Adding alternateName does not force Google to update its entity model. It gives Google more material to work with the next time it re-evaluates the entity, which happens on Google's schedule, not yours.

Every repeated string helps. Title, meta description, nav logo, footer, schema name fields, body copy. The classifier is looking for a consistent pattern. The more places the same phrase appears, the stronger the signal.

You can't test this locally. Unlike most web development, there is no way to see what Google sees. You submit, you wait, you check. Each cycle is days or weeks. It is the slowest feedback loop I have ever worked in.

The thing that is still broken
The sitemap has never been successfully fetched by Google. sitemap.xml returns a 200 in the browser, renders as valid XML, passes every validator I have tried. Search Console says "could not fetch" every single time.

It is hosted on Cloudflare Pages. I have tried removing it and resubmitting. I have tried submitting the full URL instead of the path. I have checked the content-type header, which is correct. I have disabled Bot Fight Mode in Cloudflare, which was the most common fix suggested.

Still could not fetch. If anyone has dealt with this specific problem on Cloudflare Pages, I would genuinely like to know what fixed it.

The open question
Has anyone else tried to fix a Google entity misclassification where the domain name was the primary signal. How long did it take for the entity to update. Did it stick, or did the classifier drift back to the domain signal after a while.

I am not looking for "buy a new domain." I know. I own this one and I want to make it work.

I am looking for the person who has actually walked through this and can tell me what the timeline looks like. Or the person who tried and failed, so I know when to give up and start over.

Top comments (0)