I am not a full time developer. I am a clinical psychologist and I run a non-profit rehabilitation centre in Ukraine. The website is part of the work: people look for help at two in the morning, they read whatever is in front of them, and then they either call or they do not.
That site is WordPress with Polylang, every page exists in Ukrainian and in Russian, and for two years I kept breaking it in the same ways. Not interesting ways. Boring, repeatable ways. A page shipped in one language and not the other. A meta description that was still the theme default. A hero image at 3 MB because nobody resized it. A redirect chain that quietly ate a page that used to rank.
The fix was not discipline. Discipline does not survive a bad week. The fix was turning every rule we had written down in prose into a check that runs before anything is published, and refusing to publish when a check fails.
This post is about that gate: what it checks, how it is built, and which rules only taught us anything after they had already cost us something.
The problem with a written checklist
We had a document. It said things like "every page must have a language pair", "images go in WebP, maximum 1600 px wide", "every image in the media library needs alt text, a title and a caption", "never use the em dash in body text".
A document is a suggestion. It is read once, on the day it is written, by the person who wrote it.
What actually happened is that a page would go live, look fine, and three weeks later we would discover that the Russian version had no meta description, or that the Ukrainian version of the same article pointed at a different phone number because it was written four months apart from its pair.
So the rule became: if a rule matters, it is code. If it is not code, it is not a rule, it is an opinion.
One command before anything else
Everything starts with a single entry point. You give it a URL and it does not let you past until it has looked at the page:
python gate.py https://example.com/some-page/
The gate pulls the page over HTTP, pulls the same post over the WordPress REST API, finds the language pair, and runs every check it knows about. It prints a verdict per check and exits non-zero if anything is red.
The structure is deliberately dumb. One function per rule, every function returns the same shape, and the runner does not know what any of them do:
from dataclasses import dataclass
@dataclass
class Result:
rule: str
ok: bool
detail: str = ""
fixable: bool = False
CHECKS = []
def check(name):
def wrap(fn):
fn.rule = name
CHECKS.append(fn)
return fn
return wrap
def run(page):
results = [fn(page) for fn in CHECKS]
for r in results:
print(("ok " if r.ok else "FAIL "), r.rule, r.detail)
return all(r.ok for r in results)
That is the whole framework. The value is not in the runner, it is in the fact that adding a new rule is four lines and nobody has to argue about where it goes.
The checks that earn their place
The language pair
The single most useful check. Polylang will happily let you publish a post with no translation, and the site will look complete to you because you are reading the language you just wrote.
@check("pair-exists")
def pair_exists(page):
other = page.translations.get(page.other_lang)
if not other:
return Result("pair-exists", False,
f"no {page.other_lang} version for {page.slug}")
return Result("pair-exists", True, other.url)
The second half of this rule took longer to learn. It is not enough for the pair to exist. The pair has to say the same things. We had pages where one language promised a free consultation and the other did not mention it, because one of them was edited six months later.
So the parity check compares the parts that must not drift: the phone number, the price blocks, the list of what the programme includes, the author box, the number of H2 sections.
PARITY_FIELDS = ("phone", "author_slug", "cta_blocks", "h2_count", "has_faq")
@check("pair-parity")
def pair_parity(page):
other = page.translations.get(page.other_lang)
if not other:
return Result("pair-parity", True, "skipped, no pair")
drift = [f for f in PARITY_FIELDS
if getattr(page, f) != getattr(other, f)]
if drift:
return Result("pair-parity", False, "drift in: " + ", ".join(drift))
return Result("pair-parity", True)
Nothing clever. It has caught more real problems than every other check combined.
Images
Three separate rules collapsed into one check, because they always failed together:
MAX_W = 1600
@check("images")
def images(page):
bad = []
for img in page.images:
if img.width > MAX_W:
bad.append(f"{img.name}: {img.width}px")
if not img.src.endswith(".webp"):
bad.append(f"{img.name}: not webp")
if not (img.alt and img.caption):
bad.append(f"{img.name}: no alt/caption")
return Result("images", not bad, "; ".join(bad[:5]), fixable=True)
fixable=True matters. For images the gate can do the work itself: download, resize, convert, write the attachment metadata back through the REST API. For a missing translation it cannot, and it should not pretend it can.
Typography
This one sounds petty and is not. We write in Ukrainian and Russian, and a long dash pasted from a word processor breaks inside our grid layout at narrow widths. It also happens to be the single clearest signal that text was pasted rather than written into the editor.
BANNED = {
"—": "em dash",
"–": "en dash in body text",
" ": "double nbsp",
}
@check("typography")
def typography(page):
hits = [name for ch, name in BANNED.items() if ch in page.body]
return Result("typography", not hits, ", ".join(hits), fixable=True)
A note for anyone doing this in PHP against the WordPress API: do not put the character literally in your source file. Half our tooling runs through shell on Windows and the encoding will not survive the trip. Build it with chr() or an escape and your scripts stop corrupting the content they were supposed to clean.
Redirects before anything else
The rule that cost us the most before it existed: check redirects before you touch the cache.
We merged a batch of duplicate pages, set up the 301s, flushed the cache, and then spent a day wondering why a chunk of pages returned 404 for logged out visitors. The redirect plugin had been fed a relative target in a few rows, the cache had happily stored the 404, and the live site looked fine to us because we were logged in.
@check("redirect-sanity")
def redirect_sanity(page):
chain = follow(page.url, max_hops=5)
if chain[-1].status != 200:
return Result("redirect-sanity", False,
" -> ".join(f"{c.status}" for c in chain))
if len(chain) > 2:
return Result("redirect-sanity", False,
f"{len(chain) - 1} hops before 200")
return Result("redirect-sanity", True)
And the matching operational rule, which is not in the code but is in the runbook in capital letters: never flush the whole cache to "see if that helps". Purge the URLs you touched. A full flush on a site with any traffic means every visitor for the next few minutes gets the slow path, and it hides the actual problem instead of showing it to you.
Four rules we only learned by breaking something
wpautop will rewrite your markup. If you build a layout with a grid of divs and let WordPress run its auto paragraph filter over it, you get stray <p> tags inside flex containers and a layout that is subtly wrong only on some pages. Either disable the filter for the post types where you ship real markup, or stop shipping real markup. Picking neither is the worst option, and that is what we did for about four months.
Facebook does not want your WebP. We converted everything to WebP for weight, which is correct, and then every share preview went blank. Open Graph gets a JPEG copy generated at upload time; the page itself keeps WebP. Two formats, one source image, no manual work.
The featured image is not the hero image. They look like the same thing in the admin and they are used in completely different places. We spent a while with pages where the card in the listing showed one photo and the top of the article showed another, which looks exactly as sloppy as it is.
Reindex immediately, not "later". Every time the gate passes and something goes live, the same script pings the indexing API for both language versions. Doing it by hand means doing it for the page you remember and forgetting its pair.
Did it work
The honest answer is that the gate did not improve anything by itself. It removed a class of mistakes, which is different and less exciting.
What I can measure: both homepages sit at 98 on mobile PageSpeed, which they did not before we made image rules enforceable. The number of pages that exist in one language only is zero, and it stays zero without anyone watching it. We merged a large batch of near duplicate pages that had accumulated over the years, and the thing that made that safe was being able to check every incoming link before the merge rather than after.
What I cannot claim: that traffic went up because of this. Plenty of other things changed in the same period and I am not going to pretend I can separate them.
What I would do differently
Start with the parity check. Not the image pipeline, not the typography, not the performance work. The checks that catch "this page is a lie compared to its twin" are worth more than everything else on the list, and they are the easiest to write.
Second, make every check report rather than fix by default, and only let it fix when you pass a flag. An early version silently rewrote content, which is a great way to lose a paragraph somebody spent an evening on.
Third, keep a log. Every run appends a line: URL, what was checked, what changed, verdict. It is three lines of code and it is the only reason I can answer "when did this page last change and why" six months later.
If you are curious what the actual site looks like, it is ninarkotikam.com, a Ukrainian non-profit that has been running a residential rehabilitation programme since 2006. The content is in Ukrainian and Russian, so the code is probably the more useful part for this audience.
Happy to answer questions about the WordPress and Polylang side of it. That is where almost all the pain was.
Top comments (0)