The hidden grammar of content moderation turns institutional judgments into technical facts - and makes the decision-maker disappear.

By Agustin V. Startari
TL;DR
• Moderation notices are written so that the post, not the platform, becomes the grammatical actor: “Your post violated our policy.”
• Posts do not violate policies by themselves. An institution defines the category, applies it, and - as appeals prove - can reverse it.
• Automation changes the executor, not the author. A system can detect a phrase; only an institution can define a violation.
• The disappearance is measurable: the Censor-Deletion Rate (how often the actor vanishes) and the Suppression Opacity Index (how much of the decision chain a user can reconstruct).
You open Instagram, Facebook, TikTok, or YouTube and find a message:
Your post violated our policies.
It looks factual. Almost mathematical. There was a rule, your post violated it, the platform applied the consequence. Case closed.
Except that something important has disappeared from the sentence: who decided that your post violated the rule?
The post did not read the policy. It did not classify itself, compare its own language against a database, weigh whether context mattered, or select the sanction. An institution did those things - directly, or through systems it designed, authorized, configured, and maintains. Yet the final sentence transforms an institutional judgment into something that looks like a property of the post itself.
That transformation is not a minor linguistic curiosity. It is one of the defining features of modern platform governance. I call it censorship without a censor.
The term does not mean that every act of moderation is censorship. Platforms need rules, and some material should clearly be restricted: fraud, exploitation, direct threats, targeted harassment, genuine incitement. The problem is narrower and stranger than that.
A platform can restrict speech while describing the restriction in a way that makes the institution responsible for the decision grammatically optional.
The speech remains visible as a violation. The sanction remains visible as a consequence. The censor disappears.
Watch What Happens to One Sentence
Consider three versions of the same moderation event.
Version 1: We removed your post because we determined that it violated our policy.
Everything is visible. There is an actor (we), an institutional judgment (we determined), an action (we removed), and a rule (our policy).
Version 2: Your post was removed because it violated our policy.
The post remains. The policy remains. The violation remains. The removal remains. But the actor responsible for the removal is gone. Who removed it? The answer is obvious from context - grammar simply no longer requires anyone to appear.
Version 3: Your post violated our policy.
Now something more significant has happened. The removal disappeared. The classifier disappeared. The decision disappeared. The post itself became the grammatical violator.
An entire institutional chain has been compressed into a single relationship: post → violation. But that is not how moderation works. The real structure looks more like institution → policy → classification → decision → enforcement. The sentence shown to the user reduces all of it to content → violation.
**Posts Do Not Violate Policies by Themselves
**A piece of content does not naturally belong to a category called hate speech, dangerous content, misinformation, extremism, incitement, or prohibited support. Those categories have definitions. Definitions require boundaries. Boundaries require institutional decisions.
Platforms must decide what counts and what does not, which exceptions apply, how context should be treated, which languages require different interpretation, what confidence threshold an automated system should use, and what consequence should follow from classification. Researchers have long shown that contemporary moderation is not a person reading posts and pressing delete: it is an institutional system combining policies, human reviewers, automated detection, machine-learning models, internal procedure, and enforcement machinery (Gillespie, 2018; Gorwa et al., 2020; Roberts, 2019).
So when a platform says “this post violates our policy,” the more revealing sentence would often be: “We determined that this post falls within a category that our policy defines as prohibited.” That version sounds less natural. It also exposes what the shorter one hides - a judgment occurred.
**Where the Compiled Rule Enters
**I have described a related mechanism in earlier work as the compiled rule (Startari, 2025a, 2025b). The idea is simple. An institution begins with a decision: content classified as X should receive consequence Y. But a platform cannot reconstruct the full institutional debate every time a user posts something. The decision has to become executable.
So the sequence hardens: the institution decides, the policy defines, operational criteria translate the policy, a human or automated system classifies, and the rule triggers a consequence. Eventually the whole chain can run as: if X, then Y.
The original authority is still present. It has simply moved upstream, embedded in the rule. This produces a defining feature of automated governance: authority can keep operating long after it stops appearing in any individual decision. The rule executes. The user receives the consequence. The institution never has to introduce itself again.
Authority has not vanished. It has been converted into structure.
“The System Detected a Violation” Is Not the End of the Story
Adding AI or automation to the sentence does not solve the problem. “Our systems detected a violation” looks more transparent - at least something has been named as the actor. But a question is still buried inside it: what exactly did the system detect?
Compare “our system detected this phrase” with “our system detected a policy violation.” Those are not equivalent. A technical system can detect a phrase, an image, or a hash. It can calculate similarity, assign a probability, identify a pattern learned from training data. But violation is already an institutional category. Somebody had to define what counts as one before the system could treat a detected pattern as evidence of it.
Automated moderation should therefore not be confused with authorless moderation. Automation can change the proximate executor. It does not eliminate the institution that wrote the rules within which execution occurs. Automation is not authorlessness.
**Political Speech Exposes the Problem Faster
**The distinction becomes urgent around political speech: Palestine, Iran, sanctions, war, occupation, military intervention, terrorism, resistance, Zionism and anti-Zionism, anti-imperialism, armed organizations, civilian casualties.
These are not simple lexical categories. A journalist may quote a militant organization without supporting it. A researcher may reproduce extremist rhetoric in order to analyze it. A civilian may upload graphic footage to document an attack. A person may criticize sanctions without supporting the government targeted by them. Someone may oppose Israeli government policy without expressing hatred toward Jewish people - and someone else may use similar vocabulary as a vehicle for antisemitism. The words alone do not settle the question. Context does.
That creates a hard problem for moderation at scale. Platforms need categories that can be operationalized; political meaning is often contextual. Executable systems require formalization. The result is a permanent tension between political context and moderation executability: the more politically complex the speech, the harder it becomes to translate into a stable category without losing information.
**Palestine as a Documented Stress Test
**Palestine-related moderation offers one of the clearest documented environments for studying this. Human Rights Watch examined more than one thousand reported cases involving Palestine-related content on Instagram and Facebook during October and November 2023, documenting content removals, account restrictions, limits on engagement, visibility reductions, and obstacles to appeal (Human Rights Watch, 2023).
That report does not prove that every enforcement action was politically motivated censorship. It establishes something more useful for analysis: political speech was being processed through several different visibility-control mechanisms at once.
The distinction matters, because a platform does not have to delete a post to change its political reach. It can alter distribution, recommendation, eligibility, searchability, account functionality, monetization, and discovery. That is already far more complicated than the popular image of censorship.
**One Arabic Word Demonstrates the Mechanism
**Meta’s treatment of the Arabic term shaheed became a revealing case. In 2024, Meta’s Oversight Board concluded that the company’s previous approach to the term was overbroad and disproportionately restricted expression; Meta had reported that variations of the word were associated with more removals under its Community Standards than any other single word or phrase (Oversight Board, 2024a).
Why does this matter so much? Because it demonstrates that a word does not contain its moderation status naturally. The real process runs: word → context → reference → institutional interpretation → policy category → enforcement. If the middle of that sequence disappears, the user sees only “your content violated policy.”
The Oversight Board’s intervention showed that the relationship between the word and the violation was never self-evident. It was an institutional interpretation - and institutional interpretations can change.
**Appeals Expose the Hidden Decision
**There is a simpler way to see the same thing. You receive “your post violates our policy.” You appeal. A second reviewer examines the content. The platform reverses its decision. What changed? Usually not the post. The institutional judgment changed.
Which means “your post violates our policy” was never the complete proposition. The fuller version was “we determined that your post violates our policy.” Those two sentences sound similar. They are not institutionally equivalent: one describes violation as a property, the other identifies it as a judgment.
If a determination can be reversed, someone had to determine it in the first place.
**The Most Powerful Restriction May Not Delete Anything
**Deletion is obvious: a post existed, then it disappeared. But platforms govern more than existence. They govern visibility. A post can stay online and become harder to find. An account can stay active and lose recommendation eligibility. A video can remain accessible and vanish from the systems that would have surfaced it. Distribution shrinks, search visibility shifts, monetization ends. The object survives; its probability of being seen changes.
That shift produces a linguistic transformation worth attention. Compare “we reduced the distribution of your content” with “your content has reduced distribution.” In the first, an institution acts. In the second, the result looks like a condition belonging to the content. Or compare “we no longer recommend your account” with “your account is not eligible for recommendation.” One describes institutional action; the other describes account status.
The difference is crucial, because decisions become easiest to naturalize when they stop looking like decisions. The platform no longer appears to have done anything. The account simply is ineligible.
**How a Decision Becomes an Environment
**The same mechanism appears throughout institutional language. An organization makes a choice. The choice becomes a procedure. The procedure becomes routine. Eventually the consequence appears as a condition. Nobody says “we chose this outcome.” Instead: the system requires it, policy does not allow it, the account is ineligible, the content cannot be recommended.
This is one of the central consequences of the compiled rule (Startari, 2025a, 2025b). The rule stops looking like somebody’s decision and starts looking like a feature of reality.
**Meanwhile, the User Stays Fully Visible
**Look at the other side of the sentence. Moderation systems have no difficulty identifying the governed actor: you violated our rules; your account has repeated violations; your post contains prohibited material; your content is harmful. The user is visible. The account is visible. The post is visible. The violation, the alleged risk, and the sanction are all visible. What becomes less visible is the institution making the classification.
This is what I have described elsewhere as asymmetric visibility: different actors receive different amounts and types of grammatical agency within the same discourse (Startari, 2026c). Here the asymmetry is blunt. The user appears as violator. The content appears as risk. The policy appears as authority. The system appears as detector. And the platform can disappear as decision-maker.
That changes how accountability can be reconstructed after the fact.
**There Is a Way to Measure This
**None of this has to stay at the level of philosophical interpretation. The grammar is measurable. I propose two complementary measures.
Censor-Deletion Rate (CDR)
Take every clause that describes a restrictive platform action, then ask one question: is the institutional actor responsible for the restriction explicitly present? “We removed your post” - institution visible. “Your post was removed” - institution absent.
If a corpus contained one hundred suppression clauses and the platform disappeared from seventy of them, the CDR would be 70%. That number would not tell us that 70% of the cases were illegitimate. It would tell us something narrower and testable: 70% of the restrictions were represented without an explicit institutional agent.
**Suppression Opacity Index (SOI)
**Grammar alone is not enough, so the second measure is broader. The SOI asks whether a user can reconstruct the institutional decision at all: Who acted? What action occurred? Which rule was applied? How was the content classified? Was automation involved? Was human review involved? What sanction was imposed? Can the decision be appealed?
A sentence can be passive and still explain almost everything: “Your post was removed under Rule X after automated detection and human review. You may appeal here.” The actor is grammatically missing from the first clause, but the institutional chain stays visible.
Now compare: “Your content violates our standards and is not eligible for recommendation.” The platform is missing. The exact policy is missing. The classification process, the enforcement mechanism, and the appeal pathway are all missing. The content, the violation, and the consequence remain. That is much higher suppression opacity.
CDR measures grammatical disappearance. SOI measures accountability loss.
**This Does Not Prove Political Censorship
**That distinction has to stay strict. A high CDR does not prove political bias. A high SOI does not prove that a moderation decision was wrong. An opaque explanation can accompany legitimate enforcement, a transparent explanation can accompany illegitimate enforcement, and politically sensitive content should not be presumed legitimate simply because it is politically sensitive.
The empirical test requires comparison. If Palestine-related moderation notices have a high CDR but spam notices from the same platform have exactly the same CDR, the pattern may reflect nothing more than the platform’s house style. If political content remains significantly more opaque after controlling for platform, sanction type, language, and policy category, the result becomes far more interesting. That is the difference between accusation and measurement.
**From Content Moderation to Responsibility Moderation
**The controversial claim here is not that platforms moderate content - everyone knows that. Nor is it that AI makes moderation decisions, which is too simplistic. The stronger claim is this: platforms can convert institutional judgments into procedural facts.
A company defines a category. The category becomes operational. The rule becomes executable. A system classifies the content. A sanction follows. The user sees “your content violated policy,” and the institutional history has disappeared from the surface - not necessarily because anyone wanted to hide it, but because once authority becomes executable, the full authority structure no longer needs to appear every time the rule runs.
Content moderation determines what happens to speech. Moderation language determines something else: how visible responsibility for that decision remains. Call that second process responsibility moderation. A platform governs whether speech circulates; its explanation then governs whether the platform itself remains visible as the agent of that intervention.
That leaves two separate visibility problems. The first asks: can people see the content? The second asks: can people see who changed the conditions under which the content can be seen? Those questions are not equivalent, and modern platform governance requires both.
**The Censor Did Not Disappear
**The final paradox is the most important one. When the censor disappears from the sentence, institutional power has not become weaker. It may have become more deeply embedded. The old censor needed to perform an identifiable act; the compiled system needs only a condition. If X → Y. Once X is detected or classified, Y follows, because the authority was installed upstream.
So the terminal language can become almost completely impersonal: content was removed; distribution was reduced; the account is ineligible; policy was violated. Nothing in those sentences sounds dramatic. That is exactly why they deserve analysis.
The most sophisticated institutional authority may not be the one that constantly declares “I am making this decision.” It may be the one capable of transforming its decisions into ordinary properties of the system.
The user remains. The violation remains. The policy remains. The sanction remains. The rule executes. The institution becomes grammatically optional. The censor has not stopped governing - the censor no longer needs to remain in the sentence.
**Your Turn
**Go back to the last moderation notice you received, or the last one you screenshotted in frustration. Read it as grammar rather than as verdict: who is the subject of the sentence, and who has been removed from it? If you have an example where the platform disappeared entirely - or one where it explained itself well - post it in the comments. A corpus of those notices is exactly what a CDR and SOI study needs, and readers tend to have better archives than researchers do.
**Author’s Ethos
**I do not use artificial intelligence to write what I don’t know. I use it to challenge what I do. I write to reclaim the voice in an age of automated neutrality. My work is not outsourced. It is authored.
- Agustin V. Startari
**About the Author
**Agustin V. Startari is a linguistic theorist, author, and researcher in historical studies. His work examines how language, artificial intelligence, and formal systems redistribute authority, agency, and responsibility in contemporary institutions. He is the author of Grammars of Power, Executable Power, and The Grammar of Objectivity.
Researcher ID: K-5792–2016
**References
**Gillespie, T. (2018). Custodians of the Internet: Platforms, content moderation, and the hidden decisions that shape social media. Yale University Press.
Gorwa, R., Binns, R., & Katzenbach, C. (2020). Algorithmic content moderation: Technical and political challenges in the automation of platform governance. Big Data & Society, 7(1), 1–15. https://doi.org/10.1177/2053951719897945
Human Rights Watch. (2023). Meta’s broken promises: Systemic censorship of Palestine content on Instagram and Facebook.
Oversight Board. (2024a). Referring to designated dangerous individuals as “Shaheed”. Policy advisory opinion.
Oversight Board. (2024b). Posts that include “From the River to the Sea”.
Roberts, S. T. (2019). Behind the screen: Content moderation in the shadows of social media. Yale University Press.
Startari, A. V. (2025a). Compiled norms: Towards a formal typology of executable legal speech. SSRN Electronic Journal. https://doi.org/10.2139/ssrn.5353059
Startari, A. V. (2025b). Executable power: Syntax as infrastructure in predictive societies. Zenodo. https://doi.org/10.5281/zenodo.15754714
Startari, A. V. (2025c). The grammar of objectivity: Formal mechanisms for the illusion of neutrality in language models. SSRN Electronic Journal. https://doi.org/10.2139/ssrn.5319520
Startari, A. V. (2026a). Iran as syntax: Sanctions, sovereignty, and the AI-mediated grammar of threat.
Startari, A. V. (2026b). Suffering without perpetrators: The humanitarian passive in AI-generated conflict discourse. AI Power and Discourse, 1(1), 1–10.
Startari, A. V. (2026c). The grammar of asymmetric visibility: AI, Zionism, and the reallocation of political agency. AI Power and Discourse, 1(1), 1–10.
Suggested tags: Artificial Intelligence, Content Moderation, Censorship, Platform Governance, Algorithmic Governance, Linguistics, Social Media, AI Ethics, Political Speech, Palestine
Top comments (0)