AI-generated marketing content is everywhere. The problem is not quality. The problem is homogeneity. When five frontier models rewrite 2,250 human blog posts, they produce pages that share a structural fingerprint so consistent that a classifier can identify them at 97.0% macro-F1 without reading a single word.
SlopShape, a paper from Sitefire (YC W26), shows that AI-generated commercial web content can be detected from structural signals alone: how information is presented, in what order, with what evidence, and in what voice. This is not a word-level detector. It is a shape detector. And it works even when every AI post is reworded by its own model.
Why Structure Beats Words
Word-level detectors identify unedited AI text almost perfectly. But they break under rewording. A single pass through the same model that generated the text drops detection accuracy to near-random. SlopShape sidesteps this by ignoring words entirely.
The classifier uses 203 features, 176 of which are structural:
- Section ordering patterns: Does the post start with a definition, follow with benefits, then close with a call to action?
- Evidence presentation: Are claims supported by statistics, anecdotes, or assertions?
- Voice consistency: Does the tone shift between sections, or stay uniform?
- Semantic HTML usage: How are headings, lists, and emphasis tags deployed?
- DOM depth and nesting: What is the average depth of nested elements?
These features are extracted by an LLM (not disclosed in the paper, but likely GPT-4 class) and validated against human gold annotations. Human-human agreement (kappa 0.939) and human-model agreement (kappa 0.951) are both high enough to trust the feature extraction pipeline.
Training Corpus and Ground Truth
The dataset is 2,250 pre-ChatGPT human blog posts from 268 company domains, mirrored into 11,250 AI-generated versions by five frontier models (GPT-4, Claude, Gemini, and two undisclosed). Pre-ChatGPT is critical: it establishes a clean baseline where human authorship is unambiguous.
Ground truth is still ambiguous at the edges. A human post edited by an AI assistant could carry structural fingerprints from both. The paper does not address this directly, but the high F1 score suggests the training set is clean enough that boundary cases do not dominate.
The held-out test set evaluates on unseen companies, not unseen posts from the same companies. This tests generalization across domains, not memorization of specific writing styles.
Architecture: LLM as Feature Extractor
SlopShape does not train a neural network on raw HTML. It uses an LLM to extract structured features, then trains a gradient-boosted tree (likely XGBoost or LightGBM, not specified) on those features.
This two-stage design has three advantages:
- Interpretability: Each feature is a human-readable signal (e.g., "Does the post use a numbered list in the first section?").
- Robustness: Rewording changes words but not structure. The classifier sees the same feature vector.
- Versioning: When agent output evolves, you retrain the tree, not the feature extractor.
The LLM feature extractor is the bottleneck. Running it at crawl-time on every page would add unacceptable latency to a search indexing pipeline. The paper does not describe deployment, but the obvious solution is batch processing: crawl pages, queue them for feature extraction, then classify asynchronously.
Detection Performance and Attribution
The classifier achieves 97.0% macro-F1 on held-out companies. When every AI post is reworded by its own model, F1 drops to 96.1%. This is a 0.9-point degradation, compared to the near-total collapse word-level detectors experience under rewording.
Attribution is harder but still viable. The classifier assigns 68.6% of AI posts to the correct source model against a 16.7% chance rate (six classes: five models plus human). This suggests each model has a distinct structural signature, even when generating the same content.
Human posts occupy rare structural configurations. The paper does not quantify this, but the implication is clear: human writers vary their structure more than AI models do. AI models converge on a tidy, self-announcing shape.
Adversarial Dynamics and Model Versioning
This is an arms race. Once agents know the detection signals, they can train to evade them. The paper does not address this, but the dynamics are predictable:
- Phase 1: Agents optimize for word-level diversity (already happening).
- Phase 2: Agents optimize for structural diversity (randomize section ordering, vary evidence types).
- Phase 3: Detectors shift to meta-structural signals (e.g., "Does the structural variation itself look synthetic?").
The defense is versioning. You version the detection model as agent output evolves. You version the feature extractor as new structural patterns emerge. You version the training corpus as the boundary between human and AI authorship blurs.
This is not a one-time classification problem. It is an ongoing observability problem.
Deployment Shape and Latency Constraints
A production deployment would look like this:
| Component | Role | Latency Budget |
|---|---|---|
| Crawler | Fetch HTML | 500ms per page |
| Feature Extractor (LLM) | Extract 203 features | 2-5s per page |
| Classifier (Tree) | Predict AI/human + source | <10ms per page |
| Queue | Decouple crawl from classification | N/A |
| Cache | Store features for repeat classification | <1ms per page |
The LLM feature extractor is the expensive step. You cannot run it inline during crawl. You queue pages, extract features in batch, then classify. If you need real-time detection (e.g., for a search ranking signal), you cache features and reuse them until the page changes.
The tree classifier is fast enough to run inline. 10ms is acceptable for most ranking pipelines.
When Human-AI Boundaries Blur
The paper assumes a clean split: human or AI. But production content is hybrid. A human writes a draft, an AI assistant rewrites it, a human edits the AI output. What does the classifier see?
The paper does not test this, but the structural fingerprint likely reflects the final editor. If the AI assistant rewrites the entire post, the structure will look AI-generated. If the human edits lightly, the structure will stay human.
In practice, systems like Sitefire likely flag hybrid content as high-uncertainty and route it to human review or apply a confidence threshold before making classification decisions.
This is a problem for attribution. If you want to know "Was this written by GPT-4 or Claude?", you need to know whether the human edited the output. The classifier cannot tell you that.
Code and Artifacts
The paper releases pipeline, instrument, prompts, code, and aggregate artifacts. This is rare for a detection paper. Most release nothing, or release only the trained model. SlopShape releases the entire feature extraction pipeline, which means you can reproduce the results and adapt the instrument to your own domain.
The verification URL is not included in the arxiv abstract; consult the full paper PDF for the repository link.
Technical Verdict
Real-time detection is not viable with SlopShape's architecture. The LLM feature extractor requires 2-5 seconds per page, which makes inline classification during crawl impossible. The production pattern is batch processing with an async queue: crawl pages, queue them for feature extraction, classify later, and cache results.
Use SlopShape-style structural detection when:
- You can tolerate batch latency (minutes to hours between crawl and classification).
- You need robustness against adversarial rewording that breaks word-level detectors.
- You want interpretable features for debugging, auditing, and model versioning.
- Your domain is commercial blog posts, marketing pages, or similar structured content.
- You are building an observability pipeline, not a real-time ranking signal.
Avoid it when:
- You need real-time detection with sub-100ms latency (the LLM bottleneck makes this impossible without caching).
- Your content is hybrid human-AI and you need to attribute specific sections rather than classify the whole page.
- Your domain is far from commercial blogs (fiction, academic papers, social media) where the 203-feature instrument may not generalize.
- You are in an adversarial environment where agents are already training to evade structural detection (the arms race will require continuous retraining).
The real insight is not the 97% F1 score. The real insight is that AI models converge on a structural shape, and that shape is detectable even when the words change. This is a signal about agent output homogeneity, not just a detection technique. As agents become more capable, the question is whether they will learn to vary their structure, or whether structural convergence is an intrinsic property of optimization under a shared objective.
Top comments (0)