DEV Community

Cover image for AI Writing Tells: New Quirks in Frontier Models
LuckyTaorem
LuckyTaorem

Posted on Originally published at ltdeveloperblogs.github.io

AI Writing Tells: New Quirks in Frontier Models

Overview of the Graphite Study

In a meticulously designed experiment, the marketing analytics firm Graphite set out to map the linguistic fingerprints—colloquially called tells—that differentiate text generated by today’s most advanced “frontier” AI models from human‑written prose. The research, led by Chief AI Officer Greg Druck, examined four major model families: Anthropic’s Claude Opus 5.5, OpenAI’s Astra, Google’s Gemini 3.1 Pro, and OpenAI’s GPT‑6‑derived Sol and Luna.

The methodology was rigorous:

  • Control corpus: 10,000 pre‑ChatGPT articles, representing a broad spectrum of journalistic styles.
  • Test process: Each AI model rewrote a summary of a control article, eliminating source bias while preserving core facts.
  • Scale of detection: Over 13,000 distinct phrases were flagged as appearing at least twice as often in AI‑generated text versus human text.

The findings reveal a paradox: while labs have successfully suppressed well‑known tells such as em‑dashes, every new model version spawns its own set of subtler quirks. The total count of detectable tells remains roughly constant, suggesting an equilibrium between mitigation efforts and emergent linguistic patterns.

“They are managing to remove the most well‑known tells, but other ones pop up. And every model version has its own.” – Greg Druck

Technical Breakdown of Model‑Specific Tells

Claude Opus 5.5 (Anthropic)

  • Primary word tell: dependable – appears 23× more frequently than in human samples.
  • Phrasal tells:
    • “this matters” – 116× more common.
    • “why X matters” – 92× more common.
  • Sentence template: “is more than an X, it’s a Y.”
  • Improvements: The older “it’s not X, it’s Y” construction has been largely retired, and em‑dash usage dropped by 99% compared with Opus 5.

These patterns hint at a stylistic bias toward certainty and explanatory framing, likely a by‑product of Anthropic’s alignment objectives that prioritize clarity and user reassurance.

Astra (OpenAI)

  • Phrasal tell: “another dimension.”
  • Hedging language: Frequent use of “may provide” or “can provide” when discussing benefits.
  • Core construction (“Corrective Framing”):
    • Starts with “not simply X” or “rather than relying on X.”
    • This corrective framing appears >100× more often than in human writing.
  • Em‑dash reduction: 88% lower than human baseline.

Astra’s proclivity for corrective framing reflects a defensive alignment strategy—preemptively addressing potential misconceptions before they arise.

Gemini 3.1 Pro (Google)

  • Key improvement: Near‑complete elimination of the em‑dash. This suggests Google’s fine‑tuning pipeline now includes explicit token‑level filters for punctuation patterns that were previously flagged as AI‑specific.

Sol and Luna (OpenAI GPT‑6 variants)

  • Marketing claims: “More clarity, less jargon, fewer odd turns of phrase.” While the study did not quantify specific tells for these models, the qualitative assessment aligns with the observed trend of reducing overtly mechanical phrasing.

Why Detection Matters: Safety, Trust, and Competitive Edge

1. Content Authenticity and Misinformation

Detecting AI‑generated text is a frontline defense against misinformation campaigns. If a malicious actor can flood social media with convincingly human‑like narratives, the ability to spot subtle tells becomes a critical tool for platforms, regulators, and fact‑checkers.

2. Intellectual Property and Attribution

Publishers increasingly need to verify whether a piece of copy was produced in‑house or outsourced to an AI service. Tells provide a forensic trail that can protect copyright and ensure proper attribution.

3. Model Governance and Compliance

Regulators are beginning to draft requirements for AI transparency. Demonstrating that a model’s output can be audited for tells satisfies emerging “explainability” mandates, especially in high‑stakes domains such as finance or healthcare.

4. Competitive Differentiation

Companies that can claim “fewer tells” gain a market advantage. For instance, OpenAI’s reduction of em‑dash usage is a tangible metric that can be advertised to enterprise customers concerned about brand voice consistency.

Industry Impact and Competitive Landscape

The Graphite study underscores a shifting arms race:

  • Anthropic appears to be moving toward a more human‑like distribution of words, as evidenced by the decreasing gap in word frequency metrics.
  • OpenAI’s GPT‑6 line shows a divergent trajectory, with tells becoming more pronounced—perhaps a side effect of scaling parameters without proportionate alignment data.
  • Google demonstrates a focused engineering effort on punctuation patterns, suggesting a narrower but effective mitigation strategy.

These dynamics influence partnership decisions, venture capital allocations, and even talent recruitment. Companies that can reliably produce “tell‑free” text may secure contracts for content generation, legal drafting, or customer support automation.

For a broader view of how tech giants manage risk, see the recent coverage of OpenAI’s internal safety challenges: https://ltdeveloperblogs

Implications for Detection Tools

The granular list of tells uncovered by Graphite gives detection engineers a fresh set of features to incorporate into classifiers. Traditional AI‑detector models have relied heavily on statistical anomalies such as perplexity spikes or token‑frequency mismatches. The new findings suggest a multi‑layered approach:

🔹 -----------------
• Example Tell: --------------
• How to Leverage: -----------------

🔹 *Lexical*
• Example Tell: Over‑use of dependable (Claude)
• How to Leverage: Flag unusually high term frequency relative to a baseline corpus.

🔹 *Phrasal*
• Example Tell: “this matters”, “why X matters” (Claude)
• How to Leverage: Deploy n‑gram matching with a dynamic threshold that adapts to article length.

🔹 *Syntactic*
• Example Tell: “is more than an X, it’s a Y” template (Claude)
• How to Leverage: Parse sentence structures and compare against a library of AI‑specific templates.

Read the full breakdown originally published at https://ltdeveloperblogs.github.io/posts/opus-55-loves-to-tell-you-this-matters-and-other-ai-writing-tells/

Top comments (0)