DEV Community

Costin Gheorghe
Costin Gheorghe

Posted on Originally published at letslaunch.today

How to Write a Paragraph an AI Engine Can Actually Quote

Originally published on LetsLaunch.

We wrote up the actual research
on what correlates with an AI engine citing a source, and the honest version
is shorter than most "GEO checklist" posts want it to be: one controlled
benchmark found that quotations and statistics got cited more than vague
content, and keyword stuffing measurably hurt. That's a real mechanism,
observed in a lab, and it's the most solid data point that exists.

This post is not about that research. It's about what to actually do with it
— the writing-craft side, for someone sitting down to write a paragraph today.
Not a proven-percentage technique, because we don't have one to sell you. Two
mechanical, checkable habits that follow honestly from what's known, plus one
you should adopt for a much older reason: it makes your writing better for
the person reading it, independent of whatever ChatGPT does with it.

Why structure matters even without a proven number

An extraction system — whether it's Google building an AI Overview,
Perplexity assembling an answer, or ChatGPT summarizing a page it retrieved —
is doing something specific: pulling a passage out of your page and presenting
it with reduced or no surrounding context. It might show one paragraph. It
might show one sentence. Whatever it shows has to stand on its own, because
the rest of your page usually doesn't travel with it.

That's true regardless of what the exact citation-rate lift is, which nobody
outside a handful of labs has measured on production systems. What you can
reason about, without needing a study, is whether a given passage survives
being extracted. Either it still makes a complete, correct claim once it's
alone on the screen, or it doesn't. That's a property of the text, checkable
by reading it, not a promise about traffic.

The test: copy it into a blank document

Here's the mechanical version, and you can run it on anything you've already
published.

Take one paragraph. Copy it alone into a blank document — nothing before it,
nothing after it. Read it as if you'd never seen the rest of the page. Ask:
does this still make a complete, correct, unambiguous claim?

If yes, the paragraph is structurally extractable. If it depends on "as
mentioned above," a pronoun with no visible referent, or a claim the previous
paragraph set up but this one doesn't restate, it isn't — and an extraction
system either garbles it, skips it, or quotes something that reads as
confused or wrong when separated from its context. None of that requires
knowing anything about how any particular model works. It's the same test a
human editor would run on a pull-quote.

A before-and-after

Here's a paragraph that fails the test, from a page describing a fictional
product:

It handles this automatically, which saves teams a lot of time compared to
the old way of doing it. That's part of why customers tend to stick around
longer once they switch.

Read alone, this says almost nothing. What does "it" do? What's "the old
way"? What's "a lot of time"? "Tend to stick around longer" than what, by how
much? Every noun refers outward to a sentence that isn't there anymore. A
person skimming a summary would come away with no actual information, and an
extraction system quoting it verbatim would be quoting a sentence that reads
as vague at best and misleading at worst — it implies a specific comparison
without stating one.

Here's the same claim, rewritten to stand alone:

Acme's scheduling tool re-assigns a missed shift automatically within two
minutes, instead of a manager finding a replacement by phone. Teams that
switched from manual scheduling report needing about 30% less manager time
per week on shift coverage.

Now it names the product, states the mechanism, gives a concrete number, and
makes a specific comparison. Someone could read only this paragraph, with
nothing before or after it, and walk away with a correct and complete
understanding of the claim. That's the whole test — not length, not keyword
density, just whether the sentence still works with everything around it
removed.

Notice this rewrite also happens to be the kind of claim GEO-BENCH found
correlated with more citations in its benchmark: specific, quotable, backed
by a number rather than an adjective. That's not a coincidence — a claim that
stands alone and a claim that's specific enough to quote tend to be the same
sentence. But we're stating that as a plausible connection between two
documented properties, not as a guarantee that this exact rewrite will get
you cited anywhere.

Put the answer before the build-up

The second habit: when a section implicitly answers a question — "how much
does this cost," "does it support SSO," "how long does setup take" — state
the answer in the first sentence of the section, then explain or qualify it
afterward. Not the reverse, where the section opens with context and works
toward a reveal at the end.

This isn't from a measured percentage. It's a documented, widely-recommended
practice in web writing generally, for a plain reason: both a human skimming
your page and a system scanning for a relevant passage are doing the same
thing — looking for the sentence that answers the question, not reading in
order for a payoff. A reader who wants to know if you support SSO and finds
that answer buried after three paragraphs of company background has a worse
experience whether or not any AI ever touches the page. Front-loading the
answer serves the actual reader first. That it also gives an extraction
system a clean, early sentence to lift is a second, plausible benefit — not
the reason to do it.

Compare:

Our platform was built from the ground up with enterprise needs in mind,
and over the years we've worked closely with security teams at companies of
every size to refine an authentication experience that meets a wide range
of compliance requirements. SSO is supported.

against:

Yes, this supports SSO, via SAML and OIDC. It's configured in Settings →
Security, and works with Okta, Azure AD and Google Workspace out of the
box.

The second version answers the implied question in its first four words. Everything
after that is detail a reader can keep going for or stop reading, having
already gotten what they came for.

What this doesn't promise

Neither habit above comes with a measured lift on real AI engines, and we're
not going to invent one. You may come across posts claiming a specific
percentage improvement from "answer-first" writing, sometimes down to one
decimal place, or a claim that answers need to be a precise word count to
count as citable. We looked. We couldn't trace either kind of number back to
a named study, a named researcher, or any data anyone could check — just
marketing content citing other marketing content. A blog built on "verified,
not estimated" doesn't get to make an exception for a number just because
it's specific-sounding. So we're leaving both out, and we'd suggest treating
any post that states one with unusual precision and no source the same way.

What we will say: self-contained paragraphs and answer-first sections are
good writing by any standard that predates language models entirely. They
make content easier to skim, easier to excerpt in an email, easier to quote
in a meeting, easier for a tired reader at 11pm to get what they need from
without reading the whole page. If a generative engine also finds that shape
easier to extract cleanly — which the one real study we trust suggests is at
least plausible, since specific and self-contained claims are close cousins —
that's a reasonable bonus on top of writing that was already worth doing. It
is not, and we're not going to pretend it is, a guaranteed technique with a
number attached.

None of this works if the system pulling the passage can't reach your page in
the first place — see why ChatGPT can't see your
site
if you haven't confirmed yours is
readable by a crawler before worrying about how the sentences are shaped.

And if you're publishing a product description anywhere public, including
a LetsLaunch listing, the same test applies to that paragraph too
— read it alone, with no page around it, and see if it still holds up.

Does writing self-contained paragraphs actually increase AI citations?

Nobody has published a study measuring that specific effect on production AI
engines, so we can't claim it does. What's documented is a related, narrower
finding: content with specific quotes and statistics got cited more than
vague content inside a Princeton benchmark. Self-contained, specific writing
is a reasonable, checkable habit that's consistent with that finding and
also makes content easier for a human to skim — not a proven technique with
a measured percentage behind it.

What's the quickest way to check if a paragraph is AI-citable?

Copy it alone into a blank document, with nothing before or after it, and
read it as a stranger would. If it still states a complete, correct,
unambiguous claim with no missing context, it's structurally extractable. If
it depends on a pronoun, "as mentioned above," or a fact the previous
sentence supplied, it isn't — rewrite it to carry its own context.

Should I put the answer first or build up to it?

Put it first, then explain. This isn't based on a measured citation
percentage — it's a documented practice grounded in how both readers and
extraction systems scan text: looking for the sentence that answers the
question, not reading in narrative order. A reader who wants to know if you
support SSO benefits from the answer in the first sentence whether or not any
AI system ever touches the page.

Top comments (0)