DEV Community

Faktoskop.pl
Faktoskop.pl

Posted on

What 40 analysed news articles actually look like in aggregate


I spent some time reading news articles sentence by sentence, classifying each one as a verifiable claim, an opinion, a persuasive construction, or neutral connective text. The sample is 40 articles drawn from a larger set of 99, all Polish-language online media, selected by what people submitted rather than by any sampling design.

That last point matters and I will come back to it. First, the numbers.

The distribution

Across the sample:

  • ~42% verifiable factual claims- ~32% opinion presented within a factual frame- ~5% overtly persuasive constructions- remainder neutral or structural

Mean overall assessment: 5.36 / 10.

The 5% is the surprising number

Going in, I expected the persuasive fraction to be much larger. The common intuition — mine included — is that news is saturated with manipulation. Counting does not support that. Most sentences in most articles are ordinary reporting or ordinary connective tissue.

Five percent is small. It is also not evenly distributed, which turns out to be the whole story.

Persuasive sentences cluster at the openings and closings. The lede and the final paragraph carry a disproportionate share, and those are the positions that determine how a reader remembers the piece. A reader who skims the first and last paragraphs — which is most readers — encounters a substantially higher persuasive density than the article average.

Which suggests that "how much manipulation is in this text" is the wrong question, and "where is it" is the right one. A 5% average with heavy positional clustering is a different object from a 5% average spread evenly, and the summary statistic cannot distinguish them.

The 32% is the load-bearing number

Opinion-presented-as-fact is nearly a third of the sample, and it is the category that causes the most trouble.

These are sentences that are grammatically indistinguishable from factual claims but are not checkable: "the reform will strengthen the economy", "the decision came too late", "the situation is deteriorating". No hedging, no attribution, no marker that a judgement is being offered.

A reader processing quickly has no signal that the epistemic status changed. The sentence before was checkable, the sentence after is checkable, and this one arrives in the same voice.

This is also the category that survives fact-checking untouched, because there is nothing to check. You cannot verify "too late". The claim has the form of information and the content of a position.

What lowered assessments most often

Reading the individual analyses, the recurring issues were not false statements. In the whole sample I found very few claims that were straightforwardly wrong.

What appeared repeatedly:

Numbers without provenance. A figure is stated, no source, no methodology, no date. "30,000 visitors", "40% increase". The number may well be correct; there is no way to establish that from the text.

Percentages without baselines. A 40% rise from what, over what period, against what expectation. This one is mechanically detectable — every percentage either has a denominator in the text or it does not.

Named but unquoted sources. An organisation or official is mentioned as involved, and never actually quoted. The name provides an appearance of attribution without the substance of it.

Disagreement appended rather than integrated. Opposing views present, but confined to a final paragraph as "critics say", after the argument has been made without them.

The failure mode that survives everything

The pattern I found most interesting: the articles that scored lowest contained no false statements.

The lowest-scoring piece in the sample was a promotional article about a trade fair. Twenty sentences: nine facts, three opinions, six persuasive, two neutral. Every factual claim in it was, as far as I could determine, accurate. It scored 3.5.

What produced that score was selection and framing — a promotional text wearing the format of news reporting. A reader who verifies every claim finds nothing wrong, because verification was never the relevant operation.

This is the category I think is most underserved by existing tooling. "Is it true" is answerable and mostly beside the point. Selection, omission and framing do the work, and none of them are falsifiable statements.

What this sample cannot tell you

The sampling is not random. These are articles people chose to submit, which biases toward pieces someone found suspect. The population figures for Polish media as a whole are almost certainly different — probably with a lower persuasive fraction, since submissions skew toward the contested.

Forty articles is also small. The 5% and 42% figures should be read as rough shape, not as measurements. I would not defend the second significant figure on any of them.

And there is a methodological limitation worth stating plainly: sentence-level classification depends on how much text is provided. Analyse an excerpt and you get an assessment of the excerpt. A ratio that looks alarming in three paragraphs can be an artefact of where the cut fell. Comparisons between analyses are only meaningful when the inputs are comparable.

The individual analyses are published at faktoskop.pl for anyone who wants to check the distribution against their own reading, including the ones where the method performed poorly.

Question for the community

Has anyone seen comparable sentence-level counts for English-language media? I can find plenty of work on claim verification and very little on the composition of ordinary articles — what fraction is checkable at all, before you get to whether it checks out.

The 5% figure in particular I would like to compare against. If it holds across languages and outlets, the implication for tooling is that we are mostly building detectors for a small and positionally predictable subset of the problem.

Top comments (1)

Collapse
 
topstar_ai profile image
Luis Cruz

Your analysis of the persuasive clustering in news articles is particularly insightful, especially the distinction between how information is presented versus its verifiability. It's fascinating to see how readers' perceptions can be shaped by the structure rather than just the content itself. To improve clarity, incorporating more explicit markers or context for opinion statements could guide readers better in understanding the intent behind the claims. If you're looking for help with data visualization or enhancing the analysis methodology for a deeper dive into these distinctions, I’d be glad to explore a paid collaboration. What methods do you think could be effective for highlighting these patterns to readers?