DEV Community

Cover image for The EU AI Act training-data census, corrected: 13 of 21, not 2 of 21
Pennyforge
Pennyforge

Posted on

The EU AI Act training-data census, corrected: 13 of 21, not 2 of 21

Nine days after we counted "only 2 of 21" AI providers publishing the document the EU AI Act asks for, we went back and re-verified every row in the census. The corrected number: 13 of 21 providers now publish a dedicated, template-conformant "Public Summary of Training Content" under Article 53(1)(d).

That's a big correction. Before we celebrate the providers or beat up on ourselves, it's worth being precise about what actually happened — because part of the delta is real drift, and part is us having looked too shallowly.

The document, briefly

The EU's template (C(2025) 8311 final) asks GPAI providers to publish a public summary with exactly three sections: (1) general information (provider, model IDs, EU market-placement date, per-modality data size brackets), (2) a list of data sources (public, licensed, crawled, user, synthetic), and (3) data processing aspects (TDM opt-outs, filtering, etc.). The obligation has applied since 2025-08-02; pre-existing models have until 2027-08-02.

In the issue #1 sweep we graded what a first-time visitor sees: the model card, the research page, the landing page. That undercounted. Most of the 10 providers we then promoted keep their Art. 53(1)(d) documents one click deeper — on a legal page, a trust center, or a transparency-report page that the card never links to.

The 10 promotions (all verified 2026-10-09)

Ten providers moved up to T3 (dedicated, template-structured summary):

  • OpenAI — per-model "Public Summary of Training Content" PDFs (GPT-5.2 checked live: v1, 2026-07-30, OpenAI Ireland, placed 2025-12-11, >10T tokens; the document says it covers GPT-5.2 and all subsequent releases in the lifecycle). Six or more per-model PDFs, indexed at help.openai.com.
  • Google — an 8-page "Public Summary of Training Content for General-Purpose AI Models — Gemini 3 model family" (v1, 2026-07-02, Google Ireland, placed Nov 2025, all four modalities), hosted on the AI Transparency Report's GCS bucket. A Gemma 4 report (2026-07-31) is listed too.
  • Anthropic — a "Training Data Summaries" section on the Trust Center with ten per-model documents (Opus 4.7→5.5, Sonnet 5/5.5, Mythos/Fable lines, plus an AB 2013 summary). The Opus 5.5 document (v1, 2026-09-22) is template-conformant and names low-resource languages explicitly (Basque, Breton, Korean).
  • Meta — an explicit Art. 53(1)(d) reporting program: "EU AI Act Transparency Reports published by Meta Platforms Ireland Limited pursuant to Article 53(1)(d) of Regulation (EU) 2024/1689", with per-model "Transparency Report" documents (Muse Glimmer v1, 2026-08-10 checked). Honest caveat: it's a thin 2-pager (a third-party tracker grades it C/D+).
  • Microsoft — per-model "Data Summary" PDFs for the MAI series (MAI-Thinking-1 checked live: v1.0, 2026-08-12, Microsoft Ireland, EU rep contact listed, placed 2026-08-12).
  • ByteDance — dedicated "Public Summary of Training Data Content" PDFs per model from seed.bytedance.com/transparency (Seed 2.0 Pro checked live: v1.0, 2026-07-31, EU authorised rep Mikros Ireland, placed March 2026; Seedream 5.0 Pro and Seedance 2.0/2.5 listed).
  • xAI — "Public Summary of Training Content for Grok 4.5" (V1, 2026-07-08, xAI LLC, EU rep in Tallinn, placed 2026-07-14); Grok 4.6 (2026-08-12) also listed on x.ai/legal.
  • Tencent — "Public Summary of Training Content for Hy3" (V2, 2026-08-20, OriGen Tech Singapore, EU rep Tencent International Service Europe B.V., Amsterdam, placed 2026-07-06). This closes the census's original CN-language blind spot.
  • Cohere — "Public Summary of Training Content for Command A+" (v1, 2026-07-31, Cohere Germany GmbH, placed 2026-05-20, 48 languages, data to April 2026).
  • Aleph Alpha — "Kolibri 1 — Sufficiently Detailed Summary" (v1.0, 2026-10-03): 8 pages, all three sections substantive, 8 named public datasets and 7 named synthetic-generator models.

Add the three that were already at T3 before the correction (DeepSeek, Mistral, NVIDIA) and you get 13 of 21.

The DeepSeek mystery, resolved

Last week we couldn't resolve DeepSeek's summary file: their CDN is an S3 bucket whose index 404s and whose filenames appear to rotate. This week the answer arrived from a third-party tracker's link list: the exact file DeepSeek_Template_for_the_Public_Summary_of_Training_Content_for_GeneralPurpose_AI_model.pdf is live on cdn.deepseek.com. It covers DeepSeek-V4 (incl. V4-Pro and V4-Flash), v1, last updated 2026-08-04, placed on the EU market 2026-04-24 — and it names an EU authorised representative: Prighter Group. Lesson: the bucket isn't unstable for the current filename, just for the old ones.

The remaining 8

  • T2+ (detailed reports/cards, no template summary): Alibaba Qwen (Qwen3 tech report, 36T tokens; the Qwen3.8 flagship cards are the place to watch), Baidu ERNIE 5.0 (tech report + release post, 2.4T unified multimodal), IBM Granite 4.0, Moonshot Kimi K2 (15.5T tech report), Cerebras-GPT.
  • T2: Amazon (Nova 2 Lite "AI Service Card" — one training-data paragraph).
  • T2-gated: Stability AI (card inside a login-gated Hugging Face repo).
  • T1-2: Zhipu GLM (research page only).

The gap moved. It's now about hosting, not existence.

The most interesting finding of the re-verification is that the compliance gap has changed shape. In October 2025 the question was "does anyone publish?"; now the question is "can you still find it in 90 days":

  • Cohere hosts its summary behind a signed S3 URL that expires after 7 days. Our first-party link from August is already dead; the document survives via the stable docs page that links to it (and via third-party archives).
  • Tencent's CDN returned an obscure HTTP 567 to EU fetchers all day; the document only resolved through a dated archive copy.
  • DeepSeek's old filenames 404 even as the current one works — no redirects, no index.
  • Meta's per-model report is 2 pages.

None of that disqualifies the document today. All of it will matter in 2027, when the pre-existing-models deadline hits and rightsholders start auditing.

The closest incumbent (and why we cross-reference it)

The AI Accountability Lab (aial.ie) runs "GPAI Training Transparency": 78 discovered summaries with live URLs, dated archives and A+–F grades. It's the closest thing to an incumbent of what we're building — and its dated archive copies served as our verification substitutes wherever first-party hosting was flaky (Tencent, Cohere, Anthropic, Meta). We treat it as a discovery cross-check, not ground truth: every census row now records in url_source exactly which route verified it (fetched-direct, or live-anchor-fetched + dated-archive-checked).

The machine-readable part

The census is now a small public repo with a diff tool: github.com/m0nk111-qwen-agent/gpai-transparency-census. The first real 24-hour diff (2026-10-08 → 2026-10-09) shows 17 of 21 rows changing, including the 10 promotions. That diff — not the absolute count — is the number we'll actually track going forward: how many providers move per day.

Honest caveats

  • The "2 of 21" in issue #1 was the card-level truth at publish time (what a first-time sweep sees), stated as such in the repo's README; this post is the correction, not a contradiction.
  • "Template-conformant" is a structural check (the three sections present and filled), not a quality grade — Meta's 2-pager passes structurally and is thin in practice. A quality dimension is the next iteration.
  • Chinese providers were swept with a degraded search lane plus direct domain checks; Qwen3.8, ERNIE 5.0 and MiniMax H3 are the best next-checks if anyone wants to sharpen that corner.

Pennyforge (one-person studio) · per-provider evidence as linked · AI-assisted research, studio-owned

Top comments (0)