DEV Community

Cover image for Only 2 of 21 AI Providers Publish the Document the EU AI Act Asks For
Pennyforge
Pennyforge

Posted on

Only 2 of 21 AI Providers Publish the Document the EU AI Act Asks For

The EU AI Act has been quietly asking AI model providers to do something specific since 2 August 2025: publish a "sufficiently detailed public summary of the content used for the training of the model", according to a template the Commission published. One year in, we went and looked at 21 of the largest GPAI providers. Two publish the actual document. Everyone else publishes something else that is close.

This is the first issue of a small, dated, verifiable conformance census. The table below is the deliverable; the method and caveats are below it, because we think both matter as much.

The rule, in three sentences

Article 53(1)(d) of Regulation (EU) 2024/1689 (the AI Act) obliges all providers of general-purpose AI models — including open-source and free ones, for this specific obligation — to make a public summary of training content available, "according to a template provided by the AI Office". The Commission published that template on 5 December 2025 (C(2025) 8311 final): an Explanatory Notice plus a fill-in form with three sections — 1. General information (provider, EU authorised representative, versioned model name, EU market-placement date, modality and per-modality data-size brackets, types of content), 2. List of data sources, 3. Data processing aspects. The Commission calls it "a common minimal baseline" — more detail is fine, less is the problem.

The timeline: the obligation applies as of 2 August 2025. Models placed on the EU market before that date have until 2 August 2027 to comply, and the AI Office's supervision and enforcement of GPAI rules starts 2 August 2026 — two months ago.

What we found: 2 of 21

We checked 21 major GPAI providers on 2026-10-08, reading the providers' own documents (not blogs about them). Two publish a dedicated, template-shaped "Public Summary of Training Content":

DeepSeek. Per-model summaries at cdn.deepseek.com/policies/ — e.g. the V3.2 summary (v1, updated 2026-09-03, 5 pages) is essentially the template filled in: EU authorised representative, versioned model ID, EU market-placement date (2025-08-21), per-modality size brackets, a source list (Common Crawl, Stack Exchange, licensed data, no user data, synthetic data), and a text-and-data-mining reservation.

Mistral. A legal center covering 42 models with per-model lifecycle and technical documentation, plus a dedicated "Public Summary of Training Content" for Large 3.

The remaining 19 publish model cards, technical reports, or system cards that describe training data in a paragraph, a table, or a citation to another document — detailed in some cases (Qwen3's 36-token composition, Aleph Alpha's 20T breakdown, Kimi K2's report), but none of them is a dedicated summary structured against the template. The table:

Provider Model(s) checked (10-08) Where training content is described Tier
DeepSeek V3.2, V4 Dedicated per-model "Public Summary of Training Content" (Policies) T3
Mistral Large 3 (+42 models in legal center) Dedicated Large 3 summary + per-model legal docs T3
Anthropic Claude Sonnet 4.5 149-page system card, §1.1.1 (two paragraphs: data mix, robots.txt, opt-in users) T2
OpenAI GPT-6.1 Sol System card addendum (2026-09-29), §2: one sentence — "same types of data and training as GPT-6 Astra" T2
Google Gemini 3 Pro Model card PDF (one paragraph on training data) T2
Meta Llama 4 MODEL_CARD.md (two lines + ~40T / ~22T figures) T2
xAI (SpaceXAI) Grok 4.6 Model card PDF (2026-08-12), one paragraph T2
Cohere Command A+ Docs page (three generic lines) T2
Amazon Nova 2 Lite Responsible-AI card (one paragraph) T2
Microsoft Phi-4 Model card (4 source categories, 9.8T); MAI series has no card T2+
IBM Granite 4.0 Card (SFT sources) + tech report (pretrain mixture) T2+
Alibaba (Qwen) Qwen3 arXiv 2505.09388 (36T + composition) T2+
Baidu ERNIE 5.0 arXiv 2602.04705 T2+
NVIDIA Nemotron 3 Ultra Technical report PDF T2+
Moonshot (Kimi) K2 Tech report (moonshotai.github.io, 15.5T) T2+
Cerebras Cerebras-GPT Paper + model card T2+
Aleph Alpha Kolibri-1 Tech report (20T, language split; CoP signatory) — the strongest model card we saw T2+
Stability AI SD 3.5 Card inside gated Hugging Face repo T2 (gated)
ByteDance Doubao Seed research pages only (API-first) T1–2
Zhipu GLM-4.6 Card + GLM-4.5 report T1–2
Tencent Hunyuan Unverified (blog-level only) ?

T3 = dedicated, template-conformant summary · T2+ = detailed card/report, no dedicated summary · T2 = brief section · T1–2 = thin/indirect · ? = not verifiable from public sources as of 10-08.

The US majors, checked against the primary text

A watchdog post from December 2025 (Open Future, "5+3+3=0 transparency") quoted the US majors' model-card text and found it "does not remotely resemble" the template. We re-checked against primary documents on 10-08:

  • OpenAI's current flagship card (GPT-6.1 Sol system card, dated 2026-09-29) devotes its entire "Model Data and Training" section to one sentence pointing at the GPT-6 Astra card.
  • Google's Gemini 3 Pro model card gives one paragraph.
  • Meta's Llama 4 model card gives two lines plus token counts.
  • Anthropic is the interesting counterexample in the US camp: its 149-page Claude Sonnet 4.5 system card is the most thorough safety document in the industry, and its training-data section is still two paragraphs — no authorised-representative line, no placement date, no modality brackets. Big safety effort, thin template conformance.

That is the pattern: training-data description is a sidebar in model cards everywhere except at the two providers that treated the template as the deliverable.

Method and caveats

  • Sample: the 21 largest GPAI providers by our judgment as of 10-08 (the original 20, plus Anthropic, which an early read flagged as missing — the sample is non-random and we say so).
  • Date: everything verified by fetching the provider's own document on 2026-10-08; the table is a snapshot, not a live index.
  • Grading: T3 requires a dedicated summary whose structure matches the template's three sections. We grade against the published template (C(2025) 8311 final), which the Commission frames as a minimal baseline — a provider could plausibly argue a richer card "covers" the fields; we would too, which is why T2+ exists as its own tier rather than being called non-conformant.
  • Known blind spots: Chinese-language documents (an English-only sweep is the classic census failure mode), gated repositories, and Tencent (flagged unverified, not graded). Code of Practice signatory status differs across the sample and likely predicts behaviour — worth tracking in a later issue.
  • Not legal advice. "Conformance" here is a documentary check against the published template, not a Commission assessment.

Why we're publishing a 21-row table

Because it changes. Mistral's summary was a section inside technical docs in December 2025 and a dedicated document by now. OpenAI's flagship card moved from GPT-5 to GPT-6.1 in the same period. A one-off article rots; a dated, versioned census with per-provider evidence links doesn't. That is the whole product hypothesis behind this: the value is not "who complies" (we can re-run it) but the diff — who published, who updated, who moved, against which template version, when.

Issue #2 will add machine-readable per-provider rows (version, date, URL, template section coverage) and a first diff check against this table. Corrections welcome — point us at a document we missed and we'll re-grade the row in the next issue.

Sources (all accessed 2026-10-08)

  • Template + Explanatory Notice (C(2025) 8311 final, 5.12.2025): digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models (page publication 24 Jul 2025, last update 26 Mar 2026); GPAI provider guidelines: /en/policies/guidelines-gpai-providers
  • DeepSeek: cdn.deepseek.com/policies/ (V3.2 summary v1, 2026-09-03) · Mistral: legal.mistral.ai/ai-governance/models
  • Anthropic: anthropic.com/claude-sonnet-4-5-system-card (Sept 2025, rev. 2025-12-03, 149 pp, §1.1.1)
  • OpenAI: cdn.openai.com/pdf/38e3efcf-…/oai_GPT_6_1_Sol.pdf (2026-09-29, §2)
  • Google: storage.googleapis.com/deepmind-media/Model-Cards/Gemini-3-Pro-Model-Card.pdf
  • Meta: raw.githubusercontent.com/meta-llama/llama-models/main/models/llama4/MODEL_CARD.md
  • xAI: media.x.ai/v1/website/card-4p6-4cd2dc57.pdf (2026-08-12) · Cohere: docs.cohere.com/docs/command-a-plus · Amazon: docs.aws.amazon.com/ai/responsible-ai/nova-2-lite/overview.html
  • Microsoft: huggingface.co/microsoft/phi-4 · IBM: ibm.com/granite + huggingface.co/ibm-granite/granite-4.0-tiny-preview · Qwen: arxiv.org/abs/2505.09388 · ERNIE: arxiv.org/abs/2602.04705 · NVIDIA: research.nvidia.com Nemotron-3-Ultra-Technical-Report.pdf · Kimi: moonshotai.github.io/Kimi-K2/ · Cerebras: huggingface.co/cerebras/Cerebras-GPT-256M · Aleph Alpha: aleph-alpha.com/downloads/tech-report.pdf
  • Stability: huggingface.co/stabilityai/stable-diffusion-3.5-large (gated) · ByteDance: seed.bytedance.com · Zhipu: zhipuai.cn/en/glm46v
  • Open Future, "5 + 3 + 3 = 0 transparency" (Z. Warso & P. Keller, 2025-12-09): openfuture.eu/blog/5-3-3-0-transparency/

Pennyforge (one-person studio) · full evidence links on request · AI-assisted research, studio-owned

Top comments (0)