DEV Community

Qtim
Qtim

Posted on

ChatGPT’s New EU Status Exposes an AI Architecture Problem

A product team ships a “search the web” toggle behind a chatbot. The interface barely changes. The system now decides when to browse, rewrites the user’s question, selects sources, and compresses them into one answer. That small toggle creates a much harder audit problem.

On Aug 31, 2026, the European Commission designated ChatGPT a Very Large Online Search Engine under the Digital Services Act. The Commission called it a hybrid service because it answers prompts and can search the web. The classification followed what the product does beneath the interface.

I’m Anton Fokin, CEO of Qtim. We build AI chatbots, RAG systems, and language-model integrations. The useful lesson for product teams is architectural: once a system selects information for users, the final answer stops being enough evidence.

ChatGPT crossed the 45-million-user threshold

ChatGPT reported 159.1 million average monthly users in the EU in the official list, well above the 45 million threshold for a VLOSE. After notification, the service has four months to comply with the additional obligations for the largest search engines. Those include systemic-risk assessment and mitigation, annual independent audits, and data access under defined procedures.

A 2026 Microsoft Research preprint gives the product question more weight. The authors examined 234,839 public ChatGPT conversations collected from 2023 through 2025. They classified 79% of user inputs as difficult to answer through conventional web search. For comparable searchable questions, ChatGPT responses covered less diverse information than Google results across most topics. The dataset is observational and does not represent every user. It still shows that a system can accept a broad question and return a narrower range of information.

The designation imposes obligations on the named services. It does not automatically place every AI product under the DSA. The regulation covers intermediary services offered to recipients in the EU, regardless of the provider’s location. Outside the DSA, public AI products still need traceability when they choose sources for users.

Five checks we use before release

Before an AI feature ships, we ask:

  • Which information space can the system reach: user input, company documents, partner data, or the open web?
  • Where does the product select sources and narrow the user’s view?
  • Can the team reconstruct the path from request to answer across model, instruction, retrieval, and policy versions?
  • Which users and harms could the feature affect, and what signal would reveal a problem?
  • Who can inspect the decision a month later, and which evidence will still exist?

The answers define the product boundary more accurately than the word “chatbot.” A closed knowledge base may only need document identifiers and versioning. An open-web product needs a decision log, an evaluation pipeline, and a tested way to stop the feature.

You cannot audit a screenshot

A generative-search answer is the end of at least four decisions: whether to search, which queries to send, which documents and passages to use, and which policies or filters to apply. Two identical answers can come from different sources. Two different answers can come from the same model version after web results or an index changes.

We use a common request identifier to join the evidence chain. It should connect the model and instruction versions, tool calls, generated queries, source or chunk identifiers, policy decisions, the final answer, citations, retries, human intervention, and rollback events.

The tempting shortcut is to keep every prompt, retrieved page, and response forever. That improves reproduction and creates a much larger privacy and security problem. Retention, masking, deletion, and access controls belong in the traceability design. A useful record reconstructs a decision without cloning the conversation database.

Testing also moves one layer down. A fixed “golden answer” breaks as wording changes. More stable checks ask whether the right sources were eligible, the expected tool ran, a policy fired, or a high-risk condition stopped the workflow.

Put risk review in the release path

Article 34 of the DSA requires systemic-risk assessments at least annually and before features likely to have a critical impact on identified risks. Product teams can translate that into a release gate.

Four objects need to stay connected:

the user scenario and the groups it affects;
the risk and the observable signal that would reveal it;
the mitigation, such as a limit, review step, interface change, or policy;
the rollout plan, including monitoring, stop conditions, and rollback.
The DSA does not prescribe a database schema or observability stack. This mapping is our engineering interpretation. Without it, the risk assessment sits in a document, the launch decision in a tracker, and runtime events across several logging systems. An audit becomes a reconstruction project.

A risk record should point to a release version. A metric should point to its definition and dashboard. A launch decision needs an accountable person or role. A feature flag and a tested rollback path carry more weight than a promise to “switch it off quickly.”

Give data access its own boundary

Very large platforms and search engines must provide data necessary for supervision to the Commission or the relevant Digital Services Coordinator after a reasoned request. Vetted researchers use a separate Article 40 process to obtain data for studying systemic risks.

Direct access to production databases creates risks for privacy, trade secrets, and service stability. Ad hoc exports fail more quietly: fields change, metric definitions drift, and two datasets with the same label stop meaning the same thing.

A usable access layer needs versioned schemas, a data dictionary, de-identification rules, access logs, and a reproducible sampling procedure. It also needs a boundary between evidence required for review and information whose disclosure would create another risk.

Scale controls by reach and consequence

Auditability costs time and infrastructure. Risk records, evaluation suites, instruction versioning, richer logs, and controlled exports add work before release. Building the same stack for a small internal assistant wastes time. We scale controls with the breadth of information the system can reach and the consequence of its answer.

  • User-provided context. Track model and instruction versions, failures, and clear retention boundaries.
  • A closed knowledge base. Add document identifiers, permissions, citations, retrieval evaluations, and a reliable way to remove outdated material from the index.
  • The open web and a public audience. Add query traces, source sets, diversity checks, harmful-output scenarios, staged rollout, and rapid rollback.
  • Large scale or sensitive consequences. Add formal risk assessment, independent-review readiness, reproducible reports, and a controlled data-access layer.

Our rule is simple: once AI selects sources and turns them into one answer, retrieval and generation should be designed as a single auditable decision chain.

We build AI chatbots, RAG systems, and language-model integrations. In an AI product review, we can identify the level of control a use case needs and design it into the architecture before the product scales.

Top comments (0)