
Museums and historic city centres keep adding rules about how a tour guide may talk to a group. Some require headsets above a certain group size, some ban loudspeakers, and some do not let an outside guide explain anything indoors. The rules live in PDFs, booking terms and city ordinances, in a dozen languages, and nobody had put them in one table.
We work on a small app for tour guides (Your Next Tours), so we needed that table anyway. This post is about how we built it, what the schema looks like, and the mistakes we tried to avoid. The result is open under CC BY 4.0 on Hugging Face and Kaggle.
Start wide, then throw things away
The first pass was a survey of 228 candidate rules. That list was useful for finding places, but it was not trustworthy. Some entries came from travel blogs, some quoted an old version of a rule, a few were simply wrong.
So the second pass had one job: for every candidate, open the official source again and find the quoted sentence in it. We split the list into batches and let LLM agents do the opening and searching, with one strict instruction: if the sentence is not in the official text, the record is marked not-found and does not become a rule. No search engines, only the institution's own site, its PDFs, or an archived copy of that page.
Each checked record came back as a small JSON object:
{
"id": "fr-musee-du-louvre",
"class": "A",
"threshold": 7,
"thresholdCountsGuide": null,
"devicePolicy": "own-allowed",
"status": "in-force",
"quote": "...",
"quoteLang": "fr",
"sourceUrl": "https://...",
"verifiedAt": "2026-10-02",
"verdict": "confirmed"
}
A build script then keeps a record only if the verdict is confirmed or changed, the class is A, B or C, the rule is in force or formally proposed, the source is https and the quote is not empty. Everything else is dropped. In the end 197 records passed, plus one Turkish rule (the National Palaces in Istanbul) that we had already documented separately, so the dataset has 198 rows from 36 countries.
Three classes, because "headset rule" means three different things
Reading the texts side by side, they fell into three groups:
- A, headsets required. A written rule says guided groups must use headsets or a whisper system (66 rules).
- B, loudspeakers banned. No megaphones or amplification. Headsets are not mandatory, but with a big group they are the practical way to be heard (75 rules).
- C, no outside guiding inside. Only the site's own guides or audio guides may narrate indoors (57 rules).
Keeping B separate from A mattered. Many city ordinances only ban amplification, and lumping them in with headset requirements would have inflated the headline number.
The fields that took the longest to decide
threshold. Texts say "more than 10", "10 or more", "groups of 5 to 30". We store the smallest group size the rule applies to, so "more than 8" becomes 9. An empty value means every group. Where a number exists, it ranges from 3 to 26, and the most common values are 7 and 11.
threshold_counts_guide. Some texts count the guide, most say nothing. We store true, false or empty, and empty means "the text does not say". We did not guess.
device_policy. This is the field people care about most, and the one easiest to get wrong. The values are own-allowed, unspecified, institution-only, app-banned, phone-banned and headphones-banned. 132 of the 198 rules are unspecified: the text asks for headsets but does not say whose. That does not mean any device is accepted, and the dataset card says so in plain words. Only 39 rules explicitly allow a group's own system.
quote and source_url. Every row carries the verbatim sentence in its original language and the official link, so anyone can check our reading. Rights to the quoted text stay with the issuing body.
Loading it
The CSV is a single file, so plain pandas is enough:
import pandas as pd
url = ("https://huggingface.co/datasets/first-point/tour-group-headset-rules"
"/resolve/main/data/tour-group-headset-rules-198-20261002.csv")
df = pd.read_csv(url)
print(df.shape) # (198, 19)
print(df["rule_class"].value_counts())
# Rules with a group-size threshold, smallest first
cut = df.dropna(subset=["threshold"]).sort_values("threshold")
print(cut[["place", "country", "threshold", "device_policy"]].head(10))
There is also a JSONL file with the same rows and typed values, which is easier to drop into a RAG index.
One source, three outputs
The same JSON fact packs feed three things: the human-readable list at yournext.tours/tour-group-headset-rules, the Hugging Face dataset and the Kaggle dataset. Each row's page_url points to its anchor on that page, so a row like fr-musee-du-louvre links straight to the Louvre entry. Generating all three from one place is the only way we found to keep the numbers from drifting apart.
What we would do differently
- Write the dropping rules before collecting anything. We decided late that a quote must be found word for word, and some early candidates that "felt right" did not survive that test.
- Store
effective_fromas a separate field from the start. 38 of the rules took effect in 2024 or later, and one more (canal boats in Bruges) starts in 2029. That trend was the most interesting finding, and it was almost invisible in the first draft. - Plan for updates. Rules change; we will re-check the sources every quarter and publish a new snapshot.
If you know a rule we missed or read wrong, the comments here are a good place for it. We would rather fix a row than defend it.
Top comments (0)