TL;DR
On a Hugging Face official leaderboard, the default view silently drops any model whose card declares a base_model field. That one line of YAML, added so people can see what your model was built from, moves your entry out of sight. We measured the effect live on 2026-10-07 across every official benchmark board: 461 of 1,482 entries (31.1%) are hidden in the default view, and 35 of 44 populated boards hide at least one entry. The fix is a single API parameter, base_model=false, and this post gives you a reproducible script plus an author checklist so your own submission does not vanish.
This is a measurement note, not a complaint. The filter is a reasonable default. The problem is that most authors do not know it applies to them.
What is the base_model filter and why does it hide entries?
Model cards on the Hub carry YAML metadata. One common field is base_model, which records the parent a model was fine-tuned, merged, quantized, or adapted from:
base_model: Qwen/Qwen3-8B
That field is genuinely useful. It builds the model tree, powers "derived from" links, and tells readers the lineage at a glance. But Hugging Face leaderboards use it as a signal too. The default leaderboard view treats any card that declares a base_model as a derivative and leaves it out, so that the board shows mostly from-scratch base models rather than the long tail of fine-tunes, merges, and quantized variants.
The intent is clean boards. The side effect is that fine-tunes, LoRA merges, GGUF and AWQ quantizations, and distillations all disappear from the first screen most visitors ever look at, even when they rank at or near the top.
How large is the hidden share, measured today?
We pulled every dataset tagged benchmark:official from the public Hub API, then fetched each attached leaderboard twice: once with default parameters and once with base_model=false. The difference between the two is exactly the set of entries the filter removes.
Snapshot taken 2026-10-07:
| Quantity | Value |
|---|---|
Datasets tagged benchmark:official
|
48 |
| Boards with at least one entry | 44 |
| Entries visible by default | 1,021 |
Entries with base_model=false
|
1,482 |
| Hidden by the filter | 461 |
| Hidden share | 31.1% |
| Boards hiding at least one entry | 35 of 44 |
Nearly one in three entries across the official boards sits outside the default view. This is not a few stray cards. It is a structural feature of how the ecosystem publishes derivative models, which is to say most of it.
Run the same script on another day and the exact counts will move. That is expected. The method is fixed, each run records the state of the boards on its own date, and the share has held near 30 to 31% across our recent snapshots.
How do I see the hidden entries?
The leaderboard API accepts a base_model query parameter. Setting it to false disables the derivative filter and returns the full population:
GET https://huggingface.co/api/datasets/{dataset_id}/leaderboard
GET https://huggingface.co/api/datasets/{dataset_id}/leaderboard?base_model=false
The first call is what the default web view shows. The second is everything. For every board, the larger response is the full population and the default response is the visible subset.
Can I reproduce the 31.1% number?
Yes. The script below needs only requests and finishes in under a minute. It fetches both views of every official board and computes the hidden share directly.
import requests
import concurrent.futures as cf
H = "https://huggingface.co/api"
ds = requests.get(f"{H}/datasets",
params={"filter": "benchmark:official", "limit": 200},
timeout=60).json()
ids = [d["id"] for d in ds]
def fetch(i):
def get(params):
try:
r = requests.get(f"{H}/datasets/{i}/leaderboard",
params=params, timeout=60).json()
return r if isinstance(r, list) else []
except Exception:
return []
return i, get({}), get({"base_model": "false"})
rows = list(cf.ThreadPoolExecutor(8).map(fetch, ids))
n_default = n_full = boards = affected = 0
for bid, default, full in rows:
full = full if len(full) >= len(default) else default
n_default += len(default)
n_full += len(full)
if full:
boards += 1
if len(full) > len(default):
affected += 1
print("datasets", len(ids))
print("boards with entries", boards)
print("visible", n_default, "full", n_full)
print("hidden", n_full - n_default,
"hidden %", round(100 * (1 - n_default / n_full), 1))
print("boards affected", affected, "of", boards)
On 2026-10-07 this prints 48 datasets, 44 boards with entries, 1,021 visible, 1,482 full, 461 hidden (31.1%), and 35 of 44 boards affected.
Why this matters for authors
If you fine-tune, merge, quantize, or distill a model and submit it to an official board, you probably added base_model to the card without a second thought. It is good lineage hygiene, and many upload scripts add it automatically. But the moment it is present, your result leaves the default screen. A visitor who does not know to toggle the filter, which is most of them, never sees your score. A derivative that legitimately ranks first can be invisible on the board it tops.
This is a documented ecosystem behavior, not a bug, and it is easy to work around once you know it exists. The goal is not to game the board. It is to make an honest result visible.
Checklist: keep your entry from being hidden
-
Know the rule before you submit. Any
base_modelfield in your card moves the entry out of the default view. Decide on purpose, do not discover it after the fact. -
Check your card after upload. Open the raw card and confirm whether
base_modelis present. Many trainers and quantization tools inject it silently. -
Verify both views. Fetch
/leaderboardand/leaderboard?base_model=falsefor your board and confirm which list your entry is in. -
If visibility matters, reconsider the field. If your model is a legitimate standalone result, you may choose to omit
base_model. Never fabricate lineage, and never strip a field only to deceive. Honesty first, visibility second. -
If you keep the field, say so. State in your card and README that the entry appears only with the derivative filter off, and link the
base_model=falseview so readers can find it. - Re-check on release day. Boards re-index. Confirm your entry shows where you expect on the day you announce it.
FAQ
Does declaring base_model lower my score?
No. It changes visibility, not value. The score is identical in both views. The entry is simply filtered out of the default list.
Is the filter a bug?
No. It is intended behavior that keeps boards focused on from-scratch models. The issue is that most authors do not realize it applies to their derivative, so a real result goes unseen.
How do I see every entry on a board?
Add ?base_model=false to the leaderboard API call, or use the filter toggle in the board's web UI. That returns the full population including fine-tunes, merges, and quantized variants.
Will the 31.1% change over time?
Yes. It is a live snapshot from 2026-10-07. The method stays fixed and each run records that day's state. Across recent snapshots the share has stayed near 30 to 31%.
I quantized a popular model. Is my entry hidden?
Almost certainly, if your card declares the source via base_model, which quantization tooling often adds. Quantized variants are exactly the derivative class this filter removes.
Should I remove base_model to climb the board?
Only if your model truly is a standalone result. Do not strip accurate lineage to deceive. If you keep the field, document the hidden status and link the unfiltered view instead.
Method notes
All numbers come from the public Hugging Face Hub API on 2026-10-07. We counted a board as affected when its base_model=false response contained more entries than its default response. The full population per board is the larger of the two responses. No authentication is required, and the script above is the complete method.
Disclosure: QuantID publishes measurement and diagnostics notes. VIDRAFT is a technology alliance partner.
Further reading on this account: "Who Leads Hugging Face's Official Benchmarks? Concentration, Gaps and Hidden Entries" and "How Is Quantum Computer Performance Actually Measured? Quantum Volume Explained."
Top comments (1)
Con số 31% thực sự gây bất ngờ vì nó cho thấy một lượng lớn dữ liệu chất lượng đang bị "ẩn" khỏi tầm mắt của những người đang tìm kiếm model để fine-tune. Việc khai báo base_model giúp đảm bảo tính minh bạch về nguồn gốc và giúp người dùng so sánh công bằng hơn giữa các phiên bản kế thừa, thay vì chỉ nhìn vào kết quả benchmark bề nổi. Tuy nhiên, điều này cũng vô tình tạo ra một rào cản thông tin cho những dev mới, khi họ không biết rằng có cả một kho tàng model đã được tối ưu hóa dựa trên các nền tảng lớn đang nằm ngoài danh sách mặc định.