fix: state one catalog ratio, correct the faq

The home page and the stats export counted every file carrying a
provenance record, the provenance page and the README only the system
files. The site published 553 and 566 for the same quantity, one click
apart, and the export paired the wider count with composition.systems as
its denominator. common.count_catalog_matched is now the single source,
scoped to the systems bucket.

The FAQ had drifted from the profiles it describes: MAME pinned at 0.287
against 0.289 in mame.yml, Adler-32 attributed to Dolphin's IPL rather
than the DSP ROMs that carry known_hash_adler32, and the per-emulator
verbose report named as the only content check on an existence platform,
which skips the DISCREPANCY line the platform report raises itself.

Tests read both sides: no generator may count matches inline, and each
FAQ claim is checked against the profile or the script that owns it.
This commit is contained in:
Abdessamad Derraz committed 2026-09-04 11:41:17 +02:00
1 parent 022888e9ea
commit fe77535c3b
5 files changed
+184 -52

No files matched your search

+27 -8
View File
@@ -950,6 +950,17 @@ def unique_emulator_profiles(profiles: dict[str, dict]) -> dict[str, dict]:
GAME_DATA_TOPS = ("RPG Maker", "ScummVM")
def composition_tier(path: str) -> str:
"""Which composition bucket a repository path belongs to."""
parts = path.split("/")
top = parts[1] if len(parts) > 1 else ""
if top == "Arcade":
return "arcade"
if top in GAME_DATA_TOPS:
return "game_data"
return "systems"
def compute_composition(db: dict) -> dict:
"""File and byte counts by tree area.
@@ -963,19 +974,27 @@ def compute_composition(db: dict) -> dict:
"game_data": {"files": 0, "size_bytes": 0},
}
for entry in db.get("files", {}).values():
parts = entry.get("path", "").split("/")
top = parts[1] if len(parts) > 1 else ""
if top == "Arcade":
bucket = buckets["arcade"]
elif top in GAME_DATA_TOPS:
bucket = buckets["game_data"]
else:
bucket = buckets["systems"]
bucket = buckets[composition_tier(entry.get("path", ""))]
bucket["files"] += 1
bucket["size_bytes"] += entry.get("size", 0)
return buckets
def count_catalog_matched(db: dict) -> int:
"""System files byte-identical to a dump-preservation catalog entry.
Scoped to the systems bucket, the denominator every surface pairs it
with: No-Intro, Redump and TOSEC index console and computer dumps, so
arcade ROM sets and engine data sit on neither side of the ratio.
"""
return sum(
1
for entry in db.get("files", {}).values()
if entry.get("provenance")
and composition_tier(entry.get("path", "")) == "systems"
)
def group_identical_platforms(
platforms: list[str],
platforms_dir: str,