A CI checkout omits every file over 50 MB, so verify and generate_pack
resolve those database entries against a disk that does not hold them and
report them missing. The generated README then stops matching the
committed one for a reason that has nothing to do with staleness.
restore_large_files.py writes them back from the release cache, matched by
SHA1 rather than by name, and only where the path is gitignored and
absent. The site workflow runs it, and refreshes the data directories, before
generating.
PR validation keeps its four jobs, its comment and its labels. The two
inline schema heredocs become validate_schemas.py --source-only, the
changed-file list is NUL-separated so a path with a space survives, and
the whole suite runs instead of one module.
The offline pipeline, manifest regeneration and mkdocs build stay out of
the pull-request path: they belong to the site workflow, which only runs
on main, and stacking them on every PR does not fit the free-tier
budget.
Deploy Site gains schema validation and rendered-site validation, and
checks that the committed README and CONTRIBUTING match what the
generator produces. write_if_changed compares with the timestamp line
stripped, so that check reports staleness rather than the clock.
Replace grep-based restore with SHA1 matching via database.json.
The old grep heuristic failed for assets with renamed basenames
(dsi_nand_batocera42.bin) or special characters (MAME dots vs
spaces), and only restored to the first .gitignore match when
multiple paths shared a basename.
Fix 3 broken data directory sources:
- opentyrian: buildbot URL 404, use release asset
- syobonaction: invalid git_subtree URL, use GitHub archive
- stonesoup: same fix, adds 532 game data files
generate_site.py resolves files on disk for gap analysis.
Without large files and data directories, the deployed site
showed 148 missing platform files and 207 unsourced core
complement files.