validate.yml triggered on pull_request alone, and it holds the only unittest invocation in the repository: deploy-site.yml stops at validate_schemas, generation and the freshness diff. Work lands on main by direct push far more often than by pull request, so 1,318 cases were guarding the road almost nothing takes. The suite and the schema check now run on both events. validate-bios and label-pr read pull request context and carry an event guard. The concurrency group falls back to the ref, so a push series collapses to the tip: what stays verified is the head of main. The path lists are spelled out per event because the workflow parser reads no YAML anchor, which PyYAML would have accepted in silence. Four tests hold the wiring: the suite reachable from a push, the two path lists equal, every job reading pull request context guarded, and no anchor in any workflow.
9.4 KiB
Release Process
This page documents the CI/CD pipeline: what each workflow does, how releases are built, and how to run the process manually.
CI workflows overview
The project uses 2 GitHub Actions workflows. All use only official GitHub
actions (actions/checkout, actions/setup-python, actions/upload-pages-artifact,
actions/deploy-pages). No third-party actions.
Budget target: ~175 minutes/month on the GitHub free tier.
| Workflow | File | Trigger |
|---|---|---|
| Deploy Site | deploy-site.yml |
Push to main (platforms, emulators, provenance, wiki, scripts, database.json, mkdocs.yml), manual |
| Validation | validate.yml |
PR and push to main touching bios/**, platforms/** or emulators/** |
Upstream BIOS lists are not scraped on a schedule. A maintainer runs the
scrapers by hand (see adding a scraper), reviews the
diff, and commits the refreshed platform YAML. Releases are built on the
maintainer's machine and uploaded with gh, see
cutting a release: the packs weigh 25 GB, more than a
hosted runner should rebuild and re-upload.
deploy-site.yml - Deploy Documentation Site
Trigger. Push to main when any of these paths change: platforms/,
emulators/, provenance/, wiki/, scripts/generate_site.py,
scripts/generate_readme.py, scripts/verify.py, scripts/common.py,
database.json, mkdocs.yml. Also manual dispatch.
The list is the set of inputs the site is generated from. Adding a new input to
generate_site.py means adding its path here, or the site silently goes stale.
Steps:
- Checkout, Python 3.12
- Install
pyyaml,mkdocs-material>=9.7.5,<10,pymdown-extensions>=10.14 - Restore large files from the
large-filesrelease, refresh data directories - Run
generate_site.py(converts YAML data into MkDocs pages and rewritesmkdocs.yml) - Run
generate_readme.py(rebuilds README.md and CONTRIBUTING.md) mkdocs build --strictto produce the static site- Run
validate_site.pyon the rendered HTML (metadata, headings, image alternatives, duplicate ids, local links and fragments) - Require the committed README and CONTRIBUTING to match what the generator
just produced.
write_if_changed()compares content with the timestamp line stripped, so a run that only moves the clock leaves the files untouched and the check stays meaningful - Upload artifact, deploy to GitHub Pages
Data contracts are validated with scripts/validate_schemas.py before the site
is generated: database.json, the install and target manifests, the site API
envelopes and the stats file, plus the semantic invariants those schemas cannot
express (declared totals matching their lists, no destination both installed
and omitted).
The site is deployed via the github-pages environment using the official
actions/deploy-pages action. Pages deployments are queued rather than
cancelled (cancel-in-progress: false): cancelling one mid-flight leaves the
deployment stuck and the next runs time out waiting on it.
--strict turns MkDocs warnings into failures, so a broken internal link or a
dangling anchor fails the build instead of shipping. The validation: block in
mkdocs.yml is what promotes unrecognized links and missing anchors to
warnings in the first place.
The theme version is pinned on both sides: >=9.7.5 because that is the
release which caps mkdocs < 2 (MkDocs 2.0 ships without a license), <10
so a major theme release cannot change the site without a deliberate bump.
validate.yml - Validation
Trigger. Pull requests and direct pushes to main that modify bios/**,
platforms/**, emulators/**, schemas/**, scripts/**, tests/** or
install.py. The path lists are spelled out once per event because the
workflow parser reads no YAML anchor.
Concurrency. Per-PR group on a pull request, per-ref on a push, cancel in-progress either way: a push series collapses to the tip.
Four jobs, two of which read pull request context and carry an event guard:
validate-bios (pull requests only). Diffs the PR to find changed BIOS
files, runs validate_pr.py --markdown on each, and posts the validation
report as a PR comment (hash verification, database match status).
validate-configs. Runs python scripts/validate_schemas.py --source-only,
which validates every platform YAML against schemas/platform.schema.json and
every emulator profile against schemas/emulator.schema.json. Both schemas set
additionalProperties: false, so a typo in a field name fails the job instead
of being silently ignored.
run-tests. Runs python -m unittest discover tests -v. Must pass before a
merge, and again on the commit a direct push puts at the head of main.
label-pr (pull requests only). Auto-labels the PR based on changed paths:
| Path pattern | Label |
|---|---|
bios/ |
bios |
bios/{Manufacturer}/ |
system:{manufacturer} |
platforms/ |
platform-config |
scripts/ |
automation |
Large files management
Files larger than 50 MB are stored as assets on a permanent GitHub release
named large-files (to keep the git repository lightweight).
Examples: PS3UPDAT.PUP, PSVUPDAT.PUP, PSP2UPDAT.PUP, the DSi NAND images,
maclc3.zip, Firmware.19.0.0.zip (Switch), the QEMU EDK2 firmware, the ScummVM
data bundle, the EasyRPG soundfont, the Dolphin/Ishiiruka SD card images, and
the arcade sets over 100 MB. .gitignore is the authoritative list: every
bios/ path listed there is a release asset.
Storage. Listed in .gitignore so they stay out of git history. The
large-files release is excluded from cleanup (the build workflow only
deletes version-tagged releases).
Build-time restore. The build workflow downloads all assets from
large-files into .cache/large/ and copies them to their expected paths
before pack generation.
Upload. To add or update a large file:
gh release upload large-files "bios/Sony/PS3/PS3UPDAT.PUP#PS3UPDAT.PUP"
Local cache. generate_pack.py calls fetch_large_file() which downloads
from the release and caches in .cache/large/ for subsequent runs.
Cutting a release
Releasing is deliberate and local. Nothing on GitHub builds a pack: the
pipeline runs here, the archives are checked here, and gh uploads them.
# 1. Full pipeline, online, so data directories and MAME/FBNeo hashes are fresh
python scripts/pipeline.py
# 2. RetroPie, which is archived but still served
python scripts/generate_pack.py --platform retropie --output-dir dist/
python scripts/generate_pack.py --platform retropie --verify-packs --output-dir dist/
# 3. Checksums of the full ZIPs, before splitting
(cd dist && sha256sum *.zip > SHA256SUMS.txt)
# 4. Split anything over 2 GB (GitHub asset cap); 7-Zip and PeaZip open .001 directly
for f in dist/*.zip; do
[ "$(stat -c%s "$f")" -gt 2000000000 ] || continue
split --bytes=1900M --numeric-suffixes=1 --suffix-length=3 "$f" "$f." && rm "$f"
done
# 5. The two sizes a pack has. Extracted runs well above downloaded, so the
# notes table names the one it carries.
python3 - <<'PY'
import json, pathlib, sys
sys.path.insert(0, "scripts")
from download import _match_key
parts = {}
for path in sorted(pathlib.Path("dist").glob("*_BIOS_Pack.zip*")):
parts.setdefault(path.name.split(".zip")[0] + ".zip", []).append(path)
def size(n):
return f"{n / 1024 ** 3:.1f} GB" if n >= 1024 ** 3 else f"{n / 1024 ** 2:.0f} MB"
for manifest in sorted(pathlib.Path("install").glob("*.json")):
base = next((b for b in parts if _match_key(manifest.stem) in _match_key(b)), None)
if not base:
continue
data = json.loads(manifest.read_text())
download = sum(p.stat().st_size for p in parts[base])
print(f"{base:<42} download {size(download):>8}"
f" extracted {size(data['total_size']):>8} {data['total_files']} files")
PY
# 6. Create the release as a DRAFT, upload every asset, and only then publish it.
# A public release with half its assets is a broken download for everyone
# during the whole upload.
DATE=$(date +%Y.%m.%d)
gh release create "v${DATE}" --draft --title "BIOS Pack v${DATE}" --notes-file notes.md
for f in dist/SHA256SUMS.txt dist/*.zip dist/*.zip.0*; do
[ -f "$f" ] && gh release upload "v${DATE}" "$f#$(basename "$f")" --clobber
done
gh release view "v${DATE}" --json assets --jq '.assets | length' # expect every file
gh release edit "v${DATE}" --draft=false --latest
# 7. Keep only the new release plus large-files: an older pack carries hashes
# the platforms no longer check, so it misleads more than it helps
gh release list --json tagName,createdAt \
--jq 'sort_by(.createdAt) | reverse | .[].tagName' | grep -v '^large-files$' \
| tail -n +2 | while read tag; do gh release delete "$tag" --yes --cleanup-tag; done
One pack per platform, the full one: the platform's list plus everything its
cores load. Platform-only and per-emulator packs are build options, not
release assets, since a lighter pack means a core that fails with no message.
The release notes follow the previous release: the quick install commands,
the pack table, what changed since the previous tag, and the contributors of
the closed issues. A pack has two sizes and they are far apart: Batocera
downloads as 2.4 GB and extracts to 4.0 GB. Step 5 prints both, and whichever
one the table carries, the header names it. Someone sizing a USB drive is
reading that column. The README table is the extracted size, from the install
manifests. SHA256SUMS.txt lists the checksums of the full ZIPs before
splitting.