perf: read yaml through the c loader

Loading the emulator profiles is the most expensive step of every
command here, and all forty call sites used the pure-Python scanner
while libyaml sat unused in the same wheel. One shared yaml_load picks
the C loader when pyyaml ships it: the 375 profiles parse in 0.18s
instead of 1.39s, and verify --platform retroarch drops from 2.17s to
0.73s. The loader class is the same restricted one safe_load uses.

es_bios.xml was parsed straight from the network while install.py
already refused a document declaring entities; both now share one
guard. Scrapers reach it through a single path bootstrap in the
package rather than two ad-hoc ones.
This commit is contained in:
Abdessamad Derraz committed 2026-08-11 00:55:16 +02:00
1 parent ab6a3bb26d
commit 3b8f2d75d5
18 files changed
+109 -42

No files matched your search

+5 -4
View File
@@ -16,7 +16,8 @@ Recalbox verification logic:
from __future__ import annotations
import sys
import xml.etree.ElementTree as ET
from common import parse_untrusted_xml
from .base_scraper import BaseScraper, BiosRequirement
@@ -109,7 +110,7 @@ class Scraper(BaseScraper):
def _fetch_cores(self) -> list[str]:
"""Extract unique core names from es_bios.xml bios elements."""
raw = self._fetch_raw()
root = ET.fromstring(raw)
root = parse_untrusted_xml(raw, "es_bios.xml")
cores: set[str] = set()
for bios_elem in root.findall(".//system/bios"):
raw_core = bios_elem.get("core", "").strip()
@@ -128,7 +129,7 @@ class Scraper(BaseScraper):
if not self.validate_format(raw):
raise ValueError("es_bios.xml format validation failed")
root = ET.fromstring(raw)
root = parse_untrusted_xml(raw, "es_bios.xml")
requirements = []
seen = set()
@@ -177,7 +178,7 @@ class Scraper(BaseScraper):
def fetch_full_requirements(self) -> list[dict]:
"""Parse es_bios.xml preserving all Recalbox-specific fields."""
raw = self._fetch_raw()
root = ET.fromstring(raw)
root = parse_untrusted_xml(raw, "es_bios.xml")
requirements = []
for system_elem in root.findall(".//system"):