Catalog Markers - README Auto-Update System
Source-of-truth counts driven by COURSE_CATALOG.generated.json. Markers in README files are expanded by scripts/notebook_tools/expand_catalog_markers.py and verified by CI on PRs that touch notebooks, series READMEs, or the catalog — the catalog-drift.yml workflow is filtered by paths:, so it does not run on every PR.
Overview
Instead of hardcoding notebook counts in README files, the project uses catalog markers — special HTML comments that get their values from the generated catalog. This ensures READMEs stay in sync with the actual notebook inventory automatically.
Data flow:
generate_catalog.py → COURSE_CATALOG.generated.json → expand_catalog_markers.py → README files
↑
CI checks for drift (catalog-drift.yml)
Marker Types
CATALOG-STATUS (multi-line block)
The primary marker. A multi-line HTML comment block placed near the top of a README.
Root README (MyIA.AI.Notebooks/README.md):
<!-- CATALOG-STATUS
series: ALL
total: 448
breakdown: GenAI=99, QuantConnect=93, SymbolicAI=90, ...
maturity: ALPHA=232, DRAFT=126, BETA=61, PRODUCTION=29
updated: 2026-05-02
-->Series README (e.g., MyIA.AI.Notebooks/ML/README.md):
<!-- CATALOG-STATUS
series: ML
pedagogical_count: 30
breakdown: _root=30
maturity: ALPHA=30
updated: 2026-05-02
-->Fields:
| Field | Root README | Series README |
|---|---|---|
series |
ALL |
Serie name (e.g., ML, GenAI) |
total |
Total notebook count | — |
pedagogical_count |
— | Notebooks in this serie |
breakdown |
Per-serie counts | Per-sous_serie counts |
maturity |
Global maturity distribution | Serie maturity distribution |
updated |
Last expansion date | Last expansion date |
CATALOG:counter (inline)
Inline markers for embedding counts in prose or tables. Not yet used in current READMEs but supported by the expand script.
<!-- CATALOG:counter:total --> → total notebooks
<!-- CATALOG:counter:serie=ML --> → ML notebooks
<!-- CATALOG:counter:serie=ML;status=READY --> → ML notebooks with status READY
<!-- CATALOG:counter:serie=ML;maturity=PRODUCTION --> → PRODUCTION ML notebooksScript Usage
# Expand all READMEs (idempotent)
python scripts/notebook_tools/expand_catalog_markers.py
# Dry-run (show what would change)
python scripts/notebook_tools/expand_catalog_markers.py --dry-run
# Check for drift (exit 1 if stale; local tool -- CI regenerates instead and never calls --check)
python scripts/notebook_tools/expand_catalog_markers.py --check
# Expand a specific file
python scripts/notebook_tools/expand_catalog_markers.py --file MyIA.AI.Notebooks/ML/README.md
# Use a different catalog file
python scripts/notebook_tools/expand_catalog_markers.py --catalog /path/to/catalog.jsonThe script is idempotent: running it twice produces identical output. It reads the catalog, regenerates each CATALOG-STATUS block, and only writes if the content differs.
CI Integration
The catalog-drift.yml workflow runs on PRs that touch notebooks or the catalog:
on:
pull_request:
paths:
- 'MyIA.AI.Notebooks/**/*.ipynb'
- 'MyIA.AI.Notebooks/**/README.md'
- 'COURSE_CATALOG.generated.json'Le job régénère le catalogue et les marqueurs sur le runner (rien n’est réécrit sur la branche), puis compare le résultat aux fichiers commités par un unique test git diff --cached. La dérive est remontée en annotation notice uniquement.
Ce check est advisory, non bloquant (#15998). Le marqueur advisory dans le nom du job — Notebook catalog drift (read-only, advisory) — est le contrat : pr_gate.py classe les checks par nom et ne lit pas fast_lane_registry.py. Le job peut rougir : seule l’indisponibilité des métadonnées git (rc=2 de generate_catalog.py) est absorbée en annotation notice ; tout autre échec (runner, checkout, pip, generate_catalog.py hors rc=2) exécute exit "$rc" et rend le job rouge. Mais ce rouge est exclu des causes bloquantes : le marqueur advisory du nom fait que PR gate le signale sans bloquer — une panne d’infrastructure est remontée, jamais bloquante pour une PR notebook/README (contrôle positif #16015). Le catalogue est régénéré quotidiennement sur main par catalog-cron.yml ; aucune action manuelle n’est requise sur une branche de feature (cf catalog-pr-hygiene.md, #2632).
Correction factuelle (2026-09-15) : cette section décrivait deux checks en séquence (
expand_catalog_markers.py --check,verify_catalog_readme.py) et concluait qu’un échec bloquait la PR jusqu’à mise à jour des marqueurs. Les deux affirmations étaient fausses : le workflow n’utilise pas--check(il régénère), ne fait appel àverify_catalog_readme.pydans aucun workflow, et son job est advisory depuis #15998.
Adding Markers to a New README
- Add a
<!-- CATALOG-STATUS ... -->block near the top of the README - Set the
series:field to the serie name (useALLfor root) - Run
python scripts/notebook_tools/expand_catalog_markers.pyto populate values - Commit the updated README
Regenerating the Catalog
If notebook counts change (new notebooks added/removed):
python scripts/notebook_tools/generate_catalog.py
python scripts/notebook_tools/expand_catalog_markers.py
git add COURSE_CATALOG.generated.json MyIA.AI.Notebooks/*/README.md
git commit -m "feat(catalog): regenerate catalog and update markers"