Catalog Markers - README Auto-Update System

Source-of-truth counts driven by COURSE_CATALOG.generated.json. Markers in README files are expanded by scripts/notebook_tools/expand_catalog_markers.py and verified by CI on PRs that touch notebooks, series READMEs, or the catalog — the catalog-drift.yml workflow is filtered by paths:, so it does not run on every PR.

Overview

Instead of hardcoding notebook counts in README files, the project uses catalog markers — special HTML comments that get their values from the generated catalog. This ensures READMEs stay in sync with the actual notebook inventory automatically.

Data flow:

generate_catalog.py → COURSE_CATALOG.generated.json → expand_catalog_markers.py → README files
                                                                       ↑
                                                          CI checks for drift (catalog-drift.yml)

Marker Types

CATALOG-STATUS (multi-line block)

The primary marker. A multi-line HTML comment block placed near the top of a README.

Root README (MyIA.AI.Notebooks/README.md):

<!-- CATALOG-STATUS
series: ALL
total: 448
breakdown: GenAI=99, QuantConnect=93, SymbolicAI=90, ...
maturity: ALPHA=232, DRAFT=126, BETA=61, PRODUCTION=29
updated: 2026-05-02
-->

Series README (e.g., MyIA.AI.Notebooks/ML/README.md):

<!-- CATALOG-STATUS
series: ML
pedagogical_count: 30
breakdown: _root=30
maturity: ALPHA=30
updated: 2026-05-02
-->

Fields:

Field Root README Series README
series ALL Serie name (e.g., ML, GenAI)
total Total notebook count —
pedagogical_count — Notebooks in this serie
breakdown Per-serie counts Per-sous_serie counts
maturity Global maturity distribution Serie maturity distribution
updated Last expansion date Last expansion date

CATALOG:counter (inline)

Inline markers for embedding counts in prose or tables. Not yet used in current READMEs but supported by the expand script.

<!-- CATALOG:counter:total -->                 → total notebooks
<!-- CATALOG:counter:serie=ML -->              → ML notebooks
<!-- CATALOG:counter:serie=ML;status=READY --> → ML notebooks with status READY
<!-- CATALOG:counter:serie=ML;maturity=PRODUCTION --> → PRODUCTION ML notebooks

Script Usage

# Expand all READMEs (idempotent)
python scripts/notebook_tools/expand_catalog_markers.py

# Dry-run (show what would change)
python scripts/notebook_tools/expand_catalog_markers.py --dry-run

# Check for drift (exit 1 if stale; local tool -- CI regenerates instead and never calls --check)
python scripts/notebook_tools/expand_catalog_markers.py --check

# Expand a specific file
python scripts/notebook_tools/expand_catalog_markers.py --file MyIA.AI.Notebooks/ML/README.md

# Use a different catalog file
python scripts/notebook_tools/expand_catalog_markers.py --catalog /path/to/catalog.json

The script is idempotent: running it twice produces identical output. It reads the catalog, regenerates each CATALOG-STATUS block, and only writes if the content differs.

CI Integration

The catalog-drift.yml workflow runs on PRs that touch notebooks or the catalog:

on:
  pull_request:
    paths:
      - 'MyIA.AI.Notebooks/**/*.ipynb'
      - 'MyIA.AI.Notebooks/**/README.md'
      - 'COURSE_CATALOG.generated.json'

Le job régénère le catalogue et les marqueurs sur le runner (rien n’est réécrit sur la branche), puis compare le résultat aux fichiers commités par un unique test git diff --cached. La dérive est remontée en annotation notice uniquement.

Ce check est advisory, non bloquant (#15998). Le marqueur advisory dans le nom du job — Notebook catalog drift (read-only, advisory) — est le contrat : pr_gate.py classe les checks par nom et ne lit pas fast_lane_registry.py. Le job peut rougir : seule l’indisponibilité des métadonnées git (rc=2 de generate_catalog.py) est absorbée en annotation notice ; tout autre échec (runner, checkout, pip, generate_catalog.py hors rc=2) exécute exit "$rc" et rend le job rouge. Mais ce rouge est exclu des causes bloquantes : le marqueur advisory du nom fait que PR gate le signale sans bloquer — une panne d’infrastructure est remontée, jamais bloquante pour une PR notebook/README (contrôle positif #16015). Le catalogue est régénéré quotidiennement sur main par catalog-cron.yml ; aucune action manuelle n’est requise sur une branche de feature (cf catalog-pr-hygiene.md, #2632).

Correction factuelle (2026-09-15) : cette section décrivait deux checks en séquence (expand_catalog_markers.py --check, verify_catalog_readme.py) et concluait qu’un échec bloquait la PR jusqu’à mise à jour des marqueurs. Les deux affirmations étaient fausses : le workflow n’utilise pas --check (il régénère), ne fait appel à verify_catalog_readme.py dans aucun workflow, et son job est advisory depuis #15998.

Adding Markers to a New README

  1. Add a <!-- CATALOG-STATUS ... --> block near the top of the README
  2. Set the series: field to the serie name (use ALL for root)
  3. Run python scripts/notebook_tools/expand_catalog_markers.py to populate values
  4. Commit the updated README

Regenerating the Catalog

If notebook counts change (new notebooks added/removed):

python scripts/notebook_tools/generate_catalog.py
python scripts/notebook_tools/expand_catalog_markers.py
git add COURSE_CATALOG.generated.json MyIA.AI.Notebooks/*/README.md
git commit -m "feat(catalog): regenerate catalog and update markers"
Retour au sommet