GenAI Services - ComfyUI Image Generation

Services disponibles

Service Hôte Modèle VRAM Description
Qwen Image Edit po-2023 qwen_image_edit_2509 ~29GB Edition d’images avec prompts multimodaux
Z-Image/Lumina po-2023 Lumina-Next-SFT ~10GB Generation text-to-image haute qualite
MiniMax H3 (Hailuo 3.0) non déployé (licence UE) MiniMax-H3 (open-weights) ~24+ GB VIDEO omni-modale + audio stereo natif (UE exclue par licence) — voir Architecture MiniMax H3

Architecture Qwen (Phase 29)

Workflow ComfyUI pour Qwen Image Edit 2509 :

VAELoader (qwen_image_vae.safetensors, 16 channels)
    |
CLIPLoader (qwen_2.5_vl_7b_fp8_scaled.safetensors, type: sd3)
    |
UNETLoader (qwen_image_edit_2509_fp8_e4m3fn.safetensors)
    |
ModelSamplingAuraFlow (shift=3.0)
    |
CFGNorm (strength=1.0)
    |
TextEncodeQwenImageEdit (clip, prompt, vae)
    |
ConditioningZeroOut (negative)
    |
EmptySD3LatentImage (16 channels)
    |
KSampler (scheduler=beta, cfg=1.0, sampler=euler)
    |
VAEDecode

Points critiques : - VAE 16 canaux (pas SDXL standard) - scheduler=beta obligatoire - cfg=1.0 (pas de CFG classique, utilise CFGNorm) - ModelSamplingAuraFlow avec shift=3.0

Architecture Z-Image/Lumina

Workflow ComfyUI simplifie avec LuminaDiffusersNode :

LuminaDiffusersNode (Alpha-VLLM/Lumina-Next-SFT-diffusers)
    |
VAELoader (sdxl_vae.safetensors)
    |
VAEDecode
    |
SaveImage

Paramètres LuminaDiffusersNode : - model_path: “Alpha-VLLM/Lumina-Next-SFT-diffusers” - num_inference_steps: 20-40 - guidance_scale: 3.0-5.0 - scaling_watershed: 0.3 - proportional_attn: true - max_sequence_length: 256

Note technique (Janvier 2025) : Le node utilise LuminaPipeline (diffusers 0.34+), ancien nom LuminaText2ImgPipeline obsolete.

Architecture MiniMax H3 (Hailuo 3.0) — VIDEO avec audio natif

Statut : INTRINSIC pour l’UE (Tâche 2 #10244). Le modele est open-weights et techniquement auto-hebergeable (ComfyUI workflows publies sur ai-models-lab/minimax-h3), mais la licence MiniMax H3 Community License (datee 2 aout 2026, deposee sur HuggingFace MiniMaxAI/MiniMax-H3) exclut explicitement l’usage, l’hebergement et l’affichage des Outputs sur le territoire de l’Union Europeenne (cf. Art. I.3 Applicable Territory et Art. I.5 Excluded Territories). Ce bloc documente donc l’architecture a titre de reference — pour informer la decision de deploiement — et non pour inviter au deploiement local.

Verdict SOTA pour notre contexte (France / UE, depot public, ecoles partenaires) : INTRINSIC. La voie pedagogiquement equivalente pour « video + audio natif synchronise » en UE est LTX-2 (02-5-LTX2-Audiovisual, licence permissive, executable). Le notebook 02-6-MiniMax-H3-Architecture-Licensing de la serie Video enseigne l’architecture et le raisonnement de conformite (verificateur de juridiction, matrice de decision avec Sora et LTX-2).

Specifications du modele (verifiees firsthand, source officielle minimax.io + GitHub MiniMax-AI/MiniMax-H3 + HuggingFace)

Caracteristique Valeur
Modalites d’entree Omni-modale unifiee : texte, image, video, audio
Sortie Video + audio stereo natif (32 kHz, 11 langues)
Resolution jusqu’a 2K
Duree clips de 5 a 15 s
Fps 24 fps
Modes de generation text-to-video · image-to-video · first+last frame · omni-reference (entrees mixtes)
Classement #1 Artificial Analysis video editing au lancement (Elo 1130)
VRAM (selon documentation, execution reelle hors UE) ~24 GB+ (quantization fp8 / GGUF Q4 envisageable)
Workflows ComfyUI hub communautaire ai-models-lab/minimax-h3 (non deploye ici)

Architecture du pipeline (telle que publiee par la communaute ComfyUI ; non deployee sur le cluster)

MiniMaxH3Loader (checkpoint MultiModal, fp8 ou GGUF Q4)
    |
CLIPLoader (H3 text-encoder, ~7B)
    |
OmniModalConditioning (text + image + audio refs)
    |
MiniMaxH3UNET (denoising + audio decoder couple)
    |
AudioVAEDecode (32 kHz stereo, 11 langues)
    |
VaeDecode (frames 2K, 24 fps, 5-15 s)
    |
SaveVideo + SaveAudio (synchronisation native)

Points critiques (selon documentation communautaire, non valides firsthand sur le cluster) : - VAE multi-modal (pas SDXL standard) — couples video frames + audio waveform - CFG 1.0 recommande pour eviter la degenerescence audio (semblable a Qwen Image Edit) - Scheduler beta recommande pour la convergence des clips longs - Omni-reference : permet d’injecter des refs image+audio simultanement (capacite distinctive vs Wan/SVD)

Comparaison synthetique avec les modeles Video de la serie

Modele Format Audio natif UE-ok Open-weights Notebook
HunyuanVideo text-to-video non oui oui 02-1 (~18 GB VRAM)
LTX-Video text-to-video non oui oui 02-2 (~8 GB VRAM)
Wan 2.1/2.2 text-to-video non oui oui 02-3 (~10 GB VRAM)
SVD image-to-video non oui oui 02-4 (~10 GB VRAM)
LTX-2 (Lightricks 22B) text-to-video OUI (stereo) oui oui 02-5 (~16-24 GB VRAM)
MiniMax H3 omni-modal OUI (stereo 11 langues) NON (UE exclue) oui (licence geo-restreinte) 02-6 (descriptif, INTRINSIC)
OpenAI Sora 2 text-to-video non oui (API) non (API fermee) 04-3

Etat reversible — ce qui changerait la decision

Le verdict INTRINSIC pour notre contexte n’est pas un mur infranchissable ; il est conditionne a la licence. Trois evenements le feraient basculer :

  1. MiniMax deplace l’Excluded Territory (ex. retire l’UE de la liste). Surveillance : MiniMaxAI/MiniMax-H3 releases / minimax.io/news (cf. notebook 02-6 §5).
  2. Licence commerciale UE obtenue (Art. II dernier § : « you are welcome to contact us about obtaining a license »). Demarche provider-side dependant de MiniMax.
  3. Fork communautaire re-licencie sous licence permissive (hypothese, aucun signal actuel a 2026-08-10).

Re-verification licence 2026-08-15 (firsthand, LICENSE raw HuggingFace, date 2 aout 2026) : UE toujours exclue — “Excluded Territories” means the European Union, the United Kingdom, the Republic of Korea and the United States of America. Aucun changement. Cote service cloud, la voie video-01 (/v1/video_generation) est couverte par le plan (probe 2026-08-10, notebook 04-5b) ; la serie H3 reste plan-gated (400 TokenPlan 2013, notebook 04-5 squelette idempotent).

En attendant, le notebook 02-6 evoluera vers une execution locale reelle des qu’un de ces evenements se materialise — l’architecture ci-dessus est le blueprint ComfyUI a deployer le moment venu.

Architecture Krea 2 turbo + VLM in-graph (03-4)

Workflow rejoué depuis les métadonnées d’un PNG (notebook Image/03-Orchestration/03-4) : une boucle de raisonnement dans le graphe ComfyUI — un VLM examine une image de référence quelconque et rédige lui-même le prompt de character design avant la diffusion.

LoadImage (reference quelconque)
    |
ImageScaleToTotalPixels (1.0 MP, lanczos)
    |
CLIPLoader (qwen3vl_4b_bf16.safetensors, type: qwen_image)   <- sert AUSSI de VLM
    |
TextGenerate (system prompt "visual DNA", 11235 chars, temp 0.7)
    |  -> PreviewAny (recuperation du texte genere via /history)
    v
StringConcatenate (trigger words LoRA + prompt VLM)
    |
LoraLoader (banjiesock_Krea2, model + clip)     UNETLoader (krea2_turbo_int8_convrot)
    |                                                  |
CLIPTextEncode -> KSampler (er_sde, simple, 8 steps, cfg 1.0)
                 ConditioningZeroOut (negatif nul, modele distille)
    |
VAEDecode (qwen_image_vae.safetensors, 16 canaux) -> SaveImage

Points critiques : - UNet int8 convrot 13,49 Go + VLM 8,9 Go sur RTX 3090 24 Go : offload séquentiel ComfyUI (le VLM travaille puis laisse la place au UNet) — les deux phases tiennent séparément. - TextGenerate est un node natif (comfy_extras/nodes_textgen.py) : il exige un CLIP dont la config expose stop_tokens (familles gemma4 / qwen35) ; qwen_2.5_vl_7b (type qwen_image) NE le supporte PAS (AttributeError: stop_tokens). - cfg=1.0 + 8 steps = modèle distillé (turbo) : négatif via ConditioningZeroOut, pas de prompt négatif factice. - Le texte généré se récupère via un node PreviewAny (output text dans /history) — sans node terminal texte, POST /prompt répond prompt_no_outputs. - Le workflow original (PNG auteur) utilisait gemma4_12b comme VLM et des custom nodes (LoraManager, ResolutionMaster, rgthree, KJNodes) : substitutions documentées dans le notebook (table custom -> natif). - Poids : Comfy-Org/Krea-2 (HF) pour UNet/encodeurs/VAE, Civitai modèle 2937073 pour le LoRA banjiesock (trigger words : banjiesock2style, illustration).

Approches abandonnees

Approche Raison abandon
Z-Image GGUF Incompatibilite dimensionnelle (2560 vs 2304) entre RecurrentGemma et Gemma-2
Qwen GGUF Non teste, prefer les poids fp8 pour qualite

Scripts de gestion (scripts/genai-stack/)

IMPORTANT pour agents : Utiliser le CLI unifie genai.py au lieu de demarrer des kernels MCP directement.

# CLI unifie - aide
python scripts/genai-stack/genai.py --help

# Gestion services Docker
python scripts/genai-stack/genai.py docker status          # Statut services
python scripts/genai-stack/genai.py docker start all       # Demarrer tous les services
python scripts/genai-stack/genai.py docker test --remote   # Tester endpoints (local + remote)

# Validation stack ComfyUI
python scripts/genai-stack/genai.py validate --full        # Validation complete
python scripts/genai-stack/genai.py validate --nunchaku    # Test Nunchaku INT4 Lightning

# Validation notebooks
python scripts/genai-stack/genai.py validate --notebooks   # Syntaxe notebooks GenAI
python scripts/genai-stack/genai.py notebooks              # Execution Papermill

# GPU et modeles
python scripts/genai-stack/genai.py gpu                    # Verification VRAM
python scripts/genai-stack/genai.py models list-nodes      # Custom nodes ComfyUI
python scripts/genai-stack/genai.py models list-checkpoints # Checkpoints disponibles

# Authentification
python scripts/genai-stack/genai.py auth audit             # Audit securite tokens
python scripts/genai-stack/genai.py auth sync              # Synchroniser tokens

Mapping notebooks GenAI Image -> services

Notebooks Service Prerequis
01-1, 01-3 OpenAI API (cloud) OPENAI_API_KEY
01-4, 02-3 SD Forge Service local ou myia.io
01-5, 02-1 ComfyUI Qwen COMFYUI_AUTH_TOKEN, ~29GB VRAM
02-4 Z-Image/vLLM ~10GB VRAM
03-* Multi-modèles Tous les services
03-4 ComfyUI Krea 2 (VLM in-graph) COMFYUI_AUTH_TOKEN, ~23GB VRAM (UNet int8 13,5 Go + VLM 8,9 Go, offload séquentiel)
04-* Applications Variable

Mapping notebooks Audio -> services

Notebooks Service Prerequis
Audio/01-1, 01-2 OpenAI API (TTS/STT) OPENAI_API_KEY
Audio/01-3 Local (librosa, pydub) Aucun
Audio/01-4 Whisper local GPU ~10 GB
Audio/01-5 Kokoro TTS GPU ~2 GB
Audio/02-1 Chatterbox TTS GPU ~8 GB
Audio/02-2 XTTS v2 GPU ~6 GB
Audio/02-3 MusicGen GPU ~10 GB
Audio/02-4 Demucs v4 GPU ~4 GB
Audio/03-* Multi-modèles Mixed
Audio/04-11 Kokoro TTS + FishAudio S2-Pro GPU ~5 GB (FishAudio BnB 4-bit NF4) + ~2 GB (Kokoro)
Audio/04-* Applications Mixed

Mapping notebooks Video -> services

Notebooks Service Prerequis
Video/01-1 Local (moviepy, FFmpeg) FFmpeg installe
Video/01-2 OpenAI GPT-5 OPENAI_API_KEY
Video/01-3 Qwen2.5-VL local GPU ~18 GB
Video/01-4 Real-ESRGAN/RIFE GPU ~4 GB
Video/01-5 AnimateDiff GPU ~12 GB
Video/02-1 HunyuanVideo GPU ~18 GB
Video/02-2 LTX-Video GPU ~8 GB
Video/02-3 Wan 2.1/2.2 GPU ~10 GB
Video/02-4 SVD GPU ~10 GB
Video/02-5 LTX-2 (Lightricks 22B) GPU ~16-24 GB (fp8-cast / GGUF Q4)
Video/02-6 MiniMax H3 (analyse architecture) 0 VRAM (descriptif, INTRINSIC UE)
Video/03-1 Multi-modeles locaux GPU ~18 GB
Video/03-2 Pipeline text-to-image-to-video Mixed
Video/03-3 ComfyUI Video Docker, nodes video
Video/04-1 Pipeline educatif Mixed
Video/04-2 Style transfer + music video Mixed
Video/04-3 Sora 2 API OPENAI_API_KEY
Video/04-4 Pipeline production complet Mixed

Configuration .env GenAI

Fichier : MyIA.AI.Notebooks/GenAI/.env

# Mode local (Docker) vs remote (myia.io)
LOCAL_MODE=false

# ComfyUI
COMFYUI_API_URL=https://qwen-image-edit.myia.io
# Hébergement : po-2023 (conteneur comfyui-qwen, RTX 3090 24 GB, port 8188)
COMFYUI_AUTH_TOKEN=<bearer_token_bcrypt>

# OpenAI via OpenRouter
OPENAI_API_KEY=sk-or-v1-...
OPENAI_BASE_URL=https://openrouter.ai/api/v1

# Mode batch pour execution automatisee
BATCH_MODE=false

GPU Allocation & Idle Management

GPU Layout (po-2023)

GPU Model VRAM Services
GPU0 RTX 3080Ti 16 GB comfyui-qwen, whisper-api, demucs-api
GPU1 RTX 3090 24 GB vllm-zimage, sd-forge-main, forge-turbo, tts-fishaudio, tts-kokoro, musicgen-api, qwen-asr-api, comfyui-video, whisper-webui

Quantization Settings per Service

Service Hôte Model Quantization Idle VRAM
comfyui-qwen po-2023 Qwen Image Edit 2509 fp16 (fp8 checkpoint) ~6-8 GB
whisper-api po-2023 faster-whisper-large-v3-turbo int8_float16 ~4-6 GB
demucs-api po-2023 htdemucs_ft fp16 ~4 GB
vllm-zimage po-2023 Z-Image-Turbo bfloat16 (fp8 opt via VLLM_QUANTIZATION=fp8) ~5-10 GB
sd-forge-main po-2023 SD Forge SDXL xformers fp16 ~6-10 GB
forge-turbo po-2023 SD Forge Turbo xformers fp16 ~6-10 GB
tts-fishaudio po-2023 FishAudio S2-Pro BnB 4-bit NF4 ~5 GB
tts-kokoro po-2023 Kokoro TTS fp32/fp16 default ~1-2 GB
musicgen-api po-2023 MusicGen-medium fp16 ~10 GB
qwen-asr-api po-2023 Qwen3-ASR-1.7B bfloat16 ~3-4 GB
comfyui-video po-2023 HunyuanVideoWrapper fp16 varies
whisper-webui po-2023 Whisper WebUI unknown varies

Idle Management

Built-in lazy loading (auto-unload after 5 min idle): - whisper-api, musicgen-api, demucs-api, qwen-asr-api — use shared/lazy_model.py with IDLE_TIMEOUT=300

External idle monitor sidecars: - comfyui-qwen, comfyui-video — comfyui_idle_monitor.py (calls /free to unload models). Fail-safe tri-state : si l’état du serveur est indéterminable (API injoignable / auth échouée), le check est sauté — jamais de /free à l’aveugle ; et la queue est relue à l’instant du tir (un prompt démarré entre la mesure et le /free l’annule). Avant ce garde-fou, un token stale post-restart déclenchait /free en boucle toutes les ~67 s, y compris pendant une génération active → crash serveur (exit 0, reboot 7-8 min, incident 2026-08-25). - vllm-zimage, tts-fishaudio — service_idle_monitor.py (stops container) - sd-forge-main — service_idle_monitor.py with HTTP Basic Auth (Caddy reverse proxy)

Wake-on-demand (shared/service_wake.py): a stopped container has no listener, so a notebook hitting an idle-stopped service gets HTTP 502 until a manual docker start + warm-up. service_wake.py is the demand-side counterpart to service_idle_monitor.py — it probes /health and, if down, issues docker start <container> then polls until healthy:

python docker-configurations/services/shared/service_wake.py musicgen-api \
  --health-url http://localhost:8192/health -v
from service_wake import ensure_service_up
if ensure_service_up("musicgen-api", "http://localhost:8192/health"):
    # safe to call POST /v1/generate (transparent warm-up done)
    ...

Notebook-side wrapper (MyIA.AI.Notebooks/GenAI/shared/helpers/genai_service.py): a drop-in for requests.get/post that makes the wake-on-demand transparent at the notebook level — no need to import service_wake or handle 502. It resolves a short name from a registry (12 services: ports + health paths verified), wakes the container if idle-stopped, waits for warm-up, then issues the HTTP request:

from helpers.genai_service import call_service

# Drop-in for requests.get("http://localhost:8192/v1/models")
resp = call_service("musicgen", "/v1/models")

# POST with payload + API key (Bearer derived from MUSICGEN_API_API_KEY)
resp = call_service("tts-kokoro", "/v1/audio/speech", method="POST",
                    json={"text": "bonjour"})

The registry maps short names → (container, port, health_path); an unknown name raises KeyError. For services outside the registry, use call_service_generic() with explicit container/port/health_path. Unit tests (CI-runnable, Docker mocked) lock the wake-then-call contract: shared/helpers/tests/test_genai_service.py.

VRAM economy is preserved: the idle monitor still re-stops the service after idle_timeout. See issue #2982 for the A/B/C/D decision rationale (caller-side wake chosen over reverse-proxy / always-on / native-sleep).

No idle handling (low VRAM or always-on): - tts-kokoro (~1-2 GB), forge-turbo (~6-10 GB), whisper-webui (varies)

Configuration generale

  • API keys : MyIA.AI.Notebooks/GenAI/.env (template : .env.example)
  • C# settings : MyIA.AI.Notebooks/Config/settings.json
  • Docker : docker-configurations/services/comfyui-qwen/.env
Retour au sommet