GenAI Services - ComfyUI Image Generation
Services disponibles
| Service | Hôte | Modèle | VRAM | Description |
|---|---|---|---|---|
| Qwen Image Edit | po-2023 | qwen_image_edit_2509 | ~29GB | Edition d’images avec prompts multimodaux |
| Z-Image/Lumina | po-2023 | Lumina-Next-SFT | ~10GB | Generation text-to-image haute qualite |
| MiniMax H3 (Hailuo 3.0) | non déployé (licence UE) | MiniMax-H3 (open-weights) | ~24+ GB | VIDEO omni-modale + audio stereo natif (UE exclue par licence) — voir Architecture MiniMax H3 |
Architecture Qwen (Phase 29)
Workflow ComfyUI pour Qwen Image Edit 2509 :
VAELoader (qwen_image_vae.safetensors, 16 channels)
|
CLIPLoader (qwen_2.5_vl_7b_fp8_scaled.safetensors, type: sd3)
|
UNETLoader (qwen_image_edit_2509_fp8_e4m3fn.safetensors)
|
ModelSamplingAuraFlow (shift=3.0)
|
CFGNorm (strength=1.0)
|
TextEncodeQwenImageEdit (clip, prompt, vae)
|
ConditioningZeroOut (negative)
|
EmptySD3LatentImage (16 channels)
|
KSampler (scheduler=beta, cfg=1.0, sampler=euler)
|
VAEDecode
Points critiques : - VAE 16 canaux (pas SDXL standard) - scheduler=beta obligatoire - cfg=1.0 (pas de CFG classique, utilise CFGNorm) - ModelSamplingAuraFlow avec shift=3.0
Architecture Z-Image/Lumina
Workflow ComfyUI simplifie avec LuminaDiffusersNode :
LuminaDiffusersNode (Alpha-VLLM/Lumina-Next-SFT-diffusers)
|
VAELoader (sdxl_vae.safetensors)
|
VAEDecode
|
SaveImage
Paramètres LuminaDiffusersNode : - model_path: “Alpha-VLLM/Lumina-Next-SFT-diffusers” - num_inference_steps: 20-40 - guidance_scale: 3.0-5.0 - scaling_watershed: 0.3 - proportional_attn: true - max_sequence_length: 256
Note technique (Janvier 2025) : Le node utilise LuminaPipeline (diffusers 0.34+), ancien nom LuminaText2ImgPipeline obsolete.
Architecture MiniMax H3 (Hailuo 3.0) — VIDEO avec audio natif
Statut : INTRINSIC pour l’UE (Tâche 2 #10244). Le modele est open-weights et techniquement auto-hebergeable (ComfyUI workflows publies sur ai-models-lab/minimax-h3), mais la licence MiniMax H3 Community License (datee 2 aout 2026, deposee sur HuggingFace MiniMaxAI/MiniMax-H3) exclut explicitement l’usage, l’hebergement et l’affichage des Outputs sur le territoire de l’Union Europeenne (cf. Art. I.3 Applicable Territory et Art. I.5 Excluded Territories). Ce bloc documente donc l’architecture a titre de reference — pour informer la decision de deploiement — et non pour inviter au deploiement local.
Verdict SOTA pour notre contexte (France / UE, depot public, ecoles partenaires) : INTRINSIC. La voie pedagogiquement equivalente pour « video + audio natif synchronise » en UE est LTX-2 (
02-5-LTX2-Audiovisual, licence permissive, executable). Le notebook02-6-MiniMax-H3-Architecture-Licensingde la serie Video enseigne l’architecture et le raisonnement de conformite (verificateur de juridiction, matrice de decision avec Sora et LTX-2).
Specifications du modele (verifiees firsthand, source officielle minimax.io + GitHub MiniMax-AI/MiniMax-H3 + HuggingFace)
| Caracteristique | Valeur |
|---|---|
| Modalites d’entree | Omni-modale unifiee : texte, image, video, audio |
| Sortie | Video + audio stereo natif (32 kHz, 11 langues) |
| Resolution | jusqu’a 2K |
| Duree | clips de 5 a 15 s |
| Fps | 24 fps |
| Modes de generation | text-to-video · image-to-video · first+last frame · omni-reference (entrees mixtes) |
| Classement | #1 Artificial Analysis video editing au lancement (Elo 1130) |
| VRAM (selon documentation, execution reelle hors UE) | ~24 GB+ (quantization fp8 / GGUF Q4 envisageable) |
| Workflows ComfyUI | hub communautaire ai-models-lab/minimax-h3 (non deploye ici) |
Architecture du pipeline (telle que publiee par la communaute ComfyUI ; non deployee sur le cluster)
MiniMaxH3Loader (checkpoint MultiModal, fp8 ou GGUF Q4)
|
CLIPLoader (H3 text-encoder, ~7B)
|
OmniModalConditioning (text + image + audio refs)
|
MiniMaxH3UNET (denoising + audio decoder couple)
|
AudioVAEDecode (32 kHz stereo, 11 langues)
|
VaeDecode (frames 2K, 24 fps, 5-15 s)
|
SaveVideo + SaveAudio (synchronisation native)
Points critiques (selon documentation communautaire, non valides firsthand sur le cluster) : - VAE multi-modal (pas SDXL standard) — couples video frames + audio waveform - CFG 1.0 recommande pour eviter la degenerescence audio (semblable a Qwen Image Edit) - Scheduler beta recommande pour la convergence des clips longs - Omni-reference : permet d’injecter des refs image+audio simultanement (capacite distinctive vs Wan/SVD)
Comparaison synthetique avec les modeles Video de la serie
| Modele | Format | Audio natif | UE-ok | Open-weights | Notebook |
|---|---|---|---|---|---|
| HunyuanVideo | text-to-video | non | oui | oui | 02-1 (~18 GB VRAM) |
| LTX-Video | text-to-video | non | oui | oui | 02-2 (~8 GB VRAM) |
| Wan 2.1/2.2 | text-to-video | non | oui | oui | 02-3 (~10 GB VRAM) |
| SVD | image-to-video | non | oui | oui | 02-4 (~10 GB VRAM) |
| LTX-2 (Lightricks 22B) | text-to-video | OUI (stereo) | oui | oui | 02-5 (~16-24 GB VRAM) |
| MiniMax H3 | omni-modal | OUI (stereo 11 langues) | NON (UE exclue) | oui (licence geo-restreinte) | 02-6 (descriptif, INTRINSIC) |
| OpenAI Sora 2 | text-to-video | non | oui (API) | non (API fermee) | 04-3 |
Etat reversible — ce qui changerait la decision
Le verdict INTRINSIC pour notre contexte n’est pas un mur infranchissable ; il est conditionne a la licence. Trois evenements le feraient basculer :
- MiniMax deplace l’Excluded Territory (ex. retire l’UE de la liste). Surveillance :
MiniMaxAI/MiniMax-H3releases /minimax.io/news(cf. notebook 02-6 §5). - Licence commerciale UE obtenue (Art. II dernier § : « you are welcome to contact us about obtaining a license »). Demarche provider-side dependant de MiniMax.
- Fork communautaire re-licencie sous licence permissive (hypothese, aucun signal actuel a 2026-08-10).
Re-verification licence 2026-08-15 (firsthand, LICENSE raw HuggingFace, date 2 aout 2026) : UE toujours exclue — “Excluded Territories” means the European Union, the United Kingdom, the Republic of Korea and the United States of America. Aucun changement. Cote service cloud, la voie video-01 (/v1/video_generation) est couverte par le plan (probe 2026-08-10, notebook 04-5b) ; la serie H3 reste plan-gated (400 TokenPlan 2013, notebook 04-5 squelette idempotent).
En attendant, le notebook 02-6 evoluera vers une execution locale reelle des qu’un de ces evenements se materialise — l’architecture ci-dessus est le blueprint ComfyUI a deployer le moment venu.
Architecture Krea 2 turbo + VLM in-graph (03-4)
Workflow rejoué depuis les métadonnées d’un PNG (notebook Image/03-Orchestration/03-4) : une boucle de raisonnement dans le graphe ComfyUI — un VLM examine une image de référence quelconque et rédige lui-même le prompt de character design avant la diffusion.
LoadImage (reference quelconque)
|
ImageScaleToTotalPixels (1.0 MP, lanczos)
|
CLIPLoader (qwen3vl_4b_bf16.safetensors, type: qwen_image) <- sert AUSSI de VLM
|
TextGenerate (system prompt "visual DNA", 11235 chars, temp 0.7)
| -> PreviewAny (recuperation du texte genere via /history)
v
StringConcatenate (trigger words LoRA + prompt VLM)
|
LoraLoader (banjiesock_Krea2, model + clip) UNETLoader (krea2_turbo_int8_convrot)
| |
CLIPTextEncode -> KSampler (er_sde, simple, 8 steps, cfg 1.0)
ConditioningZeroOut (negatif nul, modele distille)
|
VAEDecode (qwen_image_vae.safetensors, 16 canaux) -> SaveImage
Points critiques : - UNet int8 convrot 13,49 Go + VLM 8,9 Go sur RTX 3090 24 Go : offload séquentiel ComfyUI (le VLM travaille puis laisse la place au UNet) — les deux phases tiennent séparément. - TextGenerate est un node natif (comfy_extras/nodes_textgen.py) : il exige un CLIP dont la config expose stop_tokens (familles gemma4 / qwen35) ; qwen_2.5_vl_7b (type qwen_image) NE le supporte PAS (AttributeError: stop_tokens). - cfg=1.0 + 8 steps = modèle distillé (turbo) : négatif via ConditioningZeroOut, pas de prompt négatif factice. - Le texte généré se récupère via un node PreviewAny (output text dans /history) — sans node terminal texte, POST /prompt répond prompt_no_outputs. - Le workflow original (PNG auteur) utilisait gemma4_12b comme VLM et des custom nodes (LoraManager, ResolutionMaster, rgthree, KJNodes) : substitutions documentées dans le notebook (table custom -> natif). - Poids : Comfy-Org/Krea-2 (HF) pour UNet/encodeurs/VAE, Civitai modèle 2937073 pour le LoRA banjiesock (trigger words : banjiesock2style, illustration).
Approches abandonnees
| Approche | Raison abandon |
|---|---|
| Z-Image GGUF | Incompatibilite dimensionnelle (2560 vs 2304) entre RecurrentGemma et Gemma-2 |
| Qwen GGUF | Non teste, prefer les poids fp8 pour qualite |
Scripts de gestion (scripts/genai-stack/)
IMPORTANT pour agents : Utiliser le CLI unifie genai.py au lieu de demarrer des kernels MCP directement.
# CLI unifie - aide
python scripts/genai-stack/genai.py --help
# Gestion services Docker
python scripts/genai-stack/genai.py docker status # Statut services
python scripts/genai-stack/genai.py docker start all # Demarrer tous les services
python scripts/genai-stack/genai.py docker test --remote # Tester endpoints (local + remote)
# Validation stack ComfyUI
python scripts/genai-stack/genai.py validate --full # Validation complete
python scripts/genai-stack/genai.py validate --nunchaku # Test Nunchaku INT4 Lightning
# Validation notebooks
python scripts/genai-stack/genai.py validate --notebooks # Syntaxe notebooks GenAI
python scripts/genai-stack/genai.py notebooks # Execution Papermill
# GPU et modeles
python scripts/genai-stack/genai.py gpu # Verification VRAM
python scripts/genai-stack/genai.py models list-nodes # Custom nodes ComfyUI
python scripts/genai-stack/genai.py models list-checkpoints # Checkpoints disponibles
# Authentification
python scripts/genai-stack/genai.py auth audit # Audit securite tokens
python scripts/genai-stack/genai.py auth sync # Synchroniser tokensMapping notebooks GenAI Image -> services
| Notebooks | Service | Prerequis |
|---|---|---|
| 01-1, 01-3 | OpenAI API (cloud) | OPENAI_API_KEY |
| 01-4, 02-3 | SD Forge | Service local ou myia.io |
| 01-5, 02-1 | ComfyUI Qwen | COMFYUI_AUTH_TOKEN, ~29GB VRAM |
| 02-4 | Z-Image/vLLM | ~10GB VRAM |
| 03-* | Multi-modèles | Tous les services |
| 03-4 | ComfyUI Krea 2 (VLM in-graph) | COMFYUI_AUTH_TOKEN, ~23GB VRAM (UNet int8 13,5 Go + VLM 8,9 Go, offload séquentiel) |
| 04-* | Applications | Variable |
Mapping notebooks Audio -> services
| Notebooks | Service | Prerequis |
|---|---|---|
| Audio/01-1, 01-2 | OpenAI API (TTS/STT) | OPENAI_API_KEY |
| Audio/01-3 | Local (librosa, pydub) | Aucun |
| Audio/01-4 | Whisper local | GPU ~10 GB |
| Audio/01-5 | Kokoro TTS | GPU ~2 GB |
| Audio/02-1 | Chatterbox TTS | GPU ~8 GB |
| Audio/02-2 | XTTS v2 | GPU ~6 GB |
| Audio/02-3 | MusicGen | GPU ~10 GB |
| Audio/02-4 | Demucs v4 | GPU ~4 GB |
| Audio/03-* | Multi-modèles | Mixed |
| Audio/04-11 | Kokoro TTS + FishAudio S2-Pro | GPU ~5 GB (FishAudio BnB 4-bit NF4) + ~2 GB (Kokoro) |
| Audio/04-* | Applications | Mixed |
Mapping notebooks Video -> services
| Notebooks | Service | Prerequis |
|---|---|---|
| Video/01-1 | Local (moviepy, FFmpeg) | FFmpeg installe |
| Video/01-2 | OpenAI GPT-5 | OPENAI_API_KEY |
| Video/01-3 | Qwen2.5-VL local | GPU ~18 GB |
| Video/01-4 | Real-ESRGAN/RIFE | GPU ~4 GB |
| Video/01-5 | AnimateDiff | GPU ~12 GB |
| Video/02-1 | HunyuanVideo | GPU ~18 GB |
| Video/02-2 | LTX-Video | GPU ~8 GB |
| Video/02-3 | Wan 2.1/2.2 | GPU ~10 GB |
| Video/02-4 | SVD | GPU ~10 GB |
| Video/02-5 | LTX-2 (Lightricks 22B) | GPU ~16-24 GB (fp8-cast / GGUF Q4) |
| Video/02-6 | MiniMax H3 (analyse architecture) | 0 VRAM (descriptif, INTRINSIC UE) |
| Video/03-1 | Multi-modeles locaux | GPU ~18 GB |
| Video/03-2 | Pipeline text-to-image-to-video | Mixed |
| Video/03-3 | ComfyUI Video | Docker, nodes video |
| Video/04-1 | Pipeline educatif | Mixed |
| Video/04-2 | Style transfer + music video | Mixed |
| Video/04-3 | Sora 2 API | OPENAI_API_KEY |
| Video/04-4 | Pipeline production complet | Mixed |
Configuration .env GenAI
Fichier : MyIA.AI.Notebooks/GenAI/.env
# Mode local (Docker) vs remote (myia.io)
LOCAL_MODE=false
# ComfyUI
COMFYUI_API_URL=https://qwen-image-edit.myia.io
# Hébergement : po-2023 (conteneur comfyui-qwen, RTX 3090 24 GB, port 8188)
COMFYUI_AUTH_TOKEN=<bearer_token_bcrypt>
# OpenAI via OpenRouter
OPENAI_API_KEY=sk-or-v1-...
OPENAI_BASE_URL=https://openrouter.ai/api/v1
# Mode batch pour execution automatisee
BATCH_MODE=falseGPU Allocation & Idle Management
GPU Layout (po-2023)
| GPU | Model | VRAM | Services |
|---|---|---|---|
| GPU0 | RTX 3080Ti | 16 GB | comfyui-qwen, whisper-api, demucs-api |
| GPU1 | RTX 3090 | 24 GB | vllm-zimage, sd-forge-main, forge-turbo, tts-fishaudio, tts-kokoro, musicgen-api, qwen-asr-api, comfyui-video, whisper-webui |
Quantization Settings per Service
| Service | Hôte | Model | Quantization | Idle VRAM |
|---|---|---|---|---|
| comfyui-qwen | po-2023 | Qwen Image Edit 2509 | fp16 (fp8 checkpoint) | ~6-8 GB |
| whisper-api | po-2023 | faster-whisper-large-v3-turbo | int8_float16 | ~4-6 GB |
| demucs-api | po-2023 | htdemucs_ft | fp16 | ~4 GB |
| vllm-zimage | po-2023 | Z-Image-Turbo | bfloat16 (fp8 opt via VLLM_QUANTIZATION=fp8) |
~5-10 GB |
| sd-forge-main | po-2023 | SD Forge SDXL | xformers fp16 | ~6-10 GB |
| forge-turbo | po-2023 | SD Forge Turbo | xformers fp16 | ~6-10 GB |
| tts-fishaudio | po-2023 | FishAudio S2-Pro | BnB 4-bit NF4 | ~5 GB |
| tts-kokoro | po-2023 | Kokoro TTS | fp32/fp16 default | ~1-2 GB |
| musicgen-api | po-2023 | MusicGen-medium | fp16 | ~10 GB |
| qwen-asr-api | po-2023 | Qwen3-ASR-1.7B | bfloat16 | ~3-4 GB |
| comfyui-video | po-2023 | HunyuanVideoWrapper | fp16 | varies |
| whisper-webui | po-2023 | Whisper WebUI | unknown | varies |
Idle Management
Built-in lazy loading (auto-unload after 5 min idle): - whisper-api, musicgen-api, demucs-api, qwen-asr-api — use shared/lazy_model.py with IDLE_TIMEOUT=300
External idle monitor sidecars: - comfyui-qwen, comfyui-video — comfyui_idle_monitor.py (calls /free to unload models). Fail-safe tri-state : si l’état du serveur est indéterminable (API injoignable / auth échouée), le check est sauté — jamais de /free à l’aveugle ; et la queue est relue à l’instant du tir (un prompt démarré entre la mesure et le /free l’annule). Avant ce garde-fou, un token stale post-restart déclenchait /free en boucle toutes les ~67 s, y compris pendant une génération active → crash serveur (exit 0, reboot 7-8 min, incident 2026-08-25). - vllm-zimage, tts-fishaudio — service_idle_monitor.py (stops container) - sd-forge-main — service_idle_monitor.py with HTTP Basic Auth (Caddy reverse proxy)
Wake-on-demand (shared/service_wake.py): a stopped container has no listener, so a notebook hitting an idle-stopped service gets HTTP 502 until a manual docker start + warm-up. service_wake.py is the demand-side counterpart to service_idle_monitor.py — it probes /health and, if down, issues docker start <container> then polls until healthy:
python docker-configurations/services/shared/service_wake.py musicgen-api \
--health-url http://localhost:8192/health -vfrom service_wake import ensure_service_up
if ensure_service_up("musicgen-api", "http://localhost:8192/health"):
# safe to call POST /v1/generate (transparent warm-up done)
...Notebook-side wrapper (MyIA.AI.Notebooks/GenAI/shared/helpers/genai_service.py): a drop-in for requests.get/post that makes the wake-on-demand transparent at the notebook level — no need to import service_wake or handle 502. It resolves a short name from a registry (12 services: ports + health paths verified), wakes the container if idle-stopped, waits for warm-up, then issues the HTTP request:
from helpers.genai_service import call_service
# Drop-in for requests.get("http://localhost:8192/v1/models")
resp = call_service("musicgen", "/v1/models")
# POST with payload + API key (Bearer derived from MUSICGEN_API_API_KEY)
resp = call_service("tts-kokoro", "/v1/audio/speech", method="POST",
json={"text": "bonjour"})The registry maps short names → (container, port, health_path); an unknown name raises KeyError. For services outside the registry, use call_service_generic() with explicit container/port/health_path. Unit tests (CI-runnable, Docker mocked) lock the wake-then-call contract: shared/helpers/tests/test_genai_service.py.
VRAM economy is preserved: the idle monitor still re-stops the service after idle_timeout. See issue #2982 for the A/B/C/D decision rationale (caller-side wake chosen over reverse-proxy / always-on / native-sleep).
No idle handling (low VRAM or always-on): - tts-kokoro (~1-2 GB), forge-turbo (~6-10 GB), whisper-webui (varies)
Configuration generale
- API keys :
MyIA.AI.Notebooks/GenAI/.env(template :.env.example) - C# settings :
MyIA.AI.Notebooks/Config/settings.json - Docker :
docker-configurations/services/comfyui-qwen/.env