ComfyUI Video Pipeline
Orchestrates video generation across three engines, selecting the best one based on requirements and available resources.
Engine Selection
Pipeline 1: Wan 2.2 MoE (Highest Quality)
Image-to-Video
Prerequisites:
wan2.1_i2v_720p_14b_bf16.safetensorsinmodels/diffusion_models/umt5_xxl_fp8_e4m3fn_scaled.safetensorsinmodels/clip/open_clip_vit_h_14.safetensorsinmodels/clip_vision/wan_2.1_vae.safetensorsinmodels/vae/
Settings:
Frame count guide:
VRAM optimization:
- FP8 quantization: halves VRAM with minimal quality loss
- SageAttention: faster attention computation
- Reduce frames if OOM
Text-to-Video
Same as I2V but uses wan2.1_t2v_14b_bf16.safetensors and EmptySD3LatentImage instead of image conditioning.
First+Last Frame Control (Wan 2.2 Exclusive)
Wan 2.2 MoE allows specifying both the first and last frame, enabling precise video planning:
- Generate two hero images with consistent character
- Use first as start frame, second as end frame
- Wan interpolates the motion between them
Pipeline 2: FramePack (Long Videos, Low VRAM)
Key Innovation
VRAM usage is invariant to video length - generates 60-second videos at 30fps on just 6GB VRAM.
How it works:
- Dynamic context compression: 1536 markers for key frames, 192 for transitions
- Bidirectional memory with reverse generation prevents drift
- Frame-by-frame generation with context window
Settings
When to Use
- Videos longer than 10 seconds
- Limited VRAM systems (but RTX 5090 doesn't need this)
- When VRAM is needed for parallel operations
- Batch video generation
Pipeline 3: AnimateDiff V3 (Fast, Controllable)
Strengths
- Motion LoRAs for camera control (pan, zoom, tilt, roll)
- Effect LoRAs (shatter, smoke, explosion, liquid)
- Sliding context window for infinite length
- Very fast with Lightning model (4-8 steps)
Settings
Camera Motion LoRAs
Post-Processing Pipeline
After any video generation:
1. Frame Interpolation (RIFE)
Doubles or quadruples frame count for smoother motion:
Use rife47 or rife49 model.
2. Face Enhancement (if character video)
Apply FaceDetailer to each frame:
- denoise: 0.3-0.4 (lower than image - preserves temporal consistency)
- guide_size: 384 (speed optimization for video)
- detection_model: face_yolov8m.pt
3. Deflicker (if needed)
Reduces temporal inconsistencies between frames.
4. Color Correction
Maintain consistent color grading across frames.
5. Video Combine
Final output via VHS Video Combine:
Talking Head Pipeline
Complete pipeline for character dialogue:
Quality Checklist
Before marking video as complete:
- Character identity consistent across frames
- No flickering or temporal artifacts
- Motion looks natural (not jerky or frozen)
- Face enhancement applied if character video
- Frame rate is smooth (24+ fps for delivery)
- Audio synced (if talking head)
- Resolution matches delivery target
Reference
references/workflows.md- Workflow templates for Wan and AnimateDiffreferences/models.md- Video model download linksreferences/research-log.md- Latest video generation advancesstate/inventory.json- Available video models

