Alayarenderer Generative World

by reason-machines2384a003145aNo license83 starsListed Oct 8, 2026Updated Oct 8, 2026Repository updated 3 months ago

AI coding agent skill for AlayaRenderer — a generative world rendering framework with inverse rendering (RGB→G-buffers) and game editing (G-buffers+text→stylized video) using fine-tuned video diffusion models.

Instructions onlyDesign & Creative
AI-generated overview

Guides setup and use of AlayaRenderer, a two-stage video framework for inverse rendering and text-stylized game video.

What it does
This skill documents AlayaRenderer, a two-stage generative world rendering framework. The first stage, an inverse renderer, decomposes RGB video into five G-buffer channels (albedo, normal, depth, roughness, metallic); the second stage, game editing, synthesizes stylized RGB video from G-buffers plus a text prompt. It covers installation, model weight downloads, CLI inference parameters, G-buffer directory layout, style prompt examples, multi-GPU runs, and troubleshooting.
When to use it
Use it when setting up or running AlayaRenderer for inverse rendering of video into G-buffers or for text-driven stylization of game footage. It also fits fine-tuning or inference workflows around the Cosmos-Transfer1 and Wan2.1 based models.
Requirements
Requires cloning the AlayaRenderer repository with submodules (DiffSynth-Studio), two separate Conda environments with Python 3.10, NVIDIA CUDA GPUs, and downloaded HuggingFace model weights (Cosmos-Transfer1-DiffusionRenderer 7B and Wan2.1 1.3B). Network access is needed for repository and weight downloads. The skill ships instructions only, no scripts.

AlayaRenderer — Generative World Renderer

Skill by ara.so — Daily 2026 Skills collection.

AlayaRenderer is a two-stage framework for high-quality video rendering:

  1. Inverse Renderer (RGB → G-buffers): Extracts albedo, normal, depth, roughness, and metallic maps from RGB video using a fine-tuned Cosmos-Transfer1-DiffusionRenderer 7B model.
  2. Game Editing (G-buffers + Text → Stylized RGB): Synthesizes photorealistic, stylized RGB video from G-buffer inputs using a fine-tuned Wan2.1 1.3B model via DiffSynth-Studio.

Installation

Clone the Repository

bash
git clone --recurse-submodules https://github.com/ShandaAI/AlayaRenderer.gitcd AlayaRenderer

Important: Use --recurse-submodules — DiffSynth-Studio is a git submodule required for Game Editing.

Two Separate Conda Environments (Recommended)

The two models have conflicting dependencies. Use separate environments:

bash
# Environment 1: Inverse Rendererconda create -n inverse_renderer python=3.10 -yconda activate inverse_renderercd inverse_renderer# Follow inverse_renderer/ instructions for Cosmos-Transfer1 setup
# Environment 2: Game Editingconda create -n game_editing python=3.10 -yconda activate game_editingcd game_editing# Follow DiffSynth-Studio setup instructions

Model Weights

ModelBase ModelSizeHuggingFace Link
Inverse RendererCosmos-Transfer1-DiffusionRenderer 7B~7B paramsBrian9999/world_inverse_renderer
Game EditingWan2.1 1.3B~1.3B paramsBrian9999/stylerenderer

Download and Place Weights

bash
# Inverse Renderer — replace the base checkpointhuggingface-cli download Brian9999/world_inverse_renderer \  --local-dir inverse_renderer/checkpoints/Diffusion_Renderer_Inverse_Cosmos_7B
# Game Editing — place in game_editing models directorymkdir -p game_editing/models/train/Wan2.1-T2V-1.3B_gbufferhuggingface-cli download Brian9999/stylerenderer \  --local-dir game_editing/models/train/Wan2.1-T2V-1.3B_gbuffer

Inverse Renderer Usage

The inverse renderer decomposes an RGB video into 5 G-buffer channels: albedo, normal, depth, roughness, metallic.

Setup

bash
cd inverse_renderer# Follow Cosmos-Transfer1-DiffusionRenderer environment setup# Ensure checkpoint is at:# inverse_renderer/checkpoints/Diffusion_Renderer_Inverse_Cosmos_7B/

Inference

Refer to the inverse_renderer/ subdirectory for the full inference script. The general pattern follows Cosmos-Transfer1-DiffusionRenderer conventions:

python
# inverse_renderer/run_inverse.py (typical pattern)import torchfrom pathlib import Path
# Input: path to RGB videoinput_video = "path/to/rgb_video.mp4"output_dir = "outputs/gbuffers/"
# The model outputs 5 synchronized channels:# - albedo (diffuse color)# - normal (surface orientation)# - depth (scene geometry)# - roughness (surface roughness)# - metallic (metallic property)

Game Editing Usage

Quick Start — CLI Inference

bash
cd game_editing
CUDA_VISIBLE_DEVICES=0 python \    examples/wanvideo/model_inference/inference_gbuffer_caption.py \    --checkpoint models/train/Wan2.1-T2V-1.3B_gbuffer/model.safetensors \    --gpu 0 \    --style snowy_winter \    --prompt "the scene is set in a frozen, snow-covered environment under cold, pale winter light with falling snowflakes, creating a silent and ethereal winter wonderland atmosphere." \    --gbuffer_dir test_dataset \    --save_dir outputs/ \    --num_frames 81 \    --height 480 \    --width 832

CLI Parameters

ParameterDescriptionExample
--checkpointPath to fine-tuned .safetensors weightsmodels/train/Wan2.1-T2V-1.3B_gbuffer/model.safetensors
--gpuGPU device index0
--styleNamed style presetsnowy_winter, rainy, night, sunset
--promptText description of target lighting/atmosphereSee examples below
--gbuffer_dirDirectory containing G-buffer input frames/videotest_dataset
--save_dirOutput directory for rendered videooutputs/
--num_framesNumber of frames to generate (must be 8n+1)81
--heightOutput height in pixels480
--widthOutput width in pixels832

G-buffer Directory Structure

test_dataset/├── albedo/│   ├── frame_0000.png│   ├── frame_0001.png│   └── ...├── normal/│   ├── frame_0000.png│   └── ...├── depth/│   ├── frame_0000.png│   └── ...├── roughness/│   ├── frame_0000.png│   └── ...└── metallic/    ├── frame_0000.png    └── ...

Style Prompt Examples

bash
# Cyberpunk night scene--style night \--prompt "neon-lit urban environment at night with rain-slicked streets reflecting colorful neon signs, creating a cyberpunk noir atmosphere"
# Golden hour / sunset--style sunset \--prompt "warm golden hour lighting with long shadows and a glowing amber sky, soft cinematic atmosphere"
# Rainy urban--style rainy \--prompt "overcast rainy day with wet surfaces, soft diffuse lighting, and atmospheric fog creating a moody cinematic look"
# Fantasy / stylized--style fantasy \--prompt "magical forest environment with bioluminescent plants, ethereal blue-green lighting, and mystical particle effects"
# Foggy morning--style foggy \--prompt "early morning dense fog with soft diffused light creating a mysterious and quiet atmosphere"

Multi-GPU Inference

bash
# Run on specific GPUCUDA_VISIBLE_DEVICES=1 python \    examples/wanvideo/model_inference/inference_gbuffer_caption.py \    --checkpoint models/train/Wan2.1-T2V-1.3B_gbuffer/model.safetensors \    --gpu 1 \    --style rainy \    --prompt "heavy rainfall with dark storm clouds and dramatic lightning in the distance" \    --gbuffer_dir my_gbuffers \    --save_dir outputs/rainy_scene \    --num_frames 81 --height 480 --width 832

Full Pipeline: RGB Video → Stylized Output

bash
# Step 1: Extract G-buffers from RGB video (Inverse Renderer env)conda activate inverse_renderercd inverse_rendererpython run_inverse.py \    --input path/to/gameplay_video.mp4 \    --output_dir ../game_editing/test_dataset/
# Step 2: Apply game editing style (Game Editing env)conda activate game_editingcd ../game_editingCUDA_VISIBLE_DEVICES=0 python \    examples/wanvideo/model_inference/inference_gbuffer_caption.py \    --checkpoint models/train/Wan2.1-T2V-1.3B_gbuffer/model.safetensors \    --gpu 0 \    --style snowy_winter \    --prompt "frozen tundra with blizzard conditions, pale blue-white lighting and drifting snow" \    --gbuffer_dir test_dataset \    --save_dir outputs/final_render \    --num_frames 81 --height 480 --width 832

Online Demos


Dataset Overview

The AlayaRenderer dataset (release pending) features:

  • 4M+ frames at 720p / 30 FPS
  • 6 synchronized channels: RGB + albedo, normal, depth, metallic, roughness
  • 40 hours from Cyberpunk 2077 and Black Myth: Wukong
  • Average clip length: 8 minutes, up to 53 minutes continuous
  • Weather variants: sunny, rainy, foggy, night, sunset
  • Motion blur variant via sub-frame interpolation

Architecture Summary

RGB Video Input      │      ▼┌─────────────────────────────────────┐│  Inverse Renderer                   ││  (Cosmos-Transfer1 7B fine-tuned)   ││  RGB → [albedo, normal, depth,      ││          roughness, metallic]       │└─────────────────┬───────────────────┘                  │  G-buffers                  ▼┌─────────────────────────────────────┐│  Game Editing                       ││  (Wan2.1 1.3B fine-tuned)           ││  G-buffers + Text Prompt            ││  → Stylized RGB Video               │└─────────────────────────────────────┘

Troubleshooting

Submodule not found / DiffSynth-Studio missing

bash
# If cloned without --recurse-submodules:git submodule update --init --recursive

CUDA Out of Memory

  • Reduce --num_frames (try 41 instead of 81)
  • Reduce resolution: --height 320 --width 576
  • Ensure no other processes are using the GPU: CUDA_VISIBLE_DEVICES=0

num_frames must follow 8n+1 pattern

Valid values: 9, 17, 25, 33, 41, 49, 57, 65, 73, 81

bash
# Valid--num_frames 81   # 8*10 + 1 ✓--num_frames 41   # 8*5 + 1  ✓
# Invalid--num_frames 80   # ✗--num_frames 60   # ✗

Checkpoint not found

bash
# Verify checkpoint placementls game_editing/models/train/Wan2.1-T2V-1.3B_gbuffer/model.safetensorsls inverse_renderer/checkpoints/Diffusion_Renderer_Inverse_Cosmos_7B/

Version conflicts between models

Always use the two separate conda environments (inverse_renderer and game_editing). Do not install both models' dependencies in one environment.


Citation

bibtex
@article{huang2026generativeworldrenderer,    title={Generative World Renderer},    author={Zheng-Hui Huang and Zhixiang Wang and Jiaming Tan and Ruihan Yu and Yidan Zhang and Bo Zheng and Yu-Lun Liu and Yung-Yu Chuang and Kaipeng Zhang},    journal={arXiv preprint arXiv:2604.02329},    year={2026}}

Source and attribution

Source:reason-machines/trending-skillsinskills/alayarenderer-generative-worldat commit2384a00

License: No license

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal