v5 release with triple architecture support and prompt enhancer
This commit is contained in:
6
.gitignore
vendored
6
.gitignore
vendored
@@ -21,6 +21,7 @@
|
||||
*.html
|
||||
*.pdf
|
||||
*.whl
|
||||
*.exe
|
||||
cache
|
||||
__pycache__/
|
||||
storage/
|
||||
@@ -34,3 +35,8 @@ Wan2.1-T2V-14B/
|
||||
Wan2.1-T2V-1.3B/
|
||||
Wan2.1-I2V-14B-480P/
|
||||
Wan2.1-I2V-14B-720P/
|
||||
outputs/
|
||||
gradio_outputs/
|
||||
ckpts/
|
||||
loras/
|
||||
loras_i2v/
|
||||
|
||||
73
README.md
73
README.md
@@ -2,14 +2,35 @@
|
||||
|
||||
-----
|
||||
<p align="center">
|
||||
<b>Wan2.1 GP by DeepBeepMeep based on Wan2.1's Alibaba: Open and Advanced Large-Scale Video Generative Models for the GPU Poor</b>
|
||||
<b>WanGP by DeepBeepMeep : The best Open Source Video Generative Models Accessible to the GPU Poor</b>
|
||||
</p>
|
||||
|
||||
**NEW Discord Server to get Help from Other Users and show your Best Videos:** https://discord.gg/g7efUW9jGV
|
||||
WanGP supports the Wan (and derived models), Hunyuan Video and LTV Video models with:
|
||||
- Low VRAM requirements (as low as 6 GB of VRAM is sufficient for certain models)
|
||||
- Support for old GPUs (RTX 10XX, 20xx, ...)
|
||||
- Very Fast on the latest GPUs
|
||||
- Easy to use Full Web based interface
|
||||
- Auto download of the required model adapted to your specific architecture
|
||||
- Tools integrated to facilitate Video Generation : Mask Editor, Prompt Enhancer, Temporal and Spatial Generation
|
||||
- Loras Support to customize each model
|
||||
- Queuing system : make your shopping list of videos to generate and come back later
|
||||
|
||||
|
||||
**Discord Server to get Help from Other Users and show your Best Videos:** https://discord.gg/g7efUW9jGV
|
||||
|
||||
|
||||
|
||||
## 🔥 Latest News!!
|
||||
* May 17 2025: 👋 Wan 2.1GP v5.0 : One App to Rule Them All !\
|
||||
Added support for the other great open source architectures:
|
||||
- Hunyuan Video : text 2 video (one of the best, if not the best t2v) ,image 2 video and the recently released Hunyuan Custom (very good identify preservation when injecting a person into a video)
|
||||
- LTX Video 13B (released last week): very long video support and fast 720p generation.Wan GP version has been greatly optimzed and reduced VRAM requirements by 4 !
|
||||
|
||||
Also:
|
||||
- Added supported for the best Control Video Model, released 2 days ago : Vace 14B
|
||||
- New Integrated prompt enhancer to increase the quality of the generated videos
|
||||
You will need one more *pip install -r requirements.txt*
|
||||
|
||||
* May 5 2025: 👋 Wan 2.1GP v4.5: FantasySpeaking model, you can animate a talking head using a voice track. This works not only on people but also on objects. Also better seamless transitions between Vace sliding windows for very long videos (see recommended settings). New high quality processing features (mixed 16/32 bits calculation and 32 bitsVAE)
|
||||
* April 27 2025: 👋 Wan 2.1GP v4.4: Phantom model support, very good model to transfer people or objects into video, works quite well at 720p and with the number of steps > 30
|
||||
* April 25 2025: 👋 Wan 2.1GP v4.3: Added preview mode and support for Sky Reels v2 Diffusion Forcing for high quality "infinite length videos" (see Window Sliding section below).Note that Skyreel uses causal attention that is only supported by Sdpa attention so even if chose an other type of attention, some of the processes will use Sdpa attention.
|
||||
@@ -71,30 +92,6 @@ If you upgrade you will need to do a 'pip install -r requirements.txt' again.
|
||||
* Feb 27, 2025: 👋 Wan2.1 has been integrated into [ComfyUI](https://comfyanonymous.github.io/ComfyUI_examples/wan/). Enjoy!
|
||||
|
||||
|
||||
## Features
|
||||
*GPU Poor version by **DeepBeepMeep**. This great video generator can now run smoothly on any GPU.*
|
||||
|
||||
This version has the following improvements over the original Alibaba model:
|
||||
- Reduce greatly the RAM requirements and VRAM requirements
|
||||
- Much faster thanks to compilation and fast loading / unloading
|
||||
- Multiple profiles in order to able to run the model at a decent speed on a low end consumer config (32 GB of RAM and 12 VRAM) and to run it at a very good speed on a high end consumer config (48 GB of RAM and 24 GB of VRAM)
|
||||
- Autodownloading of the needed model files
|
||||
- Improved gradio interface with progression bar and more options
|
||||
- Multiples prompts / multiple generations per prompt
|
||||
- Support multiple pretrained Loras with 32 GB of RAM or less
|
||||
- Much simpler installation
|
||||
|
||||
|
||||
This fork by DeepBeepMeep is an integration of the mmpg module on the original model
|
||||
|
||||
It is an illustration on how one can set up on an existing model some fast and properly working CPU offloading with changing only a few lines of code in the core model.
|
||||
|
||||
For more information on how to use the mmpg module, please go to: https://github.com/deepbeepmeep/mmgp
|
||||
|
||||
You will find the original Wan2.1 Video repository here: https://github.com/Wan-Video/Wan2.1
|
||||
|
||||
|
||||
|
||||
|
||||
## Installation Guide for Linux and Windows for GPUs up to RTX40xx
|
||||
|
||||
@@ -182,11 +179,11 @@ To run the text to video generator (in Low VRAM mode):
|
||||
```bash
|
||||
python wgp.py
|
||||
#or
|
||||
python wgp.py --t2v #launch the default text 2 video model
|
||||
python wgp.py --t2v #launch the default Wan text 2 video model
|
||||
#or
|
||||
python wgp.py --t2v-14B #for the 14B model
|
||||
python wgp.py --t2v-14B #for the Wan 14B model
|
||||
#or
|
||||
python wgp.py --t2v-1-3B #for the 1.3B model
|
||||
python wgp.py --t2v-1-3B #for the Wan 1.3B model
|
||||
|
||||
```
|
||||
|
||||
@@ -227,17 +224,23 @@ python wgp.py --attention sdpa
|
||||
### Loras support
|
||||
|
||||
|
||||
Every lora stored in the subfoler 'loras' for t2v and 'loras_i2v' will be automatically loaded. You will be then able to activate / desactive any of them when running the application by selecting them in the area below "Activated Loras" .
|
||||
Lora for the Wan models are stored in the subfoler 'loras' for t2v and 'loras_i2v'. You will be then able to activate / desactive any of them when running the application by selecting them in the Advanced Tab "Loras" .
|
||||
|
||||
If you want to manage in different areas Loras for the 1.3B model and the 14B as they are not compatible, just create the following subfolders:
|
||||
If you want to manage in different areas Loras for the 1.3B model and the 14B of Wan t2v models (as they are not compatible), just create the following subfolders:
|
||||
- loras/1.3B
|
||||
- loras/14B
|
||||
|
||||
You can also put all the loras in the same place by launching the app with following command line (*path* is a path to shared loras directory):
|
||||
You can also put all the loras in the same place by launching the app with the following command line (*path* is a path to shared loras directory):
|
||||
```
|
||||
python wgp.exe --lora-dir path --lora-dir-i2v path
|
||||
```
|
||||
|
||||
Hunyuan Video and LTX Video models have also their own loras subfolders:
|
||||
-loras_hunyuan
|
||||
-loras_hunyuan_i2v
|
||||
-loras_ltxv
|
||||
|
||||
|
||||
For each activated Lora, you may specify a *multiplier* that is one float number that corresponds to its weight (default is 1.0) .The multipliers for each Lora should be separated by a space character or a carriage return. For instance:\
|
||||
*1.2 0.8* means that the first lora will have a 1.2 multiplier and the second one will have 0.8.
|
||||
|
||||
@@ -342,7 +345,11 @@ Experimental: if your prompt is broken into multiple lines (each line separated
|
||||
--i2v-1-3B : launch the Fun InP 1.3B model image to video generator\
|
||||
--vace : launch the Vace ControlNet 1.3B model image to video generator\
|
||||
--quantize-transformer bool: (default True) : enable / disable on the fly transformer quantization\
|
||||
--lora-dir path : Path of directory that contains Loras in diffusers / safetensor format\
|
||||
--lora-dir path : Path of directory that contains Wan t2v Loras\
|
||||
--lora-dir-i2v path : Path of directory that contains Wan i2v Loras\
|
||||
--lora-dir-hunyuan path : Path of directory that contains Hunyuan t2v Loras\
|
||||
--lora-dir-hunyuan-i2v path : Path of directory that contains Hunyuan i2v Loras\
|
||||
--lora-dir-ltxv path : Path of directory that contains LTX Video Loras\
|
||||
--lora-preset preset : name of preset gile (without the extension) to preload
|
||||
--verbose level : default (1) : level of information between 0 and 2\
|
||||
--server-port portno : default (7860) : Gradio port no\
|
||||
|
||||
0
hyvideo/__init__.py
Normal file
0
hyvideo/__init__.py
Normal file
534
hyvideo/config.py
Normal file
534
hyvideo/config.py
Normal file
@@ -0,0 +1,534 @@
|
||||
import argparse
|
||||
from .constants import *
|
||||
import re
|
||||
from .modules.models import HUNYUAN_VIDEO_CONFIG
|
||||
|
||||
|
||||
def parse_args(namespace=None):
|
||||
parser = argparse.ArgumentParser(description="HunyuanVideo inference script")
|
||||
|
||||
parser = add_network_args(parser)
|
||||
parser = add_extra_models_args(parser)
|
||||
parser = add_denoise_schedule_args(parser)
|
||||
parser = add_inference_args(parser)
|
||||
parser = add_parallel_args(parser)
|
||||
|
||||
args = parser.parse_args(namespace=namespace)
|
||||
args = sanity_check_args(args)
|
||||
|
||||
return args
|
||||
|
||||
|
||||
def add_network_args(parser: argparse.ArgumentParser):
|
||||
group = parser.add_argument_group(title="HunyuanVideo network args")
|
||||
|
||||
|
||||
group.add_argument(
|
||||
"--quantize-transformer",
|
||||
action="store_true",
|
||||
help="On the fly 'transformer' quantization"
|
||||
)
|
||||
|
||||
|
||||
group.add_argument(
|
||||
"--lora-dir-i2v",
|
||||
type=str,
|
||||
default="loras_i2v",
|
||||
help="Path to a directory that contains Loras for i2v"
|
||||
)
|
||||
|
||||
|
||||
group.add_argument(
|
||||
"--lora-dir",
|
||||
type=str,
|
||||
default="",
|
||||
help="Path to a directory that contains Loras"
|
||||
)
|
||||
|
||||
|
||||
group.add_argument(
|
||||
"--lora-preset",
|
||||
type=str,
|
||||
default="",
|
||||
help="Lora preset to preload"
|
||||
)
|
||||
|
||||
# group.add_argument(
|
||||
# "--lora-preset-i2v",
|
||||
# type=str,
|
||||
# default="",
|
||||
# help="Lora preset to preload for i2v"
|
||||
# )
|
||||
|
||||
group.add_argument(
|
||||
"--profile",
|
||||
type=str,
|
||||
default=-1,
|
||||
help="Profile No"
|
||||
)
|
||||
|
||||
group.add_argument(
|
||||
"--verbose",
|
||||
type=str,
|
||||
default=1,
|
||||
help="Verbose level"
|
||||
)
|
||||
|
||||
group.add_argument(
|
||||
"--server-port",
|
||||
type=str,
|
||||
default=0,
|
||||
help="Server port"
|
||||
)
|
||||
|
||||
group.add_argument(
|
||||
"--server-name",
|
||||
type=str,
|
||||
default="",
|
||||
help="Server name"
|
||||
)
|
||||
|
||||
group.add_argument(
|
||||
"--open-browser",
|
||||
action="store_true",
|
||||
help="open browser"
|
||||
)
|
||||
|
||||
group.add_argument(
|
||||
"--t2v",
|
||||
action="store_true",
|
||||
help="text to video mode"
|
||||
)
|
||||
|
||||
group.add_argument(
|
||||
"--i2v",
|
||||
action="store_true",
|
||||
help="image to video mode"
|
||||
)
|
||||
|
||||
group.add_argument(
|
||||
"--compile",
|
||||
action="store_true",
|
||||
help="Enable pytorch compilation"
|
||||
)
|
||||
|
||||
group.add_argument(
|
||||
"--fast",
|
||||
action="store_true",
|
||||
help="use Fast HunyuanVideo model"
|
||||
)
|
||||
|
||||
group.add_argument(
|
||||
"--fastest",
|
||||
action="store_true",
|
||||
help="activate the best config"
|
||||
)
|
||||
|
||||
group.add_argument(
|
||||
"--attention",
|
||||
type=str,
|
||||
default="",
|
||||
help="attention mode"
|
||||
)
|
||||
|
||||
group.add_argument(
|
||||
"--vae-config",
|
||||
type=str,
|
||||
default="",
|
||||
help="vae config mode"
|
||||
)
|
||||
|
||||
parser.add_argument(
|
||||
"--share",
|
||||
action="store_true",
|
||||
help="Create a shared URL to access webserver remotely"
|
||||
)
|
||||
|
||||
parser.add_argument(
|
||||
"--lock-config",
|
||||
action="store_true",
|
||||
help="Prevent modifying the configuration from the web interface"
|
||||
)
|
||||
|
||||
parser.add_argument(
|
||||
"--preload",
|
||||
type=str,
|
||||
default="0",
|
||||
help="Megabytes of the diffusion model to preload in VRAM"
|
||||
)
|
||||
|
||||
parser.add_argument(
|
||||
"--multiple-images",
|
||||
action="store_true",
|
||||
help="Allow inputting multiple images with image to video"
|
||||
)
|
||||
|
||||
|
||||
# Main model
|
||||
group.add_argument(
|
||||
"--model",
|
||||
type=str,
|
||||
choices=list(HUNYUAN_VIDEO_CONFIG.keys()),
|
||||
default="HYVideo-T/2-cfgdistill",
|
||||
)
|
||||
group.add_argument(
|
||||
"--latent-channels",
|
||||
type=str,
|
||||
default=16,
|
||||
help="Number of latent channels of DiT. If None, it will be determined by `vae`. If provided, "
|
||||
"it still needs to match the latent channels of the VAE model.",
|
||||
)
|
||||
group.add_argument(
|
||||
"--precision",
|
||||
type=str,
|
||||
default="bf16",
|
||||
choices=PRECISIONS,
|
||||
help="Precision mode. Options: fp32, fp16, bf16. Applied to the backbone model and optimizer.",
|
||||
)
|
||||
|
||||
# RoPE
|
||||
group.add_argument(
|
||||
"--rope-theta", type=int, default=256, help="Theta used in RoPE."
|
||||
)
|
||||
return parser
|
||||
|
||||
|
||||
def add_extra_models_args(parser: argparse.ArgumentParser):
|
||||
group = parser.add_argument_group(
|
||||
title="Extra models args, including vae, text encoders and tokenizers)"
|
||||
)
|
||||
|
||||
# - VAE
|
||||
group.add_argument(
|
||||
"--vae",
|
||||
type=str,
|
||||
default="884-16c-hy",
|
||||
choices=list(VAE_PATH),
|
||||
help="Name of the VAE model.",
|
||||
)
|
||||
group.add_argument(
|
||||
"--vae-precision",
|
||||
type=str,
|
||||
default="fp16",
|
||||
choices=PRECISIONS,
|
||||
help="Precision mode for the VAE model.",
|
||||
)
|
||||
group.add_argument(
|
||||
"--vae-tiling",
|
||||
action="store_true",
|
||||
help="Enable tiling for the VAE model to save GPU memory.",
|
||||
)
|
||||
group.set_defaults(vae_tiling=True)
|
||||
|
||||
group.add_argument(
|
||||
"--text-encoder",
|
||||
type=str,
|
||||
default="llm",
|
||||
choices=list(TEXT_ENCODER_PATH),
|
||||
help="Name of the text encoder model.",
|
||||
)
|
||||
group.add_argument(
|
||||
"--text-encoder-precision",
|
||||
type=str,
|
||||
default="fp16",
|
||||
choices=PRECISIONS,
|
||||
help="Precision mode for the text encoder model.",
|
||||
)
|
||||
group.add_argument(
|
||||
"--text-states-dim",
|
||||
type=int,
|
||||
default=4096,
|
||||
help="Dimension of the text encoder hidden states.",
|
||||
)
|
||||
group.add_argument(
|
||||
"--text-len", type=int, default=256, help="Maximum length of the text input."
|
||||
)
|
||||
group.add_argument(
|
||||
"--tokenizer",
|
||||
type=str,
|
||||
default="llm",
|
||||
choices=list(TOKENIZER_PATH),
|
||||
help="Name of the tokenizer model.",
|
||||
)
|
||||
group.add_argument(
|
||||
"--prompt-template",
|
||||
type=str,
|
||||
default="dit-llm-encode",
|
||||
choices=PROMPT_TEMPLATE,
|
||||
help="Image prompt template for the decoder-only text encoder model.",
|
||||
)
|
||||
group.add_argument(
|
||||
"--prompt-template-video",
|
||||
type=str,
|
||||
default="dit-llm-encode-video",
|
||||
choices=PROMPT_TEMPLATE,
|
||||
help="Video prompt template for the decoder-only text encoder model.",
|
||||
)
|
||||
group.add_argument(
|
||||
"--hidden-state-skip-layer",
|
||||
type=int,
|
||||
default=2,
|
||||
help="Skip layer for hidden states.",
|
||||
)
|
||||
group.add_argument(
|
||||
"--apply-final-norm",
|
||||
action="store_true",
|
||||
help="Apply final normalization to the used text encoder hidden states.",
|
||||
)
|
||||
|
||||
# - CLIP
|
||||
group.add_argument(
|
||||
"--text-encoder-2",
|
||||
type=str,
|
||||
default="clipL",
|
||||
choices=list(TEXT_ENCODER_PATH),
|
||||
help="Name of the second text encoder model.",
|
||||
)
|
||||
group.add_argument(
|
||||
"--text-encoder-precision-2",
|
||||
type=str,
|
||||
default="fp16",
|
||||
choices=PRECISIONS,
|
||||
help="Precision mode for the second text encoder model.",
|
||||
)
|
||||
group.add_argument(
|
||||
"--text-states-dim-2",
|
||||
type=int,
|
||||
default=768,
|
||||
help="Dimension of the second text encoder hidden states.",
|
||||
)
|
||||
group.add_argument(
|
||||
"--tokenizer-2",
|
||||
type=str,
|
||||
default="clipL",
|
||||
choices=list(TOKENIZER_PATH),
|
||||
help="Name of the second tokenizer model.",
|
||||
)
|
||||
group.add_argument(
|
||||
"--text-len-2",
|
||||
type=int,
|
||||
default=77,
|
||||
help="Maximum length of the second text input.",
|
||||
)
|
||||
|
||||
return parser
|
||||
|
||||
|
||||
def add_denoise_schedule_args(parser: argparse.ArgumentParser):
|
||||
group = parser.add_argument_group(title="Denoise schedule args")
|
||||
|
||||
group.add_argument(
|
||||
"--denoise-type",
|
||||
type=str,
|
||||
default="flow",
|
||||
help="Denoise type for noised inputs.",
|
||||
)
|
||||
|
||||
# Flow Matching
|
||||
group.add_argument(
|
||||
"--flow-shift",
|
||||
type=float,
|
||||
default=7.0,
|
||||
help="Shift factor for flow matching schedulers.",
|
||||
)
|
||||
group.add_argument(
|
||||
"--flow-reverse",
|
||||
action="store_true",
|
||||
help="If reverse, learning/sampling from t=1 -> t=0.",
|
||||
)
|
||||
group.add_argument(
|
||||
"--flow-solver",
|
||||
type=str,
|
||||
default="euler",
|
||||
help="Solver for flow matching.",
|
||||
)
|
||||
group.add_argument(
|
||||
"--use-linear-quadratic-schedule",
|
||||
action="store_true",
|
||||
help="Use linear quadratic schedule for flow matching."
|
||||
"Following MovieGen (https://ai.meta.com/static-resource/movie-gen-research-paper)",
|
||||
)
|
||||
group.add_argument(
|
||||
"--linear-schedule-end",
|
||||
type=int,
|
||||
default=25,
|
||||
help="End step for linear quadratic schedule for flow matching.",
|
||||
)
|
||||
|
||||
return parser
|
||||
|
||||
|
||||
def add_inference_args(parser: argparse.ArgumentParser):
|
||||
group = parser.add_argument_group(title="Inference args")
|
||||
|
||||
# ======================== Model loads ========================
|
||||
group.add_argument(
|
||||
"--model-base",
|
||||
type=str,
|
||||
default="ckpts",
|
||||
help="Root path of all the models, including t2v models and extra models.",
|
||||
)
|
||||
group.add_argument(
|
||||
"--dit-weight",
|
||||
type=str,
|
||||
default="ckpts/hunyuan-video-t2v-720p/transformers/mp_rank_00_model_states.pt",
|
||||
help="Path to the HunyuanVideo model. If None, search the model in the args.model_root."
|
||||
"1. If it is a file, load the model directly."
|
||||
"2. If it is a directory, search the model in the directory. Support two types of models: "
|
||||
"1) named `pytorch_model_*.pt`"
|
||||
"2) named `*_model_states.pt`, where * can be `mp_rank_00`.",
|
||||
)
|
||||
group.add_argument(
|
||||
"--model-resolution",
|
||||
type=str,
|
||||
default="540p",
|
||||
choices=["540p", "720p"],
|
||||
help="Root path of all the models, including t2v models and extra models.",
|
||||
)
|
||||
group.add_argument(
|
||||
"--load-key",
|
||||
type=str,
|
||||
default="module",
|
||||
help="Key to load the model states. 'module' for the main model, 'ema' for the EMA model.",
|
||||
)
|
||||
group.add_argument(
|
||||
"--use-cpu-offload",
|
||||
action="store_true",
|
||||
help="Use CPU offload for the model load.",
|
||||
)
|
||||
|
||||
# ======================== Inference general setting ========================
|
||||
group.add_argument(
|
||||
"--batch-size",
|
||||
type=int,
|
||||
default=1,
|
||||
help="Batch size for inference and evaluation.",
|
||||
)
|
||||
group.add_argument(
|
||||
"--infer-steps",
|
||||
type=int,
|
||||
default=50,
|
||||
help="Number of denoising steps for inference.",
|
||||
)
|
||||
group.add_argument(
|
||||
"--disable-autocast",
|
||||
action="store_true",
|
||||
help="Disable autocast for denoising loop and vae decoding in pipeline sampling.",
|
||||
)
|
||||
group.add_argument(
|
||||
"--save-path",
|
||||
type=str,
|
||||
default="./results",
|
||||
help="Path to save the generated samples.",
|
||||
)
|
||||
group.add_argument(
|
||||
"--save-path-suffix",
|
||||
type=str,
|
||||
default="",
|
||||
help="Suffix for the directory of saved samples.",
|
||||
)
|
||||
group.add_argument(
|
||||
"--name-suffix",
|
||||
type=str,
|
||||
default="",
|
||||
help="Suffix for the names of saved samples.",
|
||||
)
|
||||
group.add_argument(
|
||||
"--num-videos",
|
||||
type=int,
|
||||
default=1,
|
||||
help="Number of videos to generate for each prompt.",
|
||||
)
|
||||
# ---sample size---
|
||||
group.add_argument(
|
||||
"--video-size",
|
||||
type=int,
|
||||
nargs="+",
|
||||
default=(720, 1280),
|
||||
help="Video size for training. If a single value is provided, it will be used for both height "
|
||||
"and width. If two values are provided, they will be used for height and width "
|
||||
"respectively.",
|
||||
)
|
||||
group.add_argument(
|
||||
"--video-length",
|
||||
type=int,
|
||||
default=129,
|
||||
help="How many frames to sample from a video. if using 3d vae, the number should be 4n+1",
|
||||
)
|
||||
# --- prompt ---
|
||||
group.add_argument(
|
||||
"--prompt",
|
||||
type=str,
|
||||
default=None,
|
||||
help="Prompt for sampling during evaluation.",
|
||||
)
|
||||
group.add_argument(
|
||||
"--seed-type",
|
||||
type=str,
|
||||
default="auto",
|
||||
choices=["file", "random", "fixed", "auto"],
|
||||
help="Seed type for evaluation. If file, use the seed from the CSV file. If random, generate a "
|
||||
"random seed. If fixed, use the fixed seed given by `--seed`. If auto, `csv` will use the "
|
||||
"seed column if available, otherwise use the fixed `seed` value. `prompt` will use the "
|
||||
"fixed `seed` value.",
|
||||
)
|
||||
group.add_argument("--seed", type=int, default=None, help="Seed for evaluation.")
|
||||
|
||||
# Classifier-Free Guidance
|
||||
group.add_argument(
|
||||
"--neg-prompt", type=str, default=None, help="Negative prompt for sampling."
|
||||
)
|
||||
group.add_argument(
|
||||
"--cfg-scale", type=float, default=1.0, help="Classifier free guidance scale."
|
||||
)
|
||||
group.add_argument(
|
||||
"--embedded-cfg-scale",
|
||||
type=float,
|
||||
default=6.0,
|
||||
help="Embeded classifier free guidance scale.",
|
||||
)
|
||||
|
||||
group.add_argument(
|
||||
"--reproduce",
|
||||
action="store_true",
|
||||
help="Enable reproducibility by setting random seeds and deterministic algorithms.",
|
||||
)
|
||||
|
||||
return parser
|
||||
|
||||
|
||||
def add_parallel_args(parser: argparse.ArgumentParser):
|
||||
group = parser.add_argument_group(title="Parallel args")
|
||||
|
||||
# ======================== Model loads ========================
|
||||
group.add_argument(
|
||||
"--ulysses-degree",
|
||||
type=int,
|
||||
default=1,
|
||||
help="Ulysses degree.",
|
||||
)
|
||||
group.add_argument(
|
||||
"--ring-degree",
|
||||
type=int,
|
||||
default=1,
|
||||
help="Ulysses degree.",
|
||||
)
|
||||
|
||||
return parser
|
||||
|
||||
|
||||
def sanity_check_args(args):
|
||||
# VAE channels
|
||||
vae_pattern = r"\d{2,3}-\d{1,2}c-\w+"
|
||||
if not re.match(vae_pattern, args.vae):
|
||||
raise ValueError(
|
||||
f"Invalid VAE model: {args.vae}. Must be in the format of '{vae_pattern}'."
|
||||
)
|
||||
vae_channels = int(args.vae.split("-")[1][:-1])
|
||||
if args.latent_channels is None:
|
||||
args.latent_channels = vae_channels
|
||||
if vae_channels != args.latent_channels:
|
||||
raise ValueError(
|
||||
f"Latent channels ({args.latent_channels}) must match the VAE channels ({vae_channels})."
|
||||
)
|
||||
return args
|
||||
164
hyvideo/constants.py
Normal file
164
hyvideo/constants.py
Normal file
@@ -0,0 +1,164 @@
|
||||
import os
|
||||
import torch
|
||||
|
||||
__all__ = [
|
||||
"C_SCALE",
|
||||
"PROMPT_TEMPLATE",
|
||||
"MODEL_BASE",
|
||||
"PRECISIONS",
|
||||
"NORMALIZATION_TYPE",
|
||||
"ACTIVATION_TYPE",
|
||||
"VAE_PATH",
|
||||
"TEXT_ENCODER_PATH",
|
||||
"TOKENIZER_PATH",
|
||||
"TEXT_PROJECTION",
|
||||
"DATA_TYPE",
|
||||
"NEGATIVE_PROMPT",
|
||||
"NEGATIVE_PROMPT_I2V",
|
||||
"FLOW_PATH_TYPE",
|
||||
"FLOW_PREDICT_TYPE",
|
||||
"FLOW_LOSS_WEIGHT",
|
||||
"FLOW_SNR_TYPE",
|
||||
"FLOW_SOLVER",
|
||||
]
|
||||
|
||||
PRECISION_TO_TYPE = {
|
||||
'fp32': torch.float32,
|
||||
'fp16': torch.float16,
|
||||
'bf16': torch.bfloat16,
|
||||
}
|
||||
|
||||
# =================== Constant Values =====================
|
||||
# Computation scale factor, 1P = 1_000_000_000_000_000. Tensorboard will display the value in PetaFLOPS to avoid
|
||||
# overflow error when tensorboard logging values.
|
||||
C_SCALE = 1_000_000_000_000_000
|
||||
|
||||
# When using decoder-only models, we must provide a prompt template to instruct the text encoder
|
||||
# on how to generate the text.
|
||||
# --------------------------------------------------------------------
|
||||
PROMPT_TEMPLATE_ENCODE = (
|
||||
"<|start_header_id|>system<|end_header_id|>\n\nDescribe the image by detailing the color, shape, size, texture, "
|
||||
"quantity, text, spatial relationships of the objects and background:<|eot_id|>"
|
||||
"<|start_header_id|>user<|end_header_id|>\n\n{}<|eot_id|>"
|
||||
)
|
||||
PROMPT_TEMPLATE_ENCODE_VIDEO = (
|
||||
"<|start_header_id|>system<|end_header_id|>\n\nDescribe the video by detailing the following aspects: "
|
||||
"1. The main content and theme of the video."
|
||||
"2. The color, shape, size, texture, quantity, text, and spatial relationships of the objects."
|
||||
"3. Actions, events, behaviors temporal relationships, physical movement changes of the objects."
|
||||
"4. background environment, light, style and atmosphere."
|
||||
"5. camera angles, movements, and transitions used in the video:<|eot_id|>"
|
||||
"<|start_header_id|>user<|end_header_id|>\n\n{}<|eot_id|>"
|
||||
)
|
||||
|
||||
PROMPT_TEMPLATE_ENCODE_I2V = (
|
||||
"<|start_header_id|>system<|end_header_id|>\n\n<image>\nDescribe the image by detailing the color, shape, size, texture, "
|
||||
"quantity, text, spatial relationships of the objects and background:<|eot_id|>"
|
||||
"<|start_header_id|>user<|end_header_id|>\n\n{}<|eot_id|>"
|
||||
"<|start_header_id|>assistant<|end_header_id|>\n\n"
|
||||
)
|
||||
|
||||
PROMPT_TEMPLATE_ENCODE_VIDEO_I2V = (
|
||||
"<|start_header_id|>system<|end_header_id|>\n\n<image>\nDescribe the video by detailing the following aspects according to the reference image: "
|
||||
"1. The main content and theme of the video."
|
||||
"2. The color, shape, size, texture, quantity, text, and spatial relationships of the objects."
|
||||
"3. Actions, events, behaviors temporal relationships, physical movement changes of the objects."
|
||||
"4. background environment, light, style and atmosphere."
|
||||
"5. camera angles, movements, and transitions used in the video:<|eot_id|>\n\n"
|
||||
"<|start_header_id|>user<|end_header_id|>\n\n{}<|eot_id|>"
|
||||
"<|start_header_id|>assistant<|end_header_id|>\n\n"
|
||||
)
|
||||
|
||||
NEGATIVE_PROMPT = "Aerial view, aerial view, overexposed, low quality, deformation, a poor composition, bad hands, bad teeth, bad eyes, bad limbs, distortion"
|
||||
NEGATIVE_PROMPT_I2V = "deformation, a poor composition and deformed video, bad teeth, bad eyes, bad limbs"
|
||||
|
||||
PROMPT_TEMPLATE = {
|
||||
"dit-llm-encode": {
|
||||
"template": PROMPT_TEMPLATE_ENCODE,
|
||||
"crop_start": 36,
|
||||
},
|
||||
"dit-llm-encode-video": {
|
||||
"template": PROMPT_TEMPLATE_ENCODE_VIDEO,
|
||||
"crop_start": 95,
|
||||
},
|
||||
"dit-llm-encode-i2v": {
|
||||
"template": PROMPT_TEMPLATE_ENCODE_I2V,
|
||||
"crop_start": 36,
|
||||
"image_emb_start": 5,
|
||||
"image_emb_end": 581,
|
||||
"image_emb_len": 576,
|
||||
"double_return_token_id": 271
|
||||
},
|
||||
"dit-llm-encode-video-i2v": {
|
||||
"template": PROMPT_TEMPLATE_ENCODE_VIDEO_I2V,
|
||||
"crop_start": 103,
|
||||
"image_emb_start": 5,
|
||||
"image_emb_end": 581,
|
||||
"image_emb_len": 576,
|
||||
"double_return_token_id": 271
|
||||
},
|
||||
}
|
||||
|
||||
# ======================= Model ======================
|
||||
PRECISIONS = {"fp32", "fp16", "bf16"}
|
||||
NORMALIZATION_TYPE = {"layer", "rms"}
|
||||
ACTIVATION_TYPE = {"relu", "silu", "gelu", "gelu_tanh"}
|
||||
|
||||
# =================== Model Path =====================
|
||||
MODEL_BASE = os.getenv("MODEL_BASE", "./ckpts")
|
||||
|
||||
# =================== Data =======================
|
||||
DATA_TYPE = {"image", "video", "image_video"}
|
||||
|
||||
# 3D VAE
|
||||
VAE_PATH = {"884-16c-hy": f"{MODEL_BASE}/hunyuan-video-t2v-720p/vae"}
|
||||
|
||||
# Text Encoder
|
||||
TEXT_ENCODER_PATH = {
|
||||
"clipL": f"{MODEL_BASE}/clip_vit_large_patch14",
|
||||
"llm": f"{MODEL_BASE}/llava-llama-3-8b",
|
||||
"llm-i2v": f"{MODEL_BASE}/llava-llama-3-8b",
|
||||
}
|
||||
|
||||
# Tokenizer
|
||||
TOKENIZER_PATH = {
|
||||
"clipL": f"{MODEL_BASE}/clip_vit_large_patch14",
|
||||
"llm": f"{MODEL_BASE}/llava-llama-3-8b",
|
||||
"llm-i2v": f"{MODEL_BASE}/llava-llama-3-8b",
|
||||
}
|
||||
|
||||
TEXT_PROJECTION = {
|
||||
"linear", # Default, an nn.Linear() layer
|
||||
"single_refiner", # Single TokenRefiner. Refer to LI-DiT
|
||||
}
|
||||
|
||||
# Flow Matching path type
|
||||
FLOW_PATH_TYPE = {
|
||||
"linear", # Linear trajectory between noise and data
|
||||
"gvp", # Generalized variance-preserving SDE
|
||||
"vp", # Variance-preserving SDE
|
||||
}
|
||||
|
||||
# Flow Matching predict type
|
||||
FLOW_PREDICT_TYPE = {
|
||||
"velocity", # Predict velocity
|
||||
"score", # Predict score
|
||||
"noise", # Predict noise
|
||||
}
|
||||
|
||||
# Flow Matching loss weight
|
||||
FLOW_LOSS_WEIGHT = {
|
||||
"velocity", # Weight loss by velocity
|
||||
"likelihood", # Weight loss by likelihood
|
||||
}
|
||||
|
||||
# Flow Matching SNR type
|
||||
FLOW_SNR_TYPE = {
|
||||
"lognorm", # Log-normal SNR
|
||||
"uniform", # Uniform SNR
|
||||
}
|
||||
|
||||
# Flow Matching solvers
|
||||
FLOW_SOLVER = {
|
||||
"euler", # Euler solver
|
||||
}
|
||||
2
hyvideo/diffusion/__init__.py
Normal file
2
hyvideo/diffusion/__init__.py
Normal file
@@ -0,0 +1,2 @@
|
||||
from .pipelines import HunyuanVideoPipeline
|
||||
from .schedulers import FlowMatchDiscreteScheduler
|
||||
1
hyvideo/diffusion/pipelines/__init__.py
Normal file
1
hyvideo/diffusion/pipelines/__init__.py
Normal file
@@ -0,0 +1 @@
|
||||
from .pipeline_hunyuan_video import HunyuanVideoPipeline
|
||||
1419
hyvideo/diffusion/pipelines/pipeline_hunyuan_video.py
Normal file
1419
hyvideo/diffusion/pipelines/pipeline_hunyuan_video.py
Normal file
File diff suppressed because it is too large
Load Diff
1
hyvideo/diffusion/schedulers/__init__.py
Normal file
1
hyvideo/diffusion/schedulers/__init__.py
Normal file
@@ -0,0 +1 @@
|
||||
from .scheduling_flow_match_discrete import FlowMatchDiscreteScheduler
|
||||
255
hyvideo/diffusion/schedulers/scheduling_flow_match_discrete.py
Normal file
255
hyvideo/diffusion/schedulers/scheduling_flow_match_discrete.py
Normal file
@@ -0,0 +1,255 @@
|
||||
# Copyright 2024 Stability AI, Katherine Crowson and The HuggingFace Team. All rights reserved.
|
||||
#
|
||||
# Licensed under the Apache License, Version 2.0 (the "License");
|
||||
# you may not use this file except in compliance with the License.
|
||||
# You may obtain a copy of the License at
|
||||
#
|
||||
# http://www.apache.org/licenses/LICENSE-2.0
|
||||
#
|
||||
# Unless required by applicable law or agreed to in writing, software
|
||||
# distributed under the License is distributed on an "AS IS" BASIS,
|
||||
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
# ==============================================================================
|
||||
#
|
||||
# Modified from diffusers==0.29.2
|
||||
#
|
||||
# ==============================================================================
|
||||
|
||||
from dataclasses import dataclass
|
||||
from typing import Optional, Tuple, Union
|
||||
|
||||
import numpy as np
|
||||
import torch
|
||||
|
||||
from diffusers.configuration_utils import ConfigMixin, register_to_config
|
||||
from diffusers.utils import BaseOutput, logging
|
||||
from diffusers.schedulers.scheduling_utils import SchedulerMixin
|
||||
|
||||
|
||||
|
||||
@dataclass
|
||||
class FlowMatchDiscreteSchedulerOutput(BaseOutput):
|
||||
"""
|
||||
Output class for the scheduler's `step` function output.
|
||||
|
||||
Args:
|
||||
prev_sample (`torch.FloatTensor` of shape `(batch_size, num_channels, height, width)` for images):
|
||||
Computed sample `(x_{t-1})` of previous timestep. `prev_sample` should be used as next model input in the
|
||||
denoising loop.
|
||||
"""
|
||||
|
||||
prev_sample: torch.FloatTensor
|
||||
|
||||
|
||||
class FlowMatchDiscreteScheduler(SchedulerMixin, ConfigMixin):
|
||||
"""
|
||||
Euler scheduler.
|
||||
|
||||
This model inherits from [`SchedulerMixin`] and [`ConfigMixin`]. Check the superclass documentation for the generic
|
||||
methods the library implements for all schedulers such as loading and saving.
|
||||
|
||||
Args:
|
||||
num_train_timesteps (`int`, defaults to 1000):
|
||||
The number of diffusion steps to train the model.
|
||||
timestep_spacing (`str`, defaults to `"linspace"`):
|
||||
The way the timesteps should be scaled. Refer to Table 2 of the [Common Diffusion Noise Schedules and
|
||||
Sample Steps are Flawed](https://huggingface.co/papers/2305.08891) for more information.
|
||||
shift (`float`, defaults to 1.0):
|
||||
The shift value for the timestep schedule.
|
||||
reverse (`bool`, defaults to `True`):
|
||||
Whether to reverse the timestep schedule.
|
||||
"""
|
||||
|
||||
_compatibles = []
|
||||
order = 1
|
||||
|
||||
@register_to_config
|
||||
def __init__(
|
||||
self,
|
||||
num_train_timesteps: int = 1000,
|
||||
shift: float = 1.0,
|
||||
reverse: bool = True,
|
||||
solver: str = "euler",
|
||||
n_tokens: Optional[int] = None,
|
||||
):
|
||||
sigmas = torch.linspace(1, 0, num_train_timesteps + 1)
|
||||
|
||||
if not reverse:
|
||||
sigmas = sigmas.flip(0)
|
||||
|
||||
self.sigmas = sigmas
|
||||
# the value fed to model
|
||||
self.timesteps = (sigmas[:-1] * num_train_timesteps).to(dtype=torch.float32)
|
||||
|
||||
self._step_index = None
|
||||
self._begin_index = None
|
||||
|
||||
self.supported_solver = ["euler"]
|
||||
if solver not in self.supported_solver:
|
||||
raise ValueError(
|
||||
f"Solver {solver} not supported. Supported solvers: {self.supported_solver}"
|
||||
)
|
||||
|
||||
@property
|
||||
def step_index(self):
|
||||
"""
|
||||
The index counter for current timestep. It will increase 1 after each scheduler step.
|
||||
"""
|
||||
return self._step_index
|
||||
|
||||
@property
|
||||
def begin_index(self):
|
||||
"""
|
||||
The index for the first timestep. It should be set from pipeline with `set_begin_index` method.
|
||||
"""
|
||||
return self._begin_index
|
||||
|
||||
# Copied from diffusers.schedulers.scheduling_dpmsolver_multistep.DPMSolverMultistepScheduler.set_begin_index
|
||||
def set_begin_index(self, begin_index: int = 0):
|
||||
"""
|
||||
Sets the begin index for the scheduler. This function should be run from pipeline before the inference.
|
||||
|
||||
Args:
|
||||
begin_index (`int`):
|
||||
The begin index for the scheduler.
|
||||
"""
|
||||
self._begin_index = begin_index
|
||||
|
||||
def _sigma_to_t(self, sigma):
|
||||
return sigma * self.config.num_train_timesteps
|
||||
|
||||
def set_timesteps(
|
||||
self,
|
||||
num_inference_steps: int,
|
||||
device: Union[str, torch.device] = None,
|
||||
n_tokens: int = None,
|
||||
):
|
||||
"""
|
||||
Sets the discrete timesteps used for the diffusion chain (to be run before inference).
|
||||
|
||||
Args:
|
||||
num_inference_steps (`int`):
|
||||
The number of diffusion steps used when generating samples with a pre-trained model.
|
||||
device (`str` or `torch.device`, *optional*):
|
||||
The device to which the timesteps should be moved to. If `None`, the timesteps are not moved.
|
||||
n_tokens (`int`, *optional*):
|
||||
Number of tokens in the input sequence.
|
||||
"""
|
||||
self.num_inference_steps = num_inference_steps
|
||||
|
||||
sigmas = torch.linspace(1, 0, num_inference_steps + 1)
|
||||
sigmas = self.sd3_time_shift(sigmas)
|
||||
|
||||
if not self.config.reverse:
|
||||
sigmas = 1 - sigmas
|
||||
|
||||
self.sigmas = sigmas
|
||||
self.timesteps = (sigmas[:-1] * self.config.num_train_timesteps).to(
|
||||
dtype=torch.float32, device=device
|
||||
)
|
||||
|
||||
# Reset step index
|
||||
self._step_index = None
|
||||
|
||||
def index_for_timestep(self, timestep, schedule_timesteps=None):
|
||||
if schedule_timesteps is None:
|
||||
schedule_timesteps = self.timesteps
|
||||
|
||||
indices = (schedule_timesteps == timestep).nonzero()
|
||||
|
||||
# The sigma index that is taken for the **very** first `step`
|
||||
# is always the second index (or the last index if there is only 1)
|
||||
# This way we can ensure we don't accidentally skip a sigma in
|
||||
# case we start in the middle of the denoising schedule (e.g. for image-to-image)
|
||||
pos = 1 if len(indices) > 1 else 0
|
||||
|
||||
return indices[pos].item()
|
||||
|
||||
def _init_step_index(self, timestep):
|
||||
if self.begin_index is None:
|
||||
if isinstance(timestep, torch.Tensor):
|
||||
timestep = timestep.to(self.timesteps.device)
|
||||
self._step_index = self.index_for_timestep(timestep)
|
||||
else:
|
||||
self._step_index = self._begin_index
|
||||
|
||||
def scale_model_input(
|
||||
self, sample: torch.Tensor, timestep: Optional[int] = None
|
||||
) -> torch.Tensor:
|
||||
return sample
|
||||
|
||||
def sd3_time_shift(self, t: torch.Tensor):
|
||||
return (self.config.shift * t) / (1 + (self.config.shift - 1) * t)
|
||||
|
||||
def step(
|
||||
self,
|
||||
model_output: torch.FloatTensor,
|
||||
timestep: Union[float, torch.FloatTensor],
|
||||
sample: torch.FloatTensor,
|
||||
return_dict: bool = True,
|
||||
) -> Union[FlowMatchDiscreteSchedulerOutput, Tuple]:
|
||||
"""
|
||||
Predict the sample from the previous timestep by reversing the SDE. This function propagates the diffusion
|
||||
process from the learned model outputs (most often the predicted noise).
|
||||
|
||||
Args:
|
||||
model_output (`torch.FloatTensor`):
|
||||
The direct output from learned diffusion model.
|
||||
timestep (`float`):
|
||||
The current discrete timestep in the diffusion chain.
|
||||
sample (`torch.FloatTensor`):
|
||||
A current instance of a sample created by the diffusion process.
|
||||
generator (`torch.Generator`, *optional*):
|
||||
A random number generator.
|
||||
n_tokens (`int`, *optional*):
|
||||
Number of tokens in the input sequence.
|
||||
return_dict (`bool`):
|
||||
Whether or not to return a [`~schedulers.scheduling_euler_discrete.EulerDiscreteSchedulerOutput`] or
|
||||
tuple.
|
||||
|
||||
Returns:
|
||||
[`~schedulers.scheduling_euler_discrete.EulerDiscreteSchedulerOutput`] or `tuple`:
|
||||
If return_dict is `True`, [`~schedulers.scheduling_euler_discrete.EulerDiscreteSchedulerOutput`] is
|
||||
returned, otherwise a tuple is returned where the first element is the sample tensor.
|
||||
"""
|
||||
|
||||
if (
|
||||
isinstance(timestep, int)
|
||||
or isinstance(timestep, torch.IntTensor)
|
||||
or isinstance(timestep, torch.LongTensor)
|
||||
):
|
||||
raise ValueError(
|
||||
(
|
||||
"Passing integer indices (e.g. from `enumerate(timesteps)`) as timesteps to"
|
||||
" `EulerDiscreteScheduler.step()` is not supported. Make sure to pass"
|
||||
" one of the `scheduler.timesteps` as a timestep."
|
||||
),
|
||||
)
|
||||
|
||||
if self.step_index is None:
|
||||
self._init_step_index(timestep)
|
||||
|
||||
# Upcast to avoid precision issues when computing prev_sample
|
||||
sample = sample.to(torch.float32)
|
||||
|
||||
dt = self.sigmas[self.step_index + 1] - self.sigmas[self.step_index]
|
||||
|
||||
if self.config.solver == "euler":
|
||||
prev_sample = sample + model_output.to(torch.float32) * dt
|
||||
else:
|
||||
raise ValueError(
|
||||
f"Solver {self.config.solver} not supported. Supported solvers: {self.supported_solver}"
|
||||
)
|
||||
|
||||
# upon completion increase step index by one
|
||||
self._step_index += 1
|
||||
|
||||
if not return_dict:
|
||||
return (prev_sample,)
|
||||
|
||||
return FlowMatchDiscreteSchedulerOutput(prev_sample=prev_sample)
|
||||
|
||||
def __len__(self):
|
||||
return self.config.num_train_timesteps
|
||||
830
hyvideo/hunyuan.py
Normal file
830
hyvideo/hunyuan.py
Normal file
@@ -0,0 +1,830 @@
|
||||
import os
|
||||
import time
|
||||
import random
|
||||
import functools
|
||||
from typing import List, Optional, Tuple, Union
|
||||
|
||||
from pathlib import Path
|
||||
|
||||
import torch
|
||||
import torch.distributed as dist
|
||||
from hyvideo.constants import PROMPT_TEMPLATE, NEGATIVE_PROMPT, PRECISION_TO_TYPE, NEGATIVE_PROMPT_I2V
|
||||
from hyvideo.vae import load_vae
|
||||
from hyvideo.modules import load_model
|
||||
from hyvideo.text_encoder import TextEncoder
|
||||
from hyvideo.utils.data_utils import align_to, get_closest_ratio, generate_crop_size_list
|
||||
from hyvideo.modules.posemb_layers import get_nd_rotary_pos_embed, get_nd_rotary_pos_embed_new
|
||||
from hyvideo.diffusion.schedulers import FlowMatchDiscreteScheduler
|
||||
from hyvideo.diffusion.pipelines import HunyuanVideoPipeline
|
||||
from PIL import Image
|
||||
import numpy as np
|
||||
import torchvision.transforms as transforms
|
||||
import cv2
|
||||
|
||||
def pad_image(crop_img, size, color=(255, 255, 255), resize_ratio=1):
|
||||
crop_h, crop_w = crop_img.shape[:2]
|
||||
target_w, target_h = size
|
||||
scale_h, scale_w = target_h / crop_h, target_w / crop_w
|
||||
if scale_w > scale_h:
|
||||
resize_h = int(target_h*resize_ratio)
|
||||
resize_w = int(crop_w / crop_h * resize_h)
|
||||
else:
|
||||
resize_w = int(target_w*resize_ratio)
|
||||
resize_h = int(crop_h / crop_w * resize_w)
|
||||
crop_img = cv2.resize(crop_img, (resize_w, resize_h))
|
||||
pad_left = (target_w - resize_w) // 2
|
||||
pad_top = (target_h - resize_h) // 2
|
||||
pad_right = target_w - resize_w - pad_left
|
||||
pad_bottom = target_h - resize_h - pad_top
|
||||
crop_img = cv2.copyMakeBorder(crop_img, pad_top, pad_bottom, pad_left, pad_right, cv2.BORDER_CONSTANT, value=color)
|
||||
return crop_img
|
||||
|
||||
|
||||
|
||||
|
||||
def _merge_input_ids_with_image_features(self, image_features, inputs_embeds, input_ids, attention_mask, labels):
|
||||
num_images, num_image_patches, embed_dim = image_features.shape
|
||||
batch_size, sequence_length = input_ids.shape
|
||||
left_padding = not torch.sum(input_ids[:, -1] == torch.tensor(self.pad_token_id))
|
||||
# 1. Create a mask to know where special image tokens are
|
||||
special_image_token_mask = input_ids == self.config.image_token_index
|
||||
num_special_image_tokens = torch.sum(special_image_token_mask, dim=-1)
|
||||
# Compute the maximum embed dimension
|
||||
max_embed_dim = (num_special_image_tokens.max() * (num_image_patches - 1)) + sequence_length
|
||||
batch_indices, non_image_indices = torch.where(input_ids != self.config.image_token_index)
|
||||
|
||||
# 2. Compute the positions where text should be written
|
||||
# Calculate new positions for text tokens in merged image-text sequence.
|
||||
# `special_image_token_mask` identifies image tokens. Each image token will be replaced by `nb_text_tokens_per_images - 1` text tokens.
|
||||
# `torch.cumsum` computes how each image token shifts subsequent text token positions.
|
||||
# - 1 to adjust for zero-based indexing, as `cumsum` inherently increases indices by one.
|
||||
new_token_positions = torch.cumsum((special_image_token_mask * (num_image_patches - 1) + 1), -1) - 1
|
||||
nb_image_pad = max_embed_dim - 1 - new_token_positions[:, -1]
|
||||
if left_padding:
|
||||
new_token_positions += nb_image_pad[:, None] # offset for left padding
|
||||
text_to_overwrite = new_token_positions[batch_indices, non_image_indices]
|
||||
|
||||
# 3. Create the full embedding, already padded to the maximum position
|
||||
final_embedding = torch.zeros(
|
||||
batch_size, max_embed_dim, embed_dim, dtype=inputs_embeds.dtype, device=inputs_embeds.device
|
||||
)
|
||||
final_attention_mask = torch.zeros(
|
||||
batch_size, max_embed_dim, dtype=attention_mask.dtype, device=inputs_embeds.device
|
||||
)
|
||||
if labels is not None:
|
||||
final_labels = torch.full(
|
||||
(batch_size, max_embed_dim), self.config.ignore_index, dtype=input_ids.dtype, device=input_ids.device
|
||||
)
|
||||
# In case the Vision model or the Language model has been offloaded to CPU, we need to manually
|
||||
# set the corresponding tensors into their correct target device.
|
||||
target_device = inputs_embeds.device
|
||||
batch_indices, non_image_indices, text_to_overwrite = (
|
||||
batch_indices.to(target_device),
|
||||
non_image_indices.to(target_device),
|
||||
text_to_overwrite.to(target_device),
|
||||
)
|
||||
attention_mask = attention_mask.to(target_device)
|
||||
|
||||
# 4. Fill the embeddings based on the mask. If we have ["hey" "<image>", "how", "are"]
|
||||
# we need to index copy on [0, 577, 578, 579] for the text and [1:576] for the image features
|
||||
final_embedding[batch_indices, text_to_overwrite] = inputs_embeds[batch_indices, non_image_indices]
|
||||
final_attention_mask[batch_indices, text_to_overwrite] = attention_mask[batch_indices, non_image_indices]
|
||||
if labels is not None:
|
||||
final_labels[batch_indices, text_to_overwrite] = labels[batch_indices, non_image_indices]
|
||||
|
||||
# 5. Fill the embeddings corresponding to the images. Anything that is not `text_positions` needs filling (#29835)
|
||||
image_to_overwrite = torch.full(
|
||||
(batch_size, max_embed_dim), True, dtype=torch.bool, device=inputs_embeds.device
|
||||
)
|
||||
image_to_overwrite[batch_indices, text_to_overwrite] = False
|
||||
image_to_overwrite &= image_to_overwrite.cumsum(-1) - 1 >= nb_image_pad[:, None].to(target_device)
|
||||
|
||||
if image_to_overwrite.sum() != image_features.shape[:-1].numel():
|
||||
raise ValueError(
|
||||
f"The input provided to the model are wrong. The number of image tokens is {torch.sum(special_image_token_mask)} while"
|
||||
f" the number of image given to the model is {num_images}. This prevents correct indexing and breaks batch generation."
|
||||
)
|
||||
|
||||
final_embedding[image_to_overwrite] = image_features.contiguous().reshape(-1, embed_dim).to(target_device)
|
||||
final_attention_mask |= image_to_overwrite
|
||||
position_ids = (final_attention_mask.cumsum(-1) - 1).masked_fill_((final_attention_mask == 0), 1)
|
||||
|
||||
# 6. Mask out the embedding at padding positions, as we later use the past_key_value value to determine the non-attended tokens.
|
||||
batch_indices, pad_indices = torch.where(input_ids == self.pad_token_id)
|
||||
indices_to_mask = new_token_positions[batch_indices, pad_indices]
|
||||
|
||||
final_embedding[batch_indices, indices_to_mask] = 0
|
||||
|
||||
if labels is None:
|
||||
final_labels = None
|
||||
|
||||
return final_embedding, final_attention_mask, final_labels, position_ids
|
||||
|
||||
def patched_llava_forward(
|
||||
self,
|
||||
input_ids: torch.LongTensor = None,
|
||||
pixel_values: torch.FloatTensor = None,
|
||||
attention_mask: Optional[torch.Tensor] = None,
|
||||
position_ids: Optional[torch.LongTensor] = None,
|
||||
past_key_values: Optional[List[torch.FloatTensor]] = None,
|
||||
inputs_embeds: Optional[torch.FloatTensor] = None,
|
||||
vision_feature_layer: Optional[int] = None,
|
||||
vision_feature_select_strategy: Optional[str] = None,
|
||||
labels: Optional[torch.LongTensor] = None,
|
||||
use_cache: Optional[bool] = None,
|
||||
output_attentions: Optional[bool] = None,
|
||||
output_hidden_states: Optional[bool] = None,
|
||||
return_dict: Optional[bool] = None,
|
||||
cache_position: Optional[torch.LongTensor] = None,
|
||||
num_logits_to_keep: int = 0,
|
||||
):
|
||||
from transformers.models.llava.modeling_llava import LlavaCausalLMOutputWithPast
|
||||
|
||||
|
||||
output_attentions = output_attentions if output_attentions is not None else self.config.output_attentions
|
||||
output_hidden_states = (
|
||||
output_hidden_states if output_hidden_states is not None else self.config.output_hidden_states
|
||||
)
|
||||
return_dict = return_dict if return_dict is not None else self.config.use_return_dict
|
||||
vision_feature_layer = (
|
||||
vision_feature_layer if vision_feature_layer is not None else self.config.vision_feature_layer
|
||||
)
|
||||
vision_feature_select_strategy = (
|
||||
vision_feature_select_strategy
|
||||
if vision_feature_select_strategy is not None
|
||||
else self.config.vision_feature_select_strategy
|
||||
)
|
||||
|
||||
if (input_ids is None) ^ (inputs_embeds is not None):
|
||||
raise ValueError("You must specify exactly one of input_ids or inputs_embeds")
|
||||
|
||||
if pixel_values is not None and inputs_embeds is not None:
|
||||
raise ValueError(
|
||||
"You cannot specify both pixel_values and inputs_embeds at the same time, and must specify either one"
|
||||
)
|
||||
|
||||
if inputs_embeds is None:
|
||||
inputs_embeds = self.get_input_embeddings()(input_ids)
|
||||
|
||||
image_features = None
|
||||
if pixel_values is not None:
|
||||
image_features = self.get_image_features(
|
||||
pixel_values=pixel_values,
|
||||
vision_feature_layer=vision_feature_layer,
|
||||
vision_feature_select_strategy=vision_feature_select_strategy,
|
||||
)
|
||||
|
||||
|
||||
inputs_embeds, attention_mask, labels, position_ids = self._merge_input_ids_with_image_features(
|
||||
image_features, inputs_embeds, input_ids, attention_mask, labels
|
||||
)
|
||||
cache_position = torch.arange(attention_mask.shape[1], device=attention_mask.device)
|
||||
|
||||
|
||||
outputs = self.language_model(
|
||||
attention_mask=attention_mask,
|
||||
position_ids=position_ids,
|
||||
past_key_values=past_key_values,
|
||||
inputs_embeds=inputs_embeds,
|
||||
use_cache=use_cache,
|
||||
output_attentions=output_attentions,
|
||||
output_hidden_states=output_hidden_states,
|
||||
return_dict=return_dict,
|
||||
cache_position=cache_position,
|
||||
num_logits_to_keep=num_logits_to_keep,
|
||||
)
|
||||
|
||||
logits = outputs[0]
|
||||
|
||||
loss = None
|
||||
|
||||
if not return_dict:
|
||||
output = (logits,) + outputs[1:]
|
||||
return (loss,) + output if loss is not None else output
|
||||
|
||||
return LlavaCausalLMOutputWithPast(
|
||||
loss=loss,
|
||||
logits=logits,
|
||||
past_key_values=outputs.past_key_values,
|
||||
hidden_states=outputs.hidden_states,
|
||||
attentions=outputs.attentions,
|
||||
image_hidden_states=image_features if pixel_values is not None else None,
|
||||
)
|
||||
|
||||
class DataPreprocess(object):
|
||||
def __init__(self):
|
||||
self.llava_size = (336, 336)
|
||||
self.llava_transform = transforms.Compose(
|
||||
[
|
||||
transforms.Resize(self.llava_size, interpolation=transforms.InterpolationMode.BILINEAR),
|
||||
transforms.ToTensor(),
|
||||
transforms.Normalize((0.48145466, 0.4578275, 0.4082107), (0.26862954, 0.26130258, 0.27577711)),
|
||||
]
|
||||
)
|
||||
|
||||
def get_batch(self, image , size):
|
||||
image = np.asarray(image)
|
||||
llava_item_image = pad_image(image.copy(), self.llava_size)
|
||||
uncond_llava_item_image = np.ones_like(llava_item_image) * 255
|
||||
cat_item_image = pad_image(image.copy(), size)
|
||||
|
||||
llava_item_tensor = self.llava_transform(Image.fromarray(llava_item_image.astype(np.uint8)))
|
||||
uncond_llava_item_tensor = self.llava_transform(Image.fromarray(uncond_llava_item_image))
|
||||
cat_item_tensor = torch.from_numpy(cat_item_image.copy()).permute((2, 0, 1)) / 255.0
|
||||
# batch = {
|
||||
# "pixel_value_llava": llava_item_tensor.unsqueeze(0),
|
||||
# "uncond_pixel_value_llava": uncond_llava_item_tensor.unsqueeze(0),
|
||||
# 'pixel_value_ref': cat_item_tensor.unsqueeze(0),
|
||||
# }
|
||||
return llava_item_tensor.unsqueeze(0), uncond_llava_item_tensor.unsqueeze(0), cat_item_tensor.unsqueeze(0)
|
||||
|
||||
class Inference(object):
|
||||
def __init__(
|
||||
self,
|
||||
i2v,
|
||||
enable_cfg,
|
||||
vae,
|
||||
vae_kwargs,
|
||||
text_encoder,
|
||||
model,
|
||||
text_encoder_2=None,
|
||||
pipeline=None,
|
||||
device=None,
|
||||
):
|
||||
self.i2v = i2v
|
||||
self.enable_cfg = enable_cfg
|
||||
self.vae = vae
|
||||
self.vae_kwargs = vae_kwargs
|
||||
|
||||
self.text_encoder = text_encoder
|
||||
self.text_encoder_2 = text_encoder_2
|
||||
|
||||
self.model = model
|
||||
self.pipeline = pipeline
|
||||
|
||||
self.device = "cuda"
|
||||
|
||||
|
||||
|
||||
@classmethod
|
||||
def from_pretrained(cls, model_filepath, text_encoder_filepath, dtype = torch.bfloat16, VAE_dtype = torch.float16, mixed_precision_transformer =torch.bfloat16 , **kwargs):
|
||||
|
||||
device = "cuda"
|
||||
|
||||
import transformers
|
||||
transformers.models.llava.modeling_llava.LlavaForConditionalGeneration.forward = patched_llava_forward # force legacy behaviour to be able to use tansformers v>(4.47)
|
||||
transformers.models.llava.modeling_llava.LlavaForConditionalGeneration._merge_input_ids_with_image_features = _merge_input_ids_with_image_features
|
||||
|
||||
torch.set_grad_enabled(False)
|
||||
text_len = 512
|
||||
latent_channels = 16
|
||||
precision = "bf16"
|
||||
vae_precision = "fp32" if VAE_dtype == torch.float32 else "bf16"
|
||||
embedded_cfg_scale = 6
|
||||
i2v_condition_type = None
|
||||
i2v_mode = "i2v" in model_filepath[0]
|
||||
custom = False
|
||||
if i2v_mode:
|
||||
model_id = "HYVideo-T/2"
|
||||
i2v_condition_type = "token_replace"
|
||||
elif "custom" in model_filepath[0]:
|
||||
model_id = "HYVideo-T/2-custom"
|
||||
custom = True
|
||||
else:
|
||||
model_id = "HYVideo-T/2-cfgdistill"
|
||||
|
||||
if i2v_mode and i2v_condition_type == "latent_concat":
|
||||
in_channels = latent_channels * 2 + 1
|
||||
image_embed_interleave = 2
|
||||
elif i2v_mode and i2v_condition_type == "token_replace":
|
||||
in_channels = latent_channels
|
||||
image_embed_interleave = 4
|
||||
else:
|
||||
in_channels = latent_channels
|
||||
image_embed_interleave = 1
|
||||
out_channels = latent_channels
|
||||
pinToMemory = kwargs.pop("pinToMemory", False)
|
||||
partialPinning = kwargs.pop("partialPinning", False)
|
||||
factor_kwargs = kwargs | {"device": "meta", "dtype": PRECISION_TO_TYPE[precision]}
|
||||
|
||||
if embedded_cfg_scale and i2v_mode:
|
||||
factor_kwargs["guidance_embed"] = True
|
||||
|
||||
model = load_model(
|
||||
model = model_id,
|
||||
i2v_condition_type = i2v_condition_type,
|
||||
in_channels=in_channels,
|
||||
out_channels=out_channels,
|
||||
factor_kwargs=factor_kwargs,
|
||||
)
|
||||
|
||||
|
||||
from mmgp import offload
|
||||
# model = Inference.load_state_dict(args, model, model_filepath)
|
||||
|
||||
# model_filepath ="c:/temp/hc/mp_rank_00_model_states.pt"
|
||||
offload.load_model_data(model, model_filepath, pinToMemory = pinToMemory, partialPinning = partialPinning)
|
||||
pass
|
||||
# offload.save_model(model, "hunyuan_video_custom_720_bf16.safetensors")
|
||||
# offload.save_model(model, "hunyuan_video_custom_720_quanto_bf16_int8.safetensors", do_quantize= True)
|
||||
|
||||
model.mixed_precision = mixed_precision_transformer
|
||||
|
||||
if model.mixed_precision :
|
||||
model._lock_dtype = torch.float32
|
||||
model.lock_layers_dtypes(torch.float32)
|
||||
model.eval()
|
||||
|
||||
# ============================= Build extra models ========================
|
||||
# VAE
|
||||
if custom:
|
||||
vae_configpath = "ckpts/hunyuan_video_custom_VAE_config.json"
|
||||
vae_filepath = "ckpts/hunyuan_video_custom_VAE_fp32.safetensors"
|
||||
else:
|
||||
vae_configpath = "ckpts/hunyuan_video_VAE_config.json"
|
||||
vae_filepath = "ckpts/hunyuan_video_VAE_fp32.safetensors"
|
||||
|
||||
# config = AutoencoderKLCausal3D.load_config("ckpts/hunyuan_video_VAE_config.json")
|
||||
# config = AutoencoderKLCausal3D.load_config("c:/temp/hvae/config_vae.json")
|
||||
|
||||
vae, _, s_ratio, t_ratio = load_vae( "884-16c-hy", vae_path= vae_filepath, vae_config_path= vae_configpath, vae_precision= vae_precision, device= "cpu", )
|
||||
|
||||
vae._model_dtype = torch.float32 if VAE_dtype == torch.float32 else torch.bfloat16
|
||||
vae_kwargs = {"s_ratio": s_ratio, "t_ratio": t_ratio}
|
||||
enable_cfg = False
|
||||
# Text encoder
|
||||
if i2v_mode:
|
||||
text_encoder = "llm-i2v"
|
||||
tokenizer = "llm-i2v"
|
||||
prompt_template = "dit-llm-encode-i2v"
|
||||
prompt_template_video = "dit-llm-encode-video-i2v"
|
||||
elif custom :
|
||||
text_encoder = "llm-i2v"
|
||||
tokenizer = "llm-i2v"
|
||||
prompt_template = "dit-llm-encode"
|
||||
prompt_template_video = "dit-llm-encode-video"
|
||||
enable_cfg = True
|
||||
else:
|
||||
text_encoder = "llm"
|
||||
tokenizer = "llm"
|
||||
prompt_template = "dit-llm-encode"
|
||||
prompt_template_video = "dit-llm-encode-video"
|
||||
|
||||
if prompt_template_video is not None:
|
||||
crop_start = PROMPT_TEMPLATE[prompt_template_video].get( "crop_start", 0 )
|
||||
elif prompt_template is not None:
|
||||
crop_start = PROMPT_TEMPLATE[prompt_template].get("crop_start", 0)
|
||||
else:
|
||||
crop_start = 0
|
||||
max_length = text_len + crop_start
|
||||
|
||||
# prompt_template
|
||||
prompt_template = PROMPT_TEMPLATE[prompt_template] if prompt_template is not None else None
|
||||
|
||||
# prompt_template_video
|
||||
prompt_template_video = PROMPT_TEMPLATE[prompt_template_video] if prompt_template_video is not None else None
|
||||
|
||||
|
||||
text_encoder = TextEncoder(
|
||||
text_encoder_type=text_encoder,
|
||||
max_length=max_length,
|
||||
text_encoder_precision="fp16",
|
||||
tokenizer_type=tokenizer,
|
||||
i2v_mode=i2v_mode,
|
||||
prompt_template=prompt_template,
|
||||
prompt_template_video=prompt_template_video,
|
||||
hidden_state_skip_layer=2,
|
||||
apply_final_norm=False,
|
||||
reproduce=True,
|
||||
device="cpu",
|
||||
image_embed_interleave=image_embed_interleave,
|
||||
text_encoder_path = text_encoder_filepath
|
||||
)
|
||||
|
||||
text_encoder_2 = TextEncoder(
|
||||
text_encoder_type="clipL",
|
||||
max_length=77,
|
||||
text_encoder_precision="fp16",
|
||||
tokenizer_type="clipL",
|
||||
reproduce=True,
|
||||
device="cpu",
|
||||
)
|
||||
|
||||
return cls(
|
||||
i2v=i2v_mode,
|
||||
enable_cfg = enable_cfg,
|
||||
vae=vae,
|
||||
vae_kwargs=vae_kwargs,
|
||||
text_encoder=text_encoder,
|
||||
text_encoder_2=text_encoder_2,
|
||||
model=model,
|
||||
device=device,
|
||||
)
|
||||
|
||||
|
||||
|
||||
class HunyuanVideoSampler(Inference):
|
||||
def __init__(
|
||||
self,
|
||||
i2v,
|
||||
enable_cfg,
|
||||
vae,
|
||||
vae_kwargs,
|
||||
text_encoder,
|
||||
model,
|
||||
text_encoder_2=None,
|
||||
pipeline=None,
|
||||
device=0,
|
||||
):
|
||||
super().__init__(
|
||||
i2v,
|
||||
enable_cfg,
|
||||
vae,
|
||||
vae_kwargs,
|
||||
text_encoder,
|
||||
model,
|
||||
text_encoder_2=text_encoder_2,
|
||||
pipeline=pipeline,
|
||||
device=device,
|
||||
)
|
||||
|
||||
self.i2v_mode = i2v
|
||||
self.enable_cfg = enable_cfg
|
||||
self.pipeline = self.load_diffusion_pipeline(
|
||||
vae=self.vae,
|
||||
text_encoder=self.text_encoder,
|
||||
text_encoder_2=self.text_encoder_2,
|
||||
model=self.model,
|
||||
device=self.device,
|
||||
)
|
||||
|
||||
if self.i2v_mode:
|
||||
self.default_negative_prompt = NEGATIVE_PROMPT_I2V
|
||||
else:
|
||||
self.default_negative_prompt = NEGATIVE_PROMPT
|
||||
|
||||
@property
|
||||
def _interrupt(self):
|
||||
return self.pipeline._interrupt
|
||||
|
||||
@_interrupt.setter
|
||||
def _interrupt(self, value):
|
||||
self.pipeline._interrupt =value
|
||||
|
||||
def load_diffusion_pipeline(
|
||||
self,
|
||||
vae,
|
||||
text_encoder,
|
||||
text_encoder_2,
|
||||
model,
|
||||
scheduler=None,
|
||||
device=None,
|
||||
progress_bar_config=None,
|
||||
#data_type="video",
|
||||
):
|
||||
"""Load the denoising scheduler for inference."""
|
||||
if scheduler is None:
|
||||
scheduler = FlowMatchDiscreteScheduler(
|
||||
shift=6.0,
|
||||
reverse=True,
|
||||
solver="euler",
|
||||
)
|
||||
|
||||
pipeline = HunyuanVideoPipeline(
|
||||
vae=vae,
|
||||
text_encoder=text_encoder,
|
||||
text_encoder_2=text_encoder_2,
|
||||
transformer=model,
|
||||
scheduler=scheduler,
|
||||
progress_bar_config=progress_bar_config,
|
||||
)
|
||||
|
||||
return pipeline
|
||||
|
||||
def get_rotary_pos_embed_new(self, video_length, height, width, concat_dict={}):
|
||||
target_ndim = 3
|
||||
ndim = 5 - 2
|
||||
latents_size = [(video_length-1)//4+1 , height//8, width//8]
|
||||
|
||||
if isinstance(self.model.patch_size, int):
|
||||
assert all(s % self.model.patch_size == 0 for s in latents_size), \
|
||||
f"Latent size(last {ndim} dimensions) should be divisible by patch size({self.model.patch_size}), " \
|
||||
f"but got {latents_size}."
|
||||
rope_sizes = [s // self.model.patch_size for s in latents_size]
|
||||
elif isinstance(self.model.patch_size, list):
|
||||
assert all(s % self.model.patch_size[idx] == 0 for idx, s in enumerate(latents_size)), \
|
||||
f"Latent size(last {ndim} dimensions) should be divisible by patch size({self.model.patch_size}), " \
|
||||
f"but got {latents_size}."
|
||||
rope_sizes = [s // self.model.patch_size[idx] for idx, s in enumerate(latents_size)]
|
||||
|
||||
if len(rope_sizes) != target_ndim:
|
||||
rope_sizes = [1] * (target_ndim - len(rope_sizes)) + rope_sizes # time axis
|
||||
head_dim = self.model.hidden_size // self.model.heads_num
|
||||
rope_dim_list = self.model.rope_dim_list
|
||||
if rope_dim_list is None:
|
||||
rope_dim_list = [head_dim // target_ndim for _ in range(target_ndim)]
|
||||
assert sum(rope_dim_list) == head_dim, "sum(rope_dim_list) should equal to head_dim of attention layer"
|
||||
freqs_cos, freqs_sin = get_nd_rotary_pos_embed_new(rope_dim_list,
|
||||
rope_sizes,
|
||||
theta=256,
|
||||
use_real=True,
|
||||
theta_rescale_factor=1,
|
||||
concat_dict=concat_dict)
|
||||
return freqs_cos, freqs_sin
|
||||
|
||||
def get_rotary_pos_embed(self, video_length, height, width, enable_riflex = False):
|
||||
target_ndim = 3
|
||||
ndim = 5 - 2
|
||||
# 884
|
||||
vae = "884-16c-hy"
|
||||
if "884" in vae:
|
||||
latents_size = [(video_length - 1) // 4 + 1, height // 8, width // 8]
|
||||
elif "888" in vae:
|
||||
latents_size = [(video_length - 1) // 8 + 1, height // 8, width // 8]
|
||||
else:
|
||||
latents_size = [video_length, height // 8, width // 8]
|
||||
|
||||
if isinstance(self.model.patch_size, int):
|
||||
assert all(s % self.model.patch_size == 0 for s in latents_size), (
|
||||
f"Latent size(last {ndim} dimensions) should be divisible by patch size({self.model.patch_size}), "
|
||||
f"but got {latents_size}."
|
||||
)
|
||||
rope_sizes = [s // self.model.patch_size for s in latents_size]
|
||||
elif isinstance(self.model.patch_size, list):
|
||||
assert all(
|
||||
s % self.model.patch_size[idx] == 0
|
||||
for idx, s in enumerate(latents_size)
|
||||
), (
|
||||
f"Latent size(last {ndim} dimensions) should be divisible by patch size({self.model.patch_size}), "
|
||||
f"but got {latents_size}."
|
||||
)
|
||||
rope_sizes = [
|
||||
s // self.model.patch_size[idx] for idx, s in enumerate(latents_size)
|
||||
]
|
||||
|
||||
if len(rope_sizes) != target_ndim:
|
||||
rope_sizes = [1] * (target_ndim - len(rope_sizes)) + rope_sizes # time axis
|
||||
head_dim = self.model.hidden_size // self.model.heads_num
|
||||
rope_dim_list = self.model.rope_dim_list
|
||||
if rope_dim_list is None:
|
||||
rope_dim_list = [head_dim // target_ndim for _ in range(target_ndim)]
|
||||
assert (
|
||||
sum(rope_dim_list) == head_dim
|
||||
), "sum(rope_dim_list) should equal to head_dim of attention layer"
|
||||
freqs_cos, freqs_sin = get_nd_rotary_pos_embed(
|
||||
rope_dim_list,
|
||||
rope_sizes,
|
||||
theta=256,
|
||||
use_real=True,
|
||||
theta_rescale_factor=1,
|
||||
L_test = (video_length - 1) // 4 + 1,
|
||||
enable_riflex = enable_riflex
|
||||
)
|
||||
return freqs_cos, freqs_sin
|
||||
|
||||
|
||||
def generate(
|
||||
self,
|
||||
input_prompt,
|
||||
input_ref_images = None,
|
||||
height=192,
|
||||
width=336,
|
||||
frame_num=129,
|
||||
seed=None,
|
||||
n_prompt=None,
|
||||
sampling_steps=50,
|
||||
guide_scale=1.0,
|
||||
shift=5.0,
|
||||
embedded_guidance_scale=6.0,
|
||||
batch_size=1,
|
||||
num_videos_per_prompt=1,
|
||||
i2v_resolution="720p",
|
||||
image_start=None,
|
||||
enable_riflex = False,
|
||||
i2v_condition_type: str = "token_replace",
|
||||
i2v_stability=True,
|
||||
VAE_tile_size = None,
|
||||
joint_pass = False,
|
||||
cfg_star_switch = False,
|
||||
**kwargs,
|
||||
):
|
||||
|
||||
if VAE_tile_size != None:
|
||||
self.vae.tile_sample_min_tsize = VAE_tile_size["tile_sample_min_tsize"]
|
||||
self.vae.tile_latent_min_tsize = VAE_tile_size["tile_latent_min_tsize"]
|
||||
self.vae.tile_sample_min_size = VAE_tile_size["tile_sample_min_size"]
|
||||
self.vae.tile_latent_min_size = VAE_tile_size["tile_latent_min_size"]
|
||||
self.vae.tile_overlap_factor = VAE_tile_size["tile_overlap_factor"]
|
||||
|
||||
i2v_mode= self.i2v_mode
|
||||
if not self.enable_cfg:
|
||||
guide_scale=1.0
|
||||
|
||||
|
||||
out_dict = dict()
|
||||
|
||||
# ========================================================================
|
||||
# Arguments: seed
|
||||
# ========================================================================
|
||||
if isinstance(seed, torch.Tensor):
|
||||
seed = seed.tolist()
|
||||
if seed is None:
|
||||
seeds = [
|
||||
random.randint(0, 1_000_000)
|
||||
for _ in range(batch_size * num_videos_per_prompt)
|
||||
]
|
||||
elif isinstance(seed, int):
|
||||
seeds = [
|
||||
seed + i
|
||||
for _ in range(batch_size)
|
||||
for i in range(num_videos_per_prompt)
|
||||
]
|
||||
elif isinstance(seed, (list, tuple)):
|
||||
if len(seed) == batch_size:
|
||||
seeds = [
|
||||
int(seed[i]) + j
|
||||
for i in range(batch_size)
|
||||
for j in range(num_videos_per_prompt)
|
||||
]
|
||||
elif len(seed) == batch_size * num_videos_per_prompt:
|
||||
seeds = [int(s) for s in seed]
|
||||
else:
|
||||
raise ValueError(
|
||||
f"Length of seed must be equal to number of prompt(batch_size) or "
|
||||
f"batch_size * num_videos_per_prompt ({batch_size} * {num_videos_per_prompt}), got {seed}."
|
||||
)
|
||||
else:
|
||||
raise ValueError(
|
||||
f"Seed must be an integer, a list of integers, or None, got {seed}."
|
||||
)
|
||||
from wan.utils.utils import seed_everything
|
||||
seed_everything(seed)
|
||||
generator = [torch.Generator("cuda").manual_seed(seed) for seed in seeds]
|
||||
# generator = [torch.Generator(self.device).manual_seed(seed) for seed in seeds]
|
||||
out_dict["seeds"] = seeds
|
||||
|
||||
# ========================================================================
|
||||
# Arguments: target_width, target_height, target_frame_num
|
||||
# ========================================================================
|
||||
if width <= 0 or height <= 0 or frame_num <= 0:
|
||||
raise ValueError(
|
||||
f"`height` and `width` and `frame_num` must be positive integers, got height={height}, width={width}, frame_num={frame_num}"
|
||||
)
|
||||
if (frame_num - 1) % 4 != 0:
|
||||
raise ValueError(
|
||||
f"`frame_num-1` must be a multiple of 4, got {frame_num}"
|
||||
)
|
||||
|
||||
target_height = align_to(height, 16)
|
||||
target_width = align_to(width, 16)
|
||||
target_frame_num = frame_num
|
||||
|
||||
out_dict["size"] = (target_height, target_width, target_frame_num)
|
||||
|
||||
if input_ref_images != None:
|
||||
# ip_cfg_scale = 3.0
|
||||
ip_cfg_scale = 0
|
||||
denoise_strength = 1
|
||||
# guide_scale=7.5
|
||||
# shift=13
|
||||
name = "person"
|
||||
input_ref_images = input_ref_images[0]
|
||||
|
||||
# ========================================================================
|
||||
# Arguments: prompt, new_prompt, negative_prompt
|
||||
# ========================================================================
|
||||
if not isinstance(input_prompt, str):
|
||||
raise TypeError(f"`prompt` must be a string, but got {type(input_prompt)}")
|
||||
input_prompt = [input_prompt.strip()]
|
||||
|
||||
# negative prompt
|
||||
if n_prompt is None or n_prompt == "":
|
||||
n_prompt = self.default_negative_prompt
|
||||
if guide_scale == 1.0:
|
||||
n_prompt = ""
|
||||
if not isinstance(n_prompt, str):
|
||||
raise TypeError(
|
||||
f"`negative_prompt` must be a string, but got {type(n_prompt)}"
|
||||
)
|
||||
n_prompt = [n_prompt.strip()]
|
||||
|
||||
# ========================================================================
|
||||
# Scheduler
|
||||
# ========================================================================
|
||||
scheduler = FlowMatchDiscreteScheduler(
|
||||
shift=shift,
|
||||
reverse=True,
|
||||
solver="euler"
|
||||
)
|
||||
self.pipeline.scheduler = scheduler
|
||||
|
||||
# ---------------------------------
|
||||
# Reference condition
|
||||
# ---------------------------------
|
||||
img_latents = None
|
||||
semantic_images = None
|
||||
denoise_strength = 0
|
||||
ip_cfg_scale = 0
|
||||
if i2v_mode:
|
||||
if i2v_resolution == "720p":
|
||||
bucket_hw_base_size = 960
|
||||
elif i2v_resolution == "540p":
|
||||
bucket_hw_base_size = 720
|
||||
elif i2v_resolution == "360p":
|
||||
bucket_hw_base_size = 480
|
||||
else:
|
||||
raise ValueError(f"i2v_resolution: {i2v_resolution} must be in [360p, 540p, 720p]")
|
||||
|
||||
# semantic_images = [Image.open(i2v_image_path).convert('RGB')]
|
||||
semantic_images = [image_start.convert('RGB')] #
|
||||
|
||||
origin_size = semantic_images[0].size
|
||||
|
||||
crop_size_list = generate_crop_size_list(bucket_hw_base_size, 32)
|
||||
aspect_ratios = np.array([round(float(h)/float(w), 5) for h, w in crop_size_list])
|
||||
closest_size, closest_ratio = get_closest_ratio(origin_size[1], origin_size[0], aspect_ratios, crop_size_list)
|
||||
ref_image_transform = transforms.Compose([
|
||||
transforms.Resize(closest_size),
|
||||
transforms.CenterCrop(closest_size),
|
||||
transforms.ToTensor(),
|
||||
transforms.Normalize([0.5], [0.5])
|
||||
])
|
||||
|
||||
semantic_image_pixel_values = [ref_image_transform(semantic_image) for semantic_image in semantic_images]
|
||||
semantic_image_pixel_values = torch.cat(semantic_image_pixel_values).unsqueeze(0).unsqueeze(2).to(self.device)
|
||||
|
||||
with torch.autocast(device_type="cuda", dtype=torch.float16, enabled=True):
|
||||
img_latents = self.pipeline.vae.encode(semantic_image_pixel_values).latent_dist.mode() # B, C, F, H, W
|
||||
img_latents.mul_(self.pipeline.vae.config.scaling_factor)
|
||||
|
||||
target_height, target_width = closest_size
|
||||
|
||||
# ========================================================================
|
||||
# Build Rope freqs
|
||||
# ========================================================================
|
||||
|
||||
if input_ref_images == None:
|
||||
freqs_cos, freqs_sin = self.get_rotary_pos_embed(target_frame_num, target_height, target_width, enable_riflex)
|
||||
else:
|
||||
concat_dict = {'mode': 'timecat-w', 'bias': -1}
|
||||
freqs_cos, freqs_sin = self.get_rotary_pos_embed_new(target_frame_num, target_height, target_width, concat_dict)
|
||||
|
||||
n_tokens = freqs_cos.shape[0]
|
||||
|
||||
|
||||
callback = kwargs.pop("callback", None)
|
||||
callback_steps = kwargs.pop("callback_steps", None)
|
||||
# ========================================================================
|
||||
# Pipeline inference
|
||||
# ========================================================================
|
||||
start_time = time.time()
|
||||
|
||||
|
||||
# "pixel_value_llava": llava_item_tensor.unsqueeze(0),
|
||||
# "uncond_pixel_value_llava": uncond_llava_item_tensor.unsqueeze(0),
|
||||
# 'pixel_value_ref': cat_item_tensor.unsqueeze(0),
|
||||
if input_ref_images == None:
|
||||
pixel_value_llava, uncond_pixel_value_llava, pixel_value_ref = None, None, None
|
||||
name = None
|
||||
else:
|
||||
pixel_value_llava, uncond_pixel_value_llava, pixel_value_ref = DataPreprocess().get_batch(input_ref_images, (target_width, target_height))
|
||||
samples = self.pipeline(
|
||||
prompt=input_prompt,
|
||||
height=target_height,
|
||||
width=target_width,
|
||||
video_length=target_frame_num,
|
||||
num_inference_steps=sampling_steps,
|
||||
guidance_scale=guide_scale,
|
||||
negative_prompt=n_prompt,
|
||||
num_videos_per_prompt=num_videos_per_prompt,
|
||||
generator=generator,
|
||||
output_type="pil",
|
||||
name = name,
|
||||
pixel_value_llava = pixel_value_llava,
|
||||
uncond_pixel_value_llava=uncond_pixel_value_llava,
|
||||
pixel_value_ref=pixel_value_ref,
|
||||
denoise_strength=denoise_strength,
|
||||
ip_cfg_scale=ip_cfg_scale,
|
||||
freqs_cis=(freqs_cos, freqs_sin),
|
||||
n_tokens=n_tokens,
|
||||
embedded_guidance_scale=embedded_guidance_scale,
|
||||
data_type="video" if target_frame_num > 1 else "image",
|
||||
is_progress_bar=True,
|
||||
vae_ver="884-16c-hy",
|
||||
enable_tiling=True,
|
||||
i2v_mode=i2v_mode,
|
||||
i2v_condition_type=i2v_condition_type,
|
||||
i2v_stability=i2v_stability,
|
||||
img_latents=img_latents,
|
||||
semantic_images=semantic_images,
|
||||
joint_pass = joint_pass,
|
||||
cfg_star_rescale = cfg_star_switch,
|
||||
callback = callback,
|
||||
callback_steps = callback_steps,
|
||||
)[0]
|
||||
gen_time = time.time() - start_time
|
||||
if samples == None:
|
||||
return None
|
||||
samples = samples.sub_(0.5).mul_(2).squeeze(0)
|
||||
|
||||
return samples
|
||||
26
hyvideo/modules/__init__.py
Normal file
26
hyvideo/modules/__init__.py
Normal file
@@ -0,0 +1,26 @@
|
||||
from .models import HYVideoDiffusionTransformer, HUNYUAN_VIDEO_CONFIG
|
||||
|
||||
|
||||
def load_model(model, i2v_condition_type, in_channels, out_channels, factor_kwargs):
|
||||
"""load hunyuan video model
|
||||
|
||||
Args:
|
||||
args (dict): model args
|
||||
in_channels (int): input channels number
|
||||
out_channels (int): output channels number
|
||||
factor_kwargs (dict): factor kwargs
|
||||
|
||||
Returns:
|
||||
model (nn.Module): The hunyuan video model
|
||||
"""
|
||||
if model in HUNYUAN_VIDEO_CONFIG.keys():
|
||||
model = HYVideoDiffusionTransformer(
|
||||
i2v_condition_type = i2v_condition_type,
|
||||
in_channels=in_channels,
|
||||
out_channels=out_channels,
|
||||
**HUNYUAN_VIDEO_CONFIG[model],
|
||||
**factor_kwargs,
|
||||
)
|
||||
return model
|
||||
else:
|
||||
raise NotImplementedError()
|
||||
23
hyvideo/modules/activation_layers.py
Normal file
23
hyvideo/modules/activation_layers.py
Normal file
@@ -0,0 +1,23 @@
|
||||
import torch.nn as nn
|
||||
|
||||
|
||||
def get_activation_layer(act_type):
|
||||
"""get activation layer
|
||||
|
||||
Args:
|
||||
act_type (str): the activation type
|
||||
|
||||
Returns:
|
||||
torch.nn.functional: the activation layer
|
||||
"""
|
||||
if act_type == "gelu":
|
||||
return lambda: nn.GELU()
|
||||
elif act_type == "gelu_tanh":
|
||||
# Approximate `tanh` requires torch >= 1.13
|
||||
return lambda: nn.GELU(approximate="tanh")
|
||||
elif act_type == "relu":
|
||||
return nn.ReLU
|
||||
elif act_type == "silu":
|
||||
return nn.SiLU
|
||||
else:
|
||||
raise ValueError(f"Unknown activation type: {act_type}")
|
||||
362
hyvideo/modules/attenion.py
Normal file
362
hyvideo/modules/attenion.py
Normal file
@@ -0,0 +1,362 @@
|
||||
import importlib.metadata
|
||||
import math
|
||||
|
||||
import torch
|
||||
import torch.nn as nn
|
||||
import torch.nn.functional as F
|
||||
from importlib.metadata import version
|
||||
|
||||
def clear_list(l):
|
||||
for i in range(len(l)):
|
||||
l[i] = None
|
||||
|
||||
try:
|
||||
import flash_attn
|
||||
from flash_attn.flash_attn_interface import _flash_attn_forward
|
||||
from flash_attn.flash_attn_interface import flash_attn_varlen_func
|
||||
except ImportError:
|
||||
flash_attn = None
|
||||
flash_attn_varlen_func = None
|
||||
_flash_attn_forward = None
|
||||
|
||||
try:
|
||||
from xformers.ops import memory_efficient_attention
|
||||
except ImportError:
|
||||
memory_efficient_attention = None
|
||||
|
||||
try:
|
||||
from sageattention import sageattn_varlen
|
||||
def sageattn_varlen_wrapper(
|
||||
q,
|
||||
k,
|
||||
v,
|
||||
cu_seqlens_q,
|
||||
cu_seqlens_kv,
|
||||
max_seqlen_q,
|
||||
max_seqlen_kv,
|
||||
):
|
||||
return sageattn_varlen(q, k, v, cu_seqlens_q, cu_seqlens_kv, max_seqlen_q, max_seqlen_kv)
|
||||
except ImportError:
|
||||
sageattn_varlen_wrapper = None
|
||||
|
||||
try:
|
||||
from sageattention import sageattn
|
||||
@torch.compiler.disable()
|
||||
def sageattn_wrapper(
|
||||
qkv_list,
|
||||
attention_length
|
||||
):
|
||||
q,k, v = qkv_list
|
||||
padding_length = q.shape[1] -attention_length
|
||||
q = q[:, :attention_length, :, : ]
|
||||
k = k[:, :attention_length, :, : ]
|
||||
v = v[:, :attention_length, :, : ]
|
||||
|
||||
o = sageattn(q, k, v, tensor_layout="NHD")
|
||||
del q, k ,v
|
||||
clear_list(qkv_list)
|
||||
|
||||
if padding_length > 0:
|
||||
o = torch.cat([o, torch.empty( (o.shape[0], padding_length, *o.shape[-2:]), dtype= o.dtype, device=o.device ) ], 1)
|
||||
|
||||
return o
|
||||
|
||||
except ImportError:
|
||||
sageattn = None
|
||||
|
||||
|
||||
def get_attention_modes():
|
||||
ret = ["sdpa", "auto"]
|
||||
if flash_attn != None:
|
||||
ret.append("flash")
|
||||
if memory_efficient_attention != None:
|
||||
ret.append("xformers")
|
||||
if sageattn_varlen_wrapper != None:
|
||||
ret.append("sage")
|
||||
if sageattn != None and version("sageattention").startswith("2") :
|
||||
ret.append("sage2")
|
||||
|
||||
return ret
|
||||
|
||||
|
||||
|
||||
MEMORY_LAYOUT = {
|
||||
"sdpa": (
|
||||
lambda x: x.transpose(1, 2),
|
||||
lambda x: x.transpose(1, 2),
|
||||
),
|
||||
"xformers": (
|
||||
lambda x: x,
|
||||
lambda x: x,
|
||||
),
|
||||
"sage2": (
|
||||
lambda x: x,
|
||||
lambda x: x,
|
||||
),
|
||||
"sage": (
|
||||
lambda x: x.view(x.shape[0] * x.shape[1], *x.shape[2:]),
|
||||
lambda x: x,
|
||||
),
|
||||
"flash": (
|
||||
lambda x: x.view(x.shape[0] * x.shape[1], *x.shape[2:]),
|
||||
lambda x: x,
|
||||
),
|
||||
"torch": (
|
||||
lambda x: x.transpose(1, 2),
|
||||
lambda x: x.transpose(1, 2),
|
||||
),
|
||||
"vanilla": (
|
||||
lambda x: x.transpose(1, 2),
|
||||
lambda x: x.transpose(1, 2),
|
||||
),
|
||||
}
|
||||
|
||||
@torch.compiler.disable()
|
||||
def sdpa_wrapper(
|
||||
qkv_list,
|
||||
attention_length
|
||||
):
|
||||
q,k, v = qkv_list
|
||||
padding_length = q.shape[2] -attention_length
|
||||
q = q[:, :, :attention_length, :]
|
||||
k = k[:, :, :attention_length, :]
|
||||
v = v[:, :, :attention_length, :]
|
||||
|
||||
o = F.scaled_dot_product_attention(
|
||||
q, k, v, attn_mask=None, is_causal=False
|
||||
)
|
||||
del q, k ,v
|
||||
clear_list(qkv_list)
|
||||
|
||||
if padding_length > 0:
|
||||
o = torch.cat([o, torch.empty( (*o.shape[:2], padding_length, o.shape[-1]), dtype= o.dtype, device=o.device ) ], 2)
|
||||
|
||||
return o
|
||||
|
||||
def get_cu_seqlens(text_mask, img_len):
|
||||
"""Calculate cu_seqlens_q, cu_seqlens_kv using text_mask and img_len
|
||||
|
||||
Args:
|
||||
text_mask (torch.Tensor): the mask of text
|
||||
img_len (int): the length of image
|
||||
|
||||
Returns:
|
||||
torch.Tensor: the calculated cu_seqlens for flash attention
|
||||
"""
|
||||
batch_size = text_mask.shape[0]
|
||||
text_len = text_mask.sum(dim=1)
|
||||
max_len = text_mask.shape[1] + img_len
|
||||
|
||||
cu_seqlens = torch.zeros([2 * batch_size + 1], dtype=torch.int32, device="cuda")
|
||||
|
||||
for i in range(batch_size):
|
||||
s = text_len[i] + img_len
|
||||
s1 = i * max_len + s
|
||||
s2 = (i + 1) * max_len
|
||||
cu_seqlens[2 * i + 1] = s1
|
||||
cu_seqlens[2 * i + 2] = s2
|
||||
|
||||
return cu_seqlens
|
||||
|
||||
|
||||
def attention(
|
||||
qkv_list,
|
||||
mode="flash",
|
||||
drop_rate=0,
|
||||
attn_mask=None,
|
||||
causal=False,
|
||||
cu_seqlens_q=None,
|
||||
cu_seqlens_kv=None,
|
||||
max_seqlen_q=None,
|
||||
max_seqlen_kv=None,
|
||||
batch_size=1,
|
||||
):
|
||||
"""
|
||||
Perform QKV self attention.
|
||||
|
||||
Args:
|
||||
q (torch.Tensor): Query tensor with shape [b, s, a, d], where a is the number of heads.
|
||||
k (torch.Tensor): Key tensor with shape [b, s1, a, d]
|
||||
v (torch.Tensor): Value tensor with shape [b, s1, a, d]
|
||||
mode (str): Attention mode. Choose from 'self_flash', 'cross_flash', 'torch', and 'vanilla'.
|
||||
drop_rate (float): Dropout rate in attention map. (default: 0)
|
||||
attn_mask (torch.Tensor): Attention mask with shape [b, s1] (cross_attn), or [b, a, s, s1] (torch or vanilla).
|
||||
(default: None)
|
||||
causal (bool): Whether to use causal attention. (default: False)
|
||||
cu_seqlens_q (torch.Tensor): dtype torch.int32. The cumulative sequence lengths of the sequences in the batch,
|
||||
used to index into q.
|
||||
cu_seqlens_kv (torch.Tensor): dtype torch.int32. The cumulative sequence lengths of the sequences in the batch,
|
||||
used to index into kv.
|
||||
max_seqlen_q (int): The maximum sequence length in the batch of q.
|
||||
max_seqlen_kv (int): The maximum sequence length in the batch of k and v.
|
||||
|
||||
Returns:
|
||||
torch.Tensor: Output tensor after self attention with shape [b, s, ad]
|
||||
"""
|
||||
pre_attn_layout, post_attn_layout = MEMORY_LAYOUT[mode]
|
||||
q , k , v = qkv_list
|
||||
clear_list(qkv_list)
|
||||
del qkv_list
|
||||
padding_length = 0
|
||||
# if attn_mask == None and mode == "sdpa":
|
||||
# padding_length = q.shape[1] - cu_seqlens_q
|
||||
# q = q[:, :cu_seqlens_q, ... ]
|
||||
# k = k[:, :cu_seqlens_kv, ... ]
|
||||
# v = v[:, :cu_seqlens_kv, ... ]
|
||||
|
||||
q = pre_attn_layout(q)
|
||||
k = pre_attn_layout(k)
|
||||
v = pre_attn_layout(v)
|
||||
|
||||
if mode == "torch":
|
||||
if attn_mask is not None and attn_mask.dtype != torch.bool:
|
||||
attn_mask = attn_mask.to(q.dtype)
|
||||
x = F.scaled_dot_product_attention(
|
||||
q, k, v, attn_mask=attn_mask, dropout_p=drop_rate, is_causal=causal
|
||||
)
|
||||
|
||||
elif mode == "sdpa":
|
||||
# if attn_mask is not None and attn_mask.dtype != torch.bool:
|
||||
# attn_mask = attn_mask.to(q.dtype)
|
||||
# x = F.scaled_dot_product_attention(
|
||||
# q, k, v, attn_mask=attn_mask, dropout_p=drop_rate, is_causal=causal
|
||||
# )
|
||||
assert attn_mask==None
|
||||
qkv_list = [q, k, v]
|
||||
del q, k , v
|
||||
x = sdpa_wrapper( qkv_list, cu_seqlens_q )
|
||||
|
||||
elif mode == "xformers":
|
||||
x = memory_efficient_attention(
|
||||
q, k, v , attn_bias= attn_mask
|
||||
)
|
||||
|
||||
elif mode == "sage2":
|
||||
qkv_list = [q, k, v]
|
||||
del q, k , v
|
||||
x = sageattn_wrapper(qkv_list, cu_seqlens_q)
|
||||
|
||||
elif mode == "sage":
|
||||
x = sageattn_varlen_wrapper(
|
||||
q,
|
||||
k,
|
||||
v,
|
||||
cu_seqlens_q,
|
||||
cu_seqlens_kv,
|
||||
max_seqlen_q,
|
||||
max_seqlen_kv,
|
||||
)
|
||||
# x with shape [(bxs), a, d]
|
||||
x = x.view(
|
||||
batch_size, max_seqlen_q, x.shape[-2], x.shape[-1]
|
||||
) # reshape x to [b, s, a, d]
|
||||
|
||||
elif mode == "flash":
|
||||
x = flash_attn_varlen_func(
|
||||
q,
|
||||
k,
|
||||
v,
|
||||
cu_seqlens_q,
|
||||
cu_seqlens_kv,
|
||||
max_seqlen_q,
|
||||
max_seqlen_kv,
|
||||
)
|
||||
# x with shape [(bxs), a, d]
|
||||
x = x.view(
|
||||
batch_size, max_seqlen_q, x.shape[-2], x.shape[-1]
|
||||
) # reshape x to [b, s, a, d]
|
||||
elif mode == "vanilla":
|
||||
scale_factor = 1 / math.sqrt(q.size(-1))
|
||||
|
||||
b, a, s, _ = q.shape
|
||||
s1 = k.size(2)
|
||||
attn_bias = torch.zeros(b, a, s, s1, dtype=q.dtype, device=q.device)
|
||||
if causal:
|
||||
# Only applied to self attention
|
||||
assert (
|
||||
attn_mask is None
|
||||
), "Causal mask and attn_mask cannot be used together"
|
||||
temp_mask = torch.ones(b, a, s, s, dtype=torch.bool, device=q.device).tril(
|
||||
diagonal=0
|
||||
)
|
||||
attn_bias.masked_fill_(temp_mask.logical_not(), float("-inf"))
|
||||
attn_bias.to(q.dtype)
|
||||
|
||||
if attn_mask is not None:
|
||||
if attn_mask.dtype == torch.bool:
|
||||
attn_bias.masked_fill_(attn_mask.logical_not(), float("-inf"))
|
||||
else:
|
||||
attn_bias += attn_mask
|
||||
|
||||
# TODO: Maybe force q and k to be float32 to avoid numerical overflow
|
||||
attn = (q @ k.transpose(-2, -1)) * scale_factor
|
||||
attn += attn_bias
|
||||
attn = attn.softmax(dim=-1)
|
||||
attn = torch.dropout(attn, p=drop_rate, train=True)
|
||||
x = attn @ v
|
||||
else:
|
||||
raise NotImplementedError(f"Unsupported attention mode: {mode}")
|
||||
|
||||
x = post_attn_layout(x)
|
||||
b, s, a, d = x.shape
|
||||
out = x.reshape(b, s, -1)
|
||||
if padding_length > 0 :
|
||||
out = torch.cat([out, torch.empty( (out.shape[0], padding_length, out.shape[2]), dtype= out.dtype, device=out.device ) ], 1)
|
||||
|
||||
return out
|
||||
|
||||
|
||||
def parallel_attention(
|
||||
hybrid_seq_parallel_attn,
|
||||
q,
|
||||
k,
|
||||
v,
|
||||
img_q_len,
|
||||
img_kv_len,
|
||||
cu_seqlens_q,
|
||||
cu_seqlens_kv
|
||||
):
|
||||
attn1 = hybrid_seq_parallel_attn(
|
||||
None,
|
||||
q[:, :img_q_len, :, :],
|
||||
k[:, :img_kv_len, :, :],
|
||||
v[:, :img_kv_len, :, :],
|
||||
dropout_p=0.0,
|
||||
causal=False,
|
||||
joint_tensor_query=q[:,img_q_len:cu_seqlens_q[1]],
|
||||
joint_tensor_key=k[:,img_kv_len:cu_seqlens_kv[1]],
|
||||
joint_tensor_value=v[:,img_kv_len:cu_seqlens_kv[1]],
|
||||
joint_strategy="rear",
|
||||
)
|
||||
if flash_attn.__version__ >= '2.7.0':
|
||||
attn2, *_ = _flash_attn_forward(
|
||||
q[:,cu_seqlens_q[1]:],
|
||||
k[:,cu_seqlens_kv[1]:],
|
||||
v[:,cu_seqlens_kv[1]:],
|
||||
dropout_p=0.0,
|
||||
softmax_scale=q.shape[-1] ** (-0.5),
|
||||
causal=False,
|
||||
window_size_left=-1,
|
||||
window_size_right=-1,
|
||||
softcap=0.0,
|
||||
alibi_slopes=None,
|
||||
return_softmax=False,
|
||||
)
|
||||
else:
|
||||
attn2, *_ = _flash_attn_forward(
|
||||
q[:,cu_seqlens_q[1]:],
|
||||
k[:,cu_seqlens_kv[1]:],
|
||||
v[:,cu_seqlens_kv[1]:],
|
||||
dropout_p=0.0,
|
||||
softmax_scale=q.shape[-1] ** (-0.5),
|
||||
causal=False,
|
||||
window_size=(-1, -1),
|
||||
softcap=0.0,
|
||||
alibi_slopes=None,
|
||||
return_softmax=False,
|
||||
)
|
||||
attn = torch.cat([attn1, attn2], dim=1)
|
||||
b, s, a, d = attn.shape
|
||||
attn = attn.reshape(b, s, -1)
|
||||
|
||||
return attn
|
||||
157
hyvideo/modules/embed_layers.py
Normal file
157
hyvideo/modules/embed_layers.py
Normal file
@@ -0,0 +1,157 @@
|
||||
import math
|
||||
import torch
|
||||
import torch.nn as nn
|
||||
from einops import rearrange, repeat
|
||||
|
||||
from ..utils.helpers import to_2tuple
|
||||
|
||||
|
||||
class PatchEmbed(nn.Module):
|
||||
"""2D Image to Patch Embedding
|
||||
|
||||
Image to Patch Embedding using Conv2d
|
||||
|
||||
A convolution based approach to patchifying a 2D image w/ embedding projection.
|
||||
|
||||
Based on the impl in https://github.com/google-research/vision_transformer
|
||||
|
||||
Hacked together by / Copyright 2020 Ross Wightman
|
||||
|
||||
Remove the _assert function in forward function to be compatible with multi-resolution images.
|
||||
"""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
patch_size=16,
|
||||
in_chans=3,
|
||||
embed_dim=768,
|
||||
norm_layer=None,
|
||||
flatten=True,
|
||||
bias=True,
|
||||
dtype=None,
|
||||
device=None,
|
||||
):
|
||||
factory_kwargs = {"dtype": dtype, "device": device}
|
||||
super().__init__()
|
||||
patch_size = to_2tuple(patch_size)
|
||||
self.patch_size = patch_size
|
||||
self.flatten = flatten
|
||||
|
||||
self.proj = nn.Conv3d(
|
||||
in_chans,
|
||||
embed_dim,
|
||||
kernel_size=patch_size,
|
||||
stride=patch_size,
|
||||
bias=bias,
|
||||
**factory_kwargs
|
||||
)
|
||||
nn.init.xavier_uniform_(self.proj.weight.view(self.proj.weight.size(0), -1))
|
||||
if bias:
|
||||
nn.init.zeros_(self.proj.bias)
|
||||
|
||||
self.norm = norm_layer(embed_dim) if norm_layer else nn.Identity()
|
||||
|
||||
def forward(self, x):
|
||||
x = self.proj(x)
|
||||
if self.flatten:
|
||||
x = x.flatten(2).transpose(1, 2) # BCHW -> BNC
|
||||
x = self.norm(x)
|
||||
return x
|
||||
|
||||
|
||||
class TextProjection(nn.Module):
|
||||
"""
|
||||
Projects text embeddings. Also handles dropout for classifier-free guidance.
|
||||
|
||||
Adapted from https://github.com/PixArt-alpha/PixArt-alpha/blob/master/diffusion/model/nets/PixArt_blocks.py
|
||||
"""
|
||||
|
||||
def __init__(self, in_channels, hidden_size, act_layer, dtype=None, device=None):
|
||||
factory_kwargs = {"dtype": dtype, "device": device}
|
||||
super().__init__()
|
||||
self.linear_1 = nn.Linear(
|
||||
in_features=in_channels,
|
||||
out_features=hidden_size,
|
||||
bias=True,
|
||||
**factory_kwargs
|
||||
)
|
||||
self.act_1 = act_layer()
|
||||
self.linear_2 = nn.Linear(
|
||||
in_features=hidden_size,
|
||||
out_features=hidden_size,
|
||||
bias=True,
|
||||
**factory_kwargs
|
||||
)
|
||||
|
||||
def forward(self, caption):
|
||||
hidden_states = self.linear_1(caption)
|
||||
hidden_states = self.act_1(hidden_states)
|
||||
hidden_states = self.linear_2(hidden_states)
|
||||
return hidden_states
|
||||
|
||||
|
||||
def timestep_embedding(t, dim, max_period=10000):
|
||||
"""
|
||||
Create sinusoidal timestep embeddings.
|
||||
|
||||
Args:
|
||||
t (torch.Tensor): a 1-D Tensor of N indices, one per batch element. These may be fractional.
|
||||
dim (int): the dimension of the output.
|
||||
max_period (int): controls the minimum frequency of the embeddings.
|
||||
|
||||
Returns:
|
||||
embedding (torch.Tensor): An (N, D) Tensor of positional embeddings.
|
||||
|
||||
.. ref_link: https://github.com/openai/glide-text2im/blob/main/glide_text2im/nn.py
|
||||
"""
|
||||
half = dim // 2
|
||||
freqs = torch.exp(
|
||||
-math.log(max_period)
|
||||
* torch.arange(start=0, end=half, dtype=torch.float32)
|
||||
/ half
|
||||
).to(device=t.device)
|
||||
args = t[:, None].float() * freqs[None]
|
||||
embedding = torch.cat([torch.cos(args), torch.sin(args)], dim=-1)
|
||||
if dim % 2:
|
||||
embedding = torch.cat([embedding, torch.zeros_like(embedding[:, :1])], dim=-1)
|
||||
return embedding
|
||||
|
||||
|
||||
class TimestepEmbedder(nn.Module):
|
||||
"""
|
||||
Embeds scalar timesteps into vector representations.
|
||||
"""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
hidden_size,
|
||||
act_layer,
|
||||
frequency_embedding_size=256,
|
||||
max_period=10000,
|
||||
out_size=None,
|
||||
dtype=None,
|
||||
device=None,
|
||||
):
|
||||
factory_kwargs = {"dtype": dtype, "device": device}
|
||||
super().__init__()
|
||||
self.frequency_embedding_size = frequency_embedding_size
|
||||
self.max_period = max_period
|
||||
if out_size is None:
|
||||
out_size = hidden_size
|
||||
|
||||
self.mlp = nn.Sequential(
|
||||
nn.Linear(
|
||||
frequency_embedding_size, hidden_size, bias=True, **factory_kwargs
|
||||
),
|
||||
act_layer(),
|
||||
nn.Linear(hidden_size, out_size, bias=True, **factory_kwargs),
|
||||
)
|
||||
nn.init.normal_(self.mlp[0].weight, std=0.02)
|
||||
nn.init.normal_(self.mlp[2].weight, std=0.02)
|
||||
|
||||
def forward(self, t):
|
||||
t_freq = timestep_embedding(
|
||||
t, self.frequency_embedding_size, self.max_period
|
||||
).type(self.mlp[0].weight.dtype)
|
||||
t_emb = self.mlp(t_freq)
|
||||
return t_emb
|
||||
131
hyvideo/modules/mlp_layers.py
Normal file
131
hyvideo/modules/mlp_layers.py
Normal file
@@ -0,0 +1,131 @@
|
||||
# Modified from timm library:
|
||||
# https://github.com/huggingface/pytorch-image-models/blob/648aaa41233ba83eb38faf5ba9d415d574823241/timm/layers/mlp.py#L13
|
||||
|
||||
from functools import partial
|
||||
|
||||
import torch
|
||||
import torch.nn as nn
|
||||
|
||||
from .modulate_layers import modulate_
|
||||
from ..utils.helpers import to_2tuple
|
||||
|
||||
|
||||
class MLP(nn.Module):
|
||||
"""MLP as used in Vision Transformer, MLP-Mixer and related networks"""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
in_channels,
|
||||
hidden_channels=None,
|
||||
out_features=None,
|
||||
act_layer=nn.GELU,
|
||||
norm_layer=None,
|
||||
bias=True,
|
||||
drop=0.0,
|
||||
use_conv=False,
|
||||
device=None,
|
||||
dtype=None,
|
||||
):
|
||||
factory_kwargs = {"device": device, "dtype": dtype}
|
||||
super().__init__()
|
||||
out_features = out_features or in_channels
|
||||
hidden_channels = hidden_channels or in_channels
|
||||
bias = to_2tuple(bias)
|
||||
drop_probs = to_2tuple(drop)
|
||||
linear_layer = partial(nn.Conv2d, kernel_size=1) if use_conv else nn.Linear
|
||||
|
||||
self.fc1 = linear_layer(
|
||||
in_channels, hidden_channels, bias=bias[0], **factory_kwargs
|
||||
)
|
||||
self.act = act_layer()
|
||||
self.drop1 = nn.Dropout(drop_probs[0])
|
||||
self.norm = (
|
||||
norm_layer(hidden_channels, **factory_kwargs)
|
||||
if norm_layer is not None
|
||||
else nn.Identity()
|
||||
)
|
||||
self.fc2 = linear_layer(
|
||||
hidden_channels, out_features, bias=bias[1], **factory_kwargs
|
||||
)
|
||||
self.drop2 = nn.Dropout(drop_probs[1])
|
||||
|
||||
def forward(self, x):
|
||||
x = self.fc1(x)
|
||||
x = self.act(x)
|
||||
x = self.drop1(x)
|
||||
x = self.norm(x)
|
||||
x = self.fc2(x)
|
||||
x = self.drop2(x)
|
||||
return x
|
||||
|
||||
def apply_(self, x, divide = 4):
|
||||
x_shape = x.shape
|
||||
x = x.view(-1, x.shape[-1])
|
||||
chunk_size = int(x_shape[1]/divide)
|
||||
x_chunks = torch.split(x, chunk_size)
|
||||
for i, x_chunk in enumerate(x_chunks):
|
||||
mlp_chunk = self.fc1(x_chunk)
|
||||
mlp_chunk = self.act(mlp_chunk)
|
||||
mlp_chunk = self.drop1(mlp_chunk)
|
||||
mlp_chunk = self.norm(mlp_chunk)
|
||||
mlp_chunk = self.fc2(mlp_chunk)
|
||||
x_chunk[...] = self.drop2(mlp_chunk)
|
||||
return x
|
||||
|
||||
#
|
||||
class MLPEmbedder(nn.Module):
|
||||
"""copied from https://github.com/black-forest-labs/flux/blob/main/src/flux/modules/layers.py"""
|
||||
def __init__(self, in_dim: int, hidden_dim: int, device=None, dtype=None):
|
||||
factory_kwargs = {"device": device, "dtype": dtype}
|
||||
super().__init__()
|
||||
self.in_layer = nn.Linear(in_dim, hidden_dim, bias=True, **factory_kwargs)
|
||||
self.silu = nn.SiLU()
|
||||
self.out_layer = nn.Linear(hidden_dim, hidden_dim, bias=True, **factory_kwargs)
|
||||
|
||||
def forward(self, x: torch.Tensor) -> torch.Tensor:
|
||||
return self.out_layer(self.silu(self.in_layer(x)))
|
||||
|
||||
|
||||
class FinalLayer(nn.Module):
|
||||
"""The final layer of DiT."""
|
||||
|
||||
def __init__(
|
||||
self, hidden_size, patch_size, out_channels, act_layer, device=None, dtype=None
|
||||
):
|
||||
factory_kwargs = {"device": device, "dtype": dtype}
|
||||
super().__init__()
|
||||
|
||||
# Just use LayerNorm for the final layer
|
||||
self.norm_final = nn.LayerNorm(
|
||||
hidden_size, elementwise_affine=False, eps=1e-6, **factory_kwargs
|
||||
)
|
||||
if isinstance(patch_size, int):
|
||||
self.linear = nn.Linear(
|
||||
hidden_size,
|
||||
patch_size * patch_size * out_channels,
|
||||
bias=True,
|
||||
**factory_kwargs
|
||||
)
|
||||
else:
|
||||
self.linear = nn.Linear(
|
||||
hidden_size,
|
||||
patch_size[0] * patch_size[1] * patch_size[2] * out_channels,
|
||||
bias=True,
|
||||
)
|
||||
nn.init.zeros_(self.linear.weight)
|
||||
nn.init.zeros_(self.linear.bias)
|
||||
|
||||
# Here we don't distinguish between the modulate types. Just use the simple one.
|
||||
self.adaLN_modulation = nn.Sequential(
|
||||
act_layer(),
|
||||
nn.Linear(hidden_size, 2 * hidden_size, bias=True, **factory_kwargs),
|
||||
)
|
||||
# Zero-initialize the modulation
|
||||
nn.init.zeros_(self.adaLN_modulation[1].weight)
|
||||
nn.init.zeros_(self.adaLN_modulation[1].bias)
|
||||
|
||||
def forward(self, x, c):
|
||||
shift, scale = self.adaLN_modulation(c).chunk(2, dim=1)
|
||||
x = modulate_(self.norm_final(x), shift=shift, scale=scale)
|
||||
x = self.linear(x)
|
||||
return x
|
||||
1020
hyvideo/modules/models.py
Normal file
1020
hyvideo/modules/models.py
Normal file
File diff suppressed because it is too large
Load Diff
136
hyvideo/modules/modulate_layers.py
Normal file
136
hyvideo/modules/modulate_layers.py
Normal file
@@ -0,0 +1,136 @@
|
||||
from typing import Callable
|
||||
|
||||
import torch
|
||||
import torch.nn as nn
|
||||
import math
|
||||
|
||||
class ModulateDiT(nn.Module):
|
||||
"""Modulation layer for DiT."""
|
||||
def __init__(
|
||||
self,
|
||||
hidden_size: int,
|
||||
factor: int,
|
||||
act_layer: Callable,
|
||||
dtype=None,
|
||||
device=None,
|
||||
):
|
||||
factory_kwargs = {"dtype": dtype, "device": device}
|
||||
super().__init__()
|
||||
self.act = act_layer()
|
||||
self.linear = nn.Linear(
|
||||
hidden_size, factor * hidden_size, bias=True, **factory_kwargs
|
||||
)
|
||||
# Zero-initialize the modulation
|
||||
nn.init.zeros_(self.linear.weight)
|
||||
nn.init.zeros_(self.linear.bias)
|
||||
|
||||
def forward(self, x: torch.Tensor, condition_type=None, token_replace_vec=None) -> torch.Tensor:
|
||||
x_out = self.linear(self.act(x))
|
||||
|
||||
if condition_type == "token_replace":
|
||||
x_token_replace_out = self.linear(self.act(token_replace_vec))
|
||||
return x_out, x_token_replace_out
|
||||
else:
|
||||
return x_out
|
||||
|
||||
def modulate(x, shift=None, scale=None):
|
||||
"""modulate by shift and scale
|
||||
|
||||
Args:
|
||||
x (torch.Tensor): input tensor.
|
||||
shift (torch.Tensor, optional): shift tensor. Defaults to None.
|
||||
scale (torch.Tensor, optional): scale tensor. Defaults to None.
|
||||
|
||||
Returns:
|
||||
torch.Tensor: the output tensor after modulate.
|
||||
"""
|
||||
if scale is None and shift is None:
|
||||
return x
|
||||
elif shift is None:
|
||||
return x * (1 + scale.unsqueeze(1))
|
||||
elif scale is None:
|
||||
return x + shift.unsqueeze(1)
|
||||
else:
|
||||
return x * (1 + scale.unsqueeze(1)) + shift.unsqueeze(1)
|
||||
|
||||
def modulate_(x, shift=None, scale=None):
|
||||
|
||||
if scale is None and shift is None:
|
||||
return x
|
||||
elif shift is None:
|
||||
scale = scale + 1
|
||||
scale = scale.unsqueeze(1)
|
||||
return x.mul_(scale)
|
||||
elif scale is None:
|
||||
return x + shift.unsqueeze(1)
|
||||
else:
|
||||
scale = scale + 1
|
||||
scale = scale.unsqueeze(1)
|
||||
# return x * (1 + scale.unsqueeze(1)) + shift.unsqueeze(1)
|
||||
torch.addcmul(shift.unsqueeze(1), x, scale, out =x )
|
||||
return x
|
||||
|
||||
def modulate(x, shift=None, scale=None, condition_type=None,
|
||||
tr_shift=None, tr_scale=None,
|
||||
frist_frame_token_num=None):
|
||||
if condition_type == "token_replace":
|
||||
x_zero = x[:, :frist_frame_token_num] * (1 + tr_scale.unsqueeze(1)) + tr_shift.unsqueeze(1)
|
||||
x_orig = x[:, frist_frame_token_num:] * (1 + scale.unsqueeze(1)) + shift.unsqueeze(1)
|
||||
x = torch.concat((x_zero, x_orig), dim=1)
|
||||
return x
|
||||
else:
|
||||
if scale is None and shift is None:
|
||||
return x
|
||||
elif shift is None:
|
||||
return x * (1 + scale.unsqueeze(1))
|
||||
elif scale is None:
|
||||
return x + shift.unsqueeze(1)
|
||||
else:
|
||||
return x * (1 + scale.unsqueeze(1)) + shift.unsqueeze(1)
|
||||
|
||||
def apply_gate(x, gate=None, tanh=False, condition_type=None, tr_gate=None, frist_frame_token_num=None):
|
||||
"""AI is creating summary for apply_gate
|
||||
|
||||
Args:
|
||||
x (torch.Tensor): input tensor.
|
||||
gate (torch.Tensor, optional): gate tensor. Defaults to None.
|
||||
tanh (bool, optional): whether to use tanh function. Defaults to False.
|
||||
|
||||
Returns:
|
||||
torch.Tensor: the output tensor after apply gate.
|
||||
"""
|
||||
if condition_type == "token_replace":
|
||||
if gate is None:
|
||||
return x
|
||||
if tanh:
|
||||
x_zero = x[:, :frist_frame_token_num] * tr_gate.unsqueeze(1).tanh()
|
||||
x_orig = x[:, frist_frame_token_num:] * gate.unsqueeze(1).tanh()
|
||||
x = torch.concat((x_zero, x_orig), dim=1)
|
||||
return x
|
||||
else:
|
||||
x_zero = x[:, :frist_frame_token_num] * tr_gate.unsqueeze(1)
|
||||
x_orig = x[:, frist_frame_token_num:] * gate.unsqueeze(1)
|
||||
x = torch.concat((x_zero, x_orig), dim=1)
|
||||
return x
|
||||
else:
|
||||
if gate is None:
|
||||
return x
|
||||
if tanh:
|
||||
return x * gate.unsqueeze(1).tanh()
|
||||
else:
|
||||
return x * gate.unsqueeze(1)
|
||||
|
||||
def apply_gate_and_accumulate_(accumulator, x, gate=None, tanh=False):
|
||||
if gate is None:
|
||||
return accumulator
|
||||
if tanh:
|
||||
return accumulator.addcmul_(x, gate.unsqueeze(1).tanh())
|
||||
else:
|
||||
return accumulator.addcmul_(x, gate.unsqueeze(1))
|
||||
|
||||
def ckpt_wrapper(module):
|
||||
def ckpt_forward(*inputs):
|
||||
outputs = module(*inputs)
|
||||
return outputs
|
||||
|
||||
return ckpt_forward
|
||||
88
hyvideo/modules/norm_layers.py
Normal file
88
hyvideo/modules/norm_layers.py
Normal file
@@ -0,0 +1,88 @@
|
||||
import torch
|
||||
import torch.nn as nn
|
||||
|
||||
|
||||
class RMSNorm(nn.Module):
|
||||
def __init__(
|
||||
self,
|
||||
dim: int,
|
||||
elementwise_affine=True,
|
||||
eps: float = 1e-6,
|
||||
device=None,
|
||||
dtype=None,
|
||||
):
|
||||
"""
|
||||
Initialize the RMSNorm normalization layer.
|
||||
|
||||
Args:
|
||||
dim (int): The dimension of the input tensor.
|
||||
eps (float, optional): A small value added to the denominator for numerical stability. Default is 1e-6.
|
||||
|
||||
Attributes:
|
||||
eps (float): A small value added to the denominator for numerical stability.
|
||||
weight (nn.Parameter): Learnable scaling parameter.
|
||||
|
||||
"""
|
||||
factory_kwargs = {"device": device, "dtype": dtype}
|
||||
super().__init__()
|
||||
self.eps = eps
|
||||
if elementwise_affine:
|
||||
self.weight = nn.Parameter(torch.ones(dim, **factory_kwargs))
|
||||
|
||||
def _norm(self, x):
|
||||
"""
|
||||
Apply the RMSNorm normalization to the input tensor.
|
||||
|
||||
Args:
|
||||
x (torch.Tensor): The input tensor.
|
||||
|
||||
Returns:
|
||||
torch.Tensor: The normalized tensor.
|
||||
|
||||
"""
|
||||
|
||||
return x * torch.rsqrt(x.pow(2).mean(-1, keepdim=True) + self.eps)
|
||||
|
||||
def forward(self, x):
|
||||
"""
|
||||
Forward pass through the RMSNorm layer.
|
||||
|
||||
Args:
|
||||
x (torch.Tensor): The input tensor.
|
||||
|
||||
Returns:
|
||||
torch.Tensor: The output tensor after applying RMSNorm.
|
||||
|
||||
"""
|
||||
output = self._norm(x.float()).type_as(x)
|
||||
if hasattr(self, "weight"):
|
||||
output = output * self.weight
|
||||
return output
|
||||
|
||||
def apply_(self, x):
|
||||
y = x.pow(2).mean(-1, keepdim=True)
|
||||
y.add_(self.eps)
|
||||
y.rsqrt_()
|
||||
x.mul_(y)
|
||||
del y
|
||||
if hasattr(self, "weight"):
|
||||
x.mul_(self.weight)
|
||||
return x
|
||||
|
||||
|
||||
def get_norm_layer(norm_layer):
|
||||
"""
|
||||
Get the normalization layer.
|
||||
|
||||
Args:
|
||||
norm_layer (str): The type of normalization layer.
|
||||
|
||||
Returns:
|
||||
norm_layer (nn.Module): The normalization layer.
|
||||
"""
|
||||
if norm_layer == "layer":
|
||||
return nn.LayerNorm
|
||||
elif norm_layer == "rms":
|
||||
return RMSNorm
|
||||
else:
|
||||
raise NotImplementedError(f"Norm layer {norm_layer} is not implemented")
|
||||
760
hyvideo/modules/original models.py
Normal file
760
hyvideo/modules/original models.py
Normal file
@@ -0,0 +1,760 @@
|
||||
from typing import Any, List, Tuple, Optional, Union, Dict
|
||||
from einops import rearrange
|
||||
|
||||
import torch
|
||||
import torch.nn as nn
|
||||
import torch.nn.functional as F
|
||||
|
||||
from diffusers.models import ModelMixin
|
||||
from diffusers.configuration_utils import ConfigMixin, register_to_config
|
||||
|
||||
from .activation_layers import get_activation_layer
|
||||
from .norm_layers import get_norm_layer
|
||||
from .embed_layers import TimestepEmbedder, PatchEmbed, TextProjection
|
||||
from .attenion import attention, parallel_attention, get_cu_seqlens
|
||||
from .posemb_layers import apply_rotary_emb
|
||||
from .mlp_layers import MLP, MLPEmbedder, FinalLayer
|
||||
from .modulate_layers import ModulateDiT, modulate, apply_gate
|
||||
from .token_refiner import SingleTokenRefiner
|
||||
|
||||
|
||||
class MMDoubleStreamBlock(nn.Module):
|
||||
"""
|
||||
A multimodal dit block with seperate modulation for
|
||||
text and image/video, see more details (SD3): https://arxiv.org/abs/2403.03206
|
||||
(Flux.1): https://github.com/black-forest-labs/flux
|
||||
"""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
hidden_size: int,
|
||||
heads_num: int,
|
||||
mlp_width_ratio: float,
|
||||
mlp_act_type: str = "gelu_tanh",
|
||||
qk_norm: bool = True,
|
||||
qk_norm_type: str = "rms",
|
||||
qkv_bias: bool = False,
|
||||
dtype: Optional[torch.dtype] = None,
|
||||
device: Optional[torch.device] = None,
|
||||
):
|
||||
factory_kwargs = {"device": device, "dtype": dtype}
|
||||
super().__init__()
|
||||
|
||||
self.deterministic = False
|
||||
self.heads_num = heads_num
|
||||
head_dim = hidden_size // heads_num
|
||||
mlp_hidden_dim = int(hidden_size * mlp_width_ratio)
|
||||
|
||||
self.img_mod = ModulateDiT(
|
||||
hidden_size,
|
||||
factor=6,
|
||||
act_layer=get_activation_layer("silu"),
|
||||
**factory_kwargs,
|
||||
)
|
||||
self.img_norm1 = nn.LayerNorm(
|
||||
hidden_size, elementwise_affine=False, eps=1e-6, **factory_kwargs
|
||||
)
|
||||
|
||||
self.img_attn_qkv = nn.Linear(
|
||||
hidden_size, hidden_size * 3, bias=qkv_bias, **factory_kwargs
|
||||
)
|
||||
qk_norm_layer = get_norm_layer(qk_norm_type)
|
||||
self.img_attn_q_norm = (
|
||||
qk_norm_layer(head_dim, elementwise_affine=True, eps=1e-6, **factory_kwargs)
|
||||
if qk_norm
|
||||
else nn.Identity()
|
||||
)
|
||||
self.img_attn_k_norm = (
|
||||
qk_norm_layer(head_dim, elementwise_affine=True, eps=1e-6, **factory_kwargs)
|
||||
if qk_norm
|
||||
else nn.Identity()
|
||||
)
|
||||
self.img_attn_proj = nn.Linear(
|
||||
hidden_size, hidden_size, bias=qkv_bias, **factory_kwargs
|
||||
)
|
||||
|
||||
self.img_norm2 = nn.LayerNorm(
|
||||
hidden_size, elementwise_affine=False, eps=1e-6, **factory_kwargs
|
||||
)
|
||||
self.img_mlp = MLP(
|
||||
hidden_size,
|
||||
mlp_hidden_dim,
|
||||
act_layer=get_activation_layer(mlp_act_type),
|
||||
bias=True,
|
||||
**factory_kwargs,
|
||||
)
|
||||
|
||||
self.txt_mod = ModulateDiT(
|
||||
hidden_size,
|
||||
factor=6,
|
||||
act_layer=get_activation_layer("silu"),
|
||||
**factory_kwargs,
|
||||
)
|
||||
self.txt_norm1 = nn.LayerNorm(
|
||||
hidden_size, elementwise_affine=False, eps=1e-6, **factory_kwargs
|
||||
)
|
||||
|
||||
self.txt_attn_qkv = nn.Linear(
|
||||
hidden_size, hidden_size * 3, bias=qkv_bias, **factory_kwargs
|
||||
)
|
||||
self.txt_attn_q_norm = (
|
||||
qk_norm_layer(head_dim, elementwise_affine=True, eps=1e-6, **factory_kwargs)
|
||||
if qk_norm
|
||||
else nn.Identity()
|
||||
)
|
||||
self.txt_attn_k_norm = (
|
||||
qk_norm_layer(head_dim, elementwise_affine=True, eps=1e-6, **factory_kwargs)
|
||||
if qk_norm
|
||||
else nn.Identity()
|
||||
)
|
||||
self.txt_attn_proj = nn.Linear(
|
||||
hidden_size, hidden_size, bias=qkv_bias, **factory_kwargs
|
||||
)
|
||||
|
||||
self.txt_norm2 = nn.LayerNorm(
|
||||
hidden_size, elementwise_affine=False, eps=1e-6, **factory_kwargs
|
||||
)
|
||||
self.txt_mlp = MLP(
|
||||
hidden_size,
|
||||
mlp_hidden_dim,
|
||||
act_layer=get_activation_layer(mlp_act_type),
|
||||
bias=True,
|
||||
**factory_kwargs,
|
||||
)
|
||||
self.hybrid_seq_parallel_attn = None
|
||||
|
||||
def enable_deterministic(self):
|
||||
self.deterministic = True
|
||||
|
||||
def disable_deterministic(self):
|
||||
self.deterministic = False
|
||||
|
||||
def forward(
|
||||
self,
|
||||
img: torch.Tensor,
|
||||
txt: torch.Tensor,
|
||||
vec: torch.Tensor,
|
||||
cu_seqlens_q: Optional[torch.Tensor] = None,
|
||||
cu_seqlens_kv: Optional[torch.Tensor] = None,
|
||||
max_seqlen_q: Optional[int] = None,
|
||||
max_seqlen_kv: Optional[int] = None,
|
||||
freqs_cis: tuple = None,
|
||||
) -> Tuple[torch.Tensor, torch.Tensor]:
|
||||
(
|
||||
img_mod1_shift,
|
||||
img_mod1_scale,
|
||||
img_mod1_gate,
|
||||
img_mod2_shift,
|
||||
img_mod2_scale,
|
||||
img_mod2_gate,
|
||||
) = self.img_mod(vec).chunk(6, dim=-1)
|
||||
(
|
||||
txt_mod1_shift,
|
||||
txt_mod1_scale,
|
||||
txt_mod1_gate,
|
||||
txt_mod2_shift,
|
||||
txt_mod2_scale,
|
||||
txt_mod2_gate,
|
||||
) = self.txt_mod(vec).chunk(6, dim=-1)
|
||||
|
||||
# Prepare image for attention.
|
||||
img_modulated = self.img_norm1(img)
|
||||
img_modulated = modulate(
|
||||
img_modulated, shift=img_mod1_shift, scale=img_mod1_scale
|
||||
)
|
||||
img_qkv = self.img_attn_qkv(img_modulated)
|
||||
img_q, img_k, img_v = rearrange(
|
||||
img_qkv, "B L (K H D) -> K B L H D", K=3, H=self.heads_num
|
||||
)
|
||||
# Apply QK-Norm if needed
|
||||
img_q = self.img_attn_q_norm(img_q).to(img_v)
|
||||
img_k = self.img_attn_k_norm(img_k).to(img_v)
|
||||
|
||||
# Apply RoPE if needed.
|
||||
if freqs_cis is not None:
|
||||
img_qq, img_kk = apply_rotary_emb(img_q, img_k, freqs_cis, head_first=False)
|
||||
assert (
|
||||
img_qq.shape == img_q.shape and img_kk.shape == img_k.shape
|
||||
), f"img_kk: {img_qq.shape}, img_q: {img_q.shape}, img_kk: {img_kk.shape}, img_k: {img_k.shape}"
|
||||
img_q, img_k = img_qq, img_kk
|
||||
|
||||
# Prepare txt for attention.
|
||||
txt_modulated = self.txt_norm1(txt)
|
||||
txt_modulated = modulate(
|
||||
txt_modulated, shift=txt_mod1_shift, scale=txt_mod1_scale
|
||||
)
|
||||
txt_qkv = self.txt_attn_qkv(txt_modulated)
|
||||
txt_q, txt_k, txt_v = rearrange(
|
||||
txt_qkv, "B L (K H D) -> K B L H D", K=3, H=self.heads_num
|
||||
)
|
||||
# Apply QK-Norm if needed.
|
||||
txt_q = self.txt_attn_q_norm(txt_q).to(txt_v)
|
||||
txt_k = self.txt_attn_k_norm(txt_k).to(txt_v)
|
||||
|
||||
# Run actual attention.
|
||||
q = torch.cat((img_q, txt_q), dim=1)
|
||||
k = torch.cat((img_k, txt_k), dim=1)
|
||||
v = torch.cat((img_v, txt_v), dim=1)
|
||||
assert (
|
||||
cu_seqlens_q.shape[0] == 2 * img.shape[0] + 1
|
||||
), f"cu_seqlens_q.shape:{cu_seqlens_q.shape}, img.shape[0]:{img.shape[0]}"
|
||||
|
||||
# attention computation start
|
||||
if not self.hybrid_seq_parallel_attn:
|
||||
attn = attention(
|
||||
q,
|
||||
k,
|
||||
v,
|
||||
cu_seqlens_q=cu_seqlens_q,
|
||||
cu_seqlens_kv=cu_seqlens_kv,
|
||||
max_seqlen_q=max_seqlen_q,
|
||||
max_seqlen_kv=max_seqlen_kv,
|
||||
batch_size=img_k.shape[0],
|
||||
)
|
||||
else:
|
||||
attn = parallel_attention(
|
||||
self.hybrid_seq_parallel_attn,
|
||||
q,
|
||||
k,
|
||||
v,
|
||||
img_q_len=img_q.shape[1],
|
||||
img_kv_len=img_k.shape[1],
|
||||
cu_seqlens_q=cu_seqlens_q,
|
||||
cu_seqlens_kv=cu_seqlens_kv
|
||||
)
|
||||
|
||||
# attention computation end
|
||||
|
||||
img_attn, txt_attn = attn[:, : img.shape[1]], attn[:, img.shape[1] :]
|
||||
|
||||
# Calculate the img bloks.
|
||||
img = img + apply_gate(self.img_attn_proj(img_attn), gate=img_mod1_gate)
|
||||
img = img + apply_gate(
|
||||
self.img_mlp(
|
||||
modulate(
|
||||
self.img_norm2(img), shift=img_mod2_shift, scale=img_mod2_scale
|
||||
)
|
||||
),
|
||||
gate=img_mod2_gate,
|
||||
)
|
||||
|
||||
# Calculate the txt bloks.
|
||||
txt = txt + apply_gate(self.txt_attn_proj(txt_attn), gate=txt_mod1_gate)
|
||||
txt = txt + apply_gate(
|
||||
self.txt_mlp(
|
||||
modulate(
|
||||
self.txt_norm2(txt), shift=txt_mod2_shift, scale=txt_mod2_scale
|
||||
)
|
||||
),
|
||||
gate=txt_mod2_gate,
|
||||
)
|
||||
|
||||
return img, txt
|
||||
|
||||
|
||||
class MMSingleStreamBlock(nn.Module):
|
||||
"""
|
||||
A DiT block with parallel linear layers as described in
|
||||
https://arxiv.org/abs/2302.05442 and adapted modulation interface.
|
||||
Also refer to (SD3): https://arxiv.org/abs/2403.03206
|
||||
(Flux.1): https://github.com/black-forest-labs/flux
|
||||
"""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
hidden_size: int,
|
||||
heads_num: int,
|
||||
mlp_width_ratio: float = 4.0,
|
||||
mlp_act_type: str = "gelu_tanh",
|
||||
qk_norm: bool = True,
|
||||
qk_norm_type: str = "rms",
|
||||
qk_scale: float = None,
|
||||
dtype: Optional[torch.dtype] = None,
|
||||
device: Optional[torch.device] = None,
|
||||
):
|
||||
factory_kwargs = {"device": device, "dtype": dtype}
|
||||
super().__init__()
|
||||
|
||||
self.deterministic = False
|
||||
self.hidden_size = hidden_size
|
||||
self.heads_num = heads_num
|
||||
head_dim = hidden_size // heads_num
|
||||
mlp_hidden_dim = int(hidden_size * mlp_width_ratio)
|
||||
self.mlp_hidden_dim = mlp_hidden_dim
|
||||
self.scale = qk_scale or head_dim ** -0.5
|
||||
|
||||
# qkv and mlp_in
|
||||
self.linear1 = nn.Linear(
|
||||
hidden_size, hidden_size * 3 + mlp_hidden_dim, **factory_kwargs
|
||||
)
|
||||
# proj and mlp_out
|
||||
self.linear2 = nn.Linear(
|
||||
hidden_size + mlp_hidden_dim, hidden_size, **factory_kwargs
|
||||
)
|
||||
|
||||
qk_norm_layer = get_norm_layer(qk_norm_type)
|
||||
self.q_norm = (
|
||||
qk_norm_layer(head_dim, elementwise_affine=True, eps=1e-6, **factory_kwargs)
|
||||
if qk_norm
|
||||
else nn.Identity()
|
||||
)
|
||||
self.k_norm = (
|
||||
qk_norm_layer(head_dim, elementwise_affine=True, eps=1e-6, **factory_kwargs)
|
||||
if qk_norm
|
||||
else nn.Identity()
|
||||
)
|
||||
|
||||
self.pre_norm = nn.LayerNorm(
|
||||
hidden_size, elementwise_affine=False, eps=1e-6, **factory_kwargs
|
||||
)
|
||||
|
||||
self.mlp_act = get_activation_layer(mlp_act_type)()
|
||||
self.modulation = ModulateDiT(
|
||||
hidden_size,
|
||||
factor=3,
|
||||
act_layer=get_activation_layer("silu"),
|
||||
**factory_kwargs,
|
||||
)
|
||||
self.hybrid_seq_parallel_attn = None
|
||||
|
||||
def enable_deterministic(self):
|
||||
self.deterministic = True
|
||||
|
||||
def disable_deterministic(self):
|
||||
self.deterministic = False
|
||||
|
||||
def forward(
|
||||
self,
|
||||
x: torch.Tensor,
|
||||
vec: torch.Tensor,
|
||||
txt_len: int,
|
||||
cu_seqlens_q: Optional[torch.Tensor] = None,
|
||||
cu_seqlens_kv: Optional[torch.Tensor] = None,
|
||||
max_seqlen_q: Optional[int] = None,
|
||||
max_seqlen_kv: Optional[int] = None,
|
||||
freqs_cis: Tuple[torch.Tensor, torch.Tensor] = None,
|
||||
) -> torch.Tensor:
|
||||
mod_shift, mod_scale, mod_gate = self.modulation(vec).chunk(3, dim=-1)
|
||||
x_mod = modulate(self.pre_norm(x), shift=mod_shift, scale=mod_scale)
|
||||
qkv, mlp = torch.split(
|
||||
self.linear1(x_mod), [3 * self.hidden_size, self.mlp_hidden_dim], dim=-1
|
||||
)
|
||||
|
||||
q, k, v = rearrange(qkv, "B L (K H D) -> K B L H D", K=3, H=self.heads_num)
|
||||
|
||||
# Apply QK-Norm if needed.
|
||||
q = self.q_norm(q).to(v)
|
||||
k = self.k_norm(k).to(v)
|
||||
|
||||
# Apply RoPE if needed.
|
||||
if freqs_cis is not None:
|
||||
img_q, txt_q = q[:, :-txt_len, :, :], q[:, -txt_len:, :, :]
|
||||
img_k, txt_k = k[:, :-txt_len, :, :], k[:, -txt_len:, :, :]
|
||||
img_qq, img_kk = apply_rotary_emb(img_q, img_k, freqs_cis, head_first=False)
|
||||
assert (
|
||||
img_qq.shape == img_q.shape and img_kk.shape == img_k.shape
|
||||
), f"img_kk: {img_qq.shape}, img_q: {img_q.shape}, img_kk: {img_kk.shape}, img_k: {img_k.shape}"
|
||||
img_q, img_k = img_qq, img_kk
|
||||
q = torch.cat((img_q, txt_q), dim=1)
|
||||
k = torch.cat((img_k, txt_k), dim=1)
|
||||
|
||||
# Compute attention.
|
||||
assert (
|
||||
cu_seqlens_q.shape[0] == 2 * x.shape[0] + 1
|
||||
), f"cu_seqlens_q.shape:{cu_seqlens_q.shape}, x.shape[0]:{x.shape[0]}"
|
||||
|
||||
# attention computation start
|
||||
if not self.hybrid_seq_parallel_attn:
|
||||
attn = attention(
|
||||
q,
|
||||
k,
|
||||
v,
|
||||
cu_seqlens_q=cu_seqlens_q,
|
||||
cu_seqlens_kv=cu_seqlens_kv,
|
||||
max_seqlen_q=max_seqlen_q,
|
||||
max_seqlen_kv=max_seqlen_kv,
|
||||
batch_size=x.shape[0],
|
||||
)
|
||||
else:
|
||||
attn = parallel_attention(
|
||||
self.hybrid_seq_parallel_attn,
|
||||
q,
|
||||
k,
|
||||
v,
|
||||
img_q_len=img_q.shape[1],
|
||||
img_kv_len=img_k.shape[1],
|
||||
cu_seqlens_q=cu_seqlens_q,
|
||||
cu_seqlens_kv=cu_seqlens_kv
|
||||
)
|
||||
# attention computation end
|
||||
|
||||
# Compute activation in mlp stream, cat again and run second linear layer.
|
||||
output = self.linear2(torch.cat((attn, self.mlp_act(mlp)), 2))
|
||||
return x + apply_gate(output, gate=mod_gate)
|
||||
|
||||
|
||||
class HYVideoDiffusionTransformer(ModelMixin, ConfigMixin):
|
||||
"""
|
||||
HunyuanVideo Transformer backbone
|
||||
|
||||
Inherited from ModelMixin and ConfigMixin for compatibility with diffusers' sampler StableDiffusionPipeline.
|
||||
|
||||
Reference:
|
||||
[1] Flux.1: https://github.com/black-forest-labs/flux
|
||||
[2] MMDiT: http://arxiv.org/abs/2403.03206
|
||||
|
||||
Parameters
|
||||
----------
|
||||
args: argparse.Namespace
|
||||
The arguments parsed by argparse.
|
||||
patch_size: list
|
||||
The size of the patch.
|
||||
in_channels: int
|
||||
The number of input channels.
|
||||
out_channels: int
|
||||
The number of output channels.
|
||||
hidden_size: int
|
||||
The hidden size of the transformer backbone.
|
||||
heads_num: int
|
||||
The number of attention heads.
|
||||
mlp_width_ratio: float
|
||||
The ratio of the hidden size of the MLP in the transformer block.
|
||||
mlp_act_type: str
|
||||
The activation function of the MLP in the transformer block.
|
||||
depth_double_blocks: int
|
||||
The number of transformer blocks in the double blocks.
|
||||
depth_single_blocks: int
|
||||
The number of transformer blocks in the single blocks.
|
||||
rope_dim_list: list
|
||||
The dimension of the rotary embedding for t, h, w.
|
||||
qkv_bias: bool
|
||||
Whether to use bias in the qkv linear layer.
|
||||
qk_norm: bool
|
||||
Whether to use qk norm.
|
||||
qk_norm_type: str
|
||||
The type of qk norm.
|
||||
guidance_embed: bool
|
||||
Whether to use guidance embedding for distillation.
|
||||
text_projection: str
|
||||
The type of the text projection, default is single_refiner.
|
||||
use_attention_mask: bool
|
||||
Whether to use attention mask for text encoder.
|
||||
dtype: torch.dtype
|
||||
The dtype of the model.
|
||||
device: torch.device
|
||||
The device of the model.
|
||||
"""
|
||||
|
||||
@register_to_config
|
||||
def __init__(
|
||||
self,
|
||||
args: Any,
|
||||
patch_size: list = [1, 2, 2],
|
||||
in_channels: int = 4, # Should be VAE.config.latent_channels.
|
||||
out_channels: int = None,
|
||||
hidden_size: int = 3072,
|
||||
heads_num: int = 24,
|
||||
mlp_width_ratio: float = 4.0,
|
||||
mlp_act_type: str = "gelu_tanh",
|
||||
mm_double_blocks_depth: int = 20,
|
||||
mm_single_blocks_depth: int = 40,
|
||||
rope_dim_list: List[int] = [16, 56, 56],
|
||||
qkv_bias: bool = True,
|
||||
qk_norm: bool = True,
|
||||
qk_norm_type: str = "rms",
|
||||
guidance_embed: bool = False, # For modulation.
|
||||
text_projection: str = "single_refiner",
|
||||
use_attention_mask: bool = True,
|
||||
dtype: Optional[torch.dtype] = None,
|
||||
device: Optional[torch.device] = None,
|
||||
):
|
||||
factory_kwargs = {"device": device, "dtype": dtype}
|
||||
super().__init__()
|
||||
|
||||
self.patch_size = patch_size
|
||||
self.in_channels = in_channels
|
||||
self.out_channels = in_channels if out_channels is None else out_channels
|
||||
self.unpatchify_channels = self.out_channels
|
||||
self.guidance_embed = guidance_embed
|
||||
self.rope_dim_list = rope_dim_list
|
||||
|
||||
# Text projection. Default to linear projection.
|
||||
# Alternative: TokenRefiner. See more details (LI-DiT): http://arxiv.org/abs/2406.11831
|
||||
self.use_attention_mask = use_attention_mask
|
||||
self.text_projection = text_projection
|
||||
|
||||
self.text_states_dim = args.text_states_dim
|
||||
self.text_states_dim_2 = args.text_states_dim_2
|
||||
|
||||
if hidden_size % heads_num != 0:
|
||||
raise ValueError(
|
||||
f"Hidden size {hidden_size} must be divisible by heads_num {heads_num}"
|
||||
)
|
||||
pe_dim = hidden_size // heads_num
|
||||
if sum(rope_dim_list) != pe_dim:
|
||||
raise ValueError(
|
||||
f"Got {rope_dim_list} but expected positional dim {pe_dim}"
|
||||
)
|
||||
self.hidden_size = hidden_size
|
||||
self.heads_num = heads_num
|
||||
|
||||
# image projection
|
||||
self.img_in = PatchEmbed(
|
||||
self.patch_size, self.in_channels, self.hidden_size, **factory_kwargs
|
||||
)
|
||||
|
||||
# text projection
|
||||
if self.text_projection == "linear":
|
||||
self.txt_in = TextProjection(
|
||||
self.text_states_dim,
|
||||
self.hidden_size,
|
||||
get_activation_layer("silu"),
|
||||
**factory_kwargs,
|
||||
)
|
||||
elif self.text_projection == "single_refiner":
|
||||
self.txt_in = SingleTokenRefiner(
|
||||
self.text_states_dim, hidden_size, heads_num, depth=2, **factory_kwargs
|
||||
)
|
||||
else:
|
||||
raise NotImplementedError(
|
||||
f"Unsupported text_projection: {self.text_projection}"
|
||||
)
|
||||
|
||||
# time modulation
|
||||
self.time_in = TimestepEmbedder(
|
||||
self.hidden_size, get_activation_layer("silu"), **factory_kwargs
|
||||
)
|
||||
|
||||
# text modulation
|
||||
self.vector_in = MLPEmbedder(
|
||||
self.text_states_dim_2, self.hidden_size, **factory_kwargs
|
||||
)
|
||||
|
||||
# guidance modulation
|
||||
self.guidance_in = (
|
||||
TimestepEmbedder(
|
||||
self.hidden_size, get_activation_layer("silu"), **factory_kwargs
|
||||
)
|
||||
if guidance_embed
|
||||
else None
|
||||
)
|
||||
|
||||
# double blocks
|
||||
self.double_blocks = nn.ModuleList(
|
||||
[
|
||||
MMDoubleStreamBlock(
|
||||
self.hidden_size,
|
||||
self.heads_num,
|
||||
mlp_width_ratio=mlp_width_ratio,
|
||||
mlp_act_type=mlp_act_type,
|
||||
qk_norm=qk_norm,
|
||||
qk_norm_type=qk_norm_type,
|
||||
qkv_bias=qkv_bias,
|
||||
**factory_kwargs,
|
||||
)
|
||||
for _ in range(mm_double_blocks_depth)
|
||||
]
|
||||
)
|
||||
|
||||
# single blocks
|
||||
self.single_blocks = nn.ModuleList(
|
||||
[
|
||||
MMSingleStreamBlock(
|
||||
self.hidden_size,
|
||||
self.heads_num,
|
||||
mlp_width_ratio=mlp_width_ratio,
|
||||
mlp_act_type=mlp_act_type,
|
||||
qk_norm=qk_norm,
|
||||
qk_norm_type=qk_norm_type,
|
||||
**factory_kwargs,
|
||||
)
|
||||
for _ in range(mm_single_blocks_depth)
|
||||
]
|
||||
)
|
||||
|
||||
self.final_layer = FinalLayer(
|
||||
self.hidden_size,
|
||||
self.patch_size,
|
||||
self.out_channels,
|
||||
get_activation_layer("silu"),
|
||||
**factory_kwargs,
|
||||
)
|
||||
|
||||
def enable_deterministic(self):
|
||||
for block in self.double_blocks:
|
||||
block.enable_deterministic()
|
||||
for block in self.single_blocks:
|
||||
block.enable_deterministic()
|
||||
|
||||
def disable_deterministic(self):
|
||||
for block in self.double_blocks:
|
||||
block.disable_deterministic()
|
||||
for block in self.single_blocks:
|
||||
block.disable_deterministic()
|
||||
|
||||
def forward(
|
||||
self,
|
||||
x: torch.Tensor,
|
||||
t: torch.Tensor, # Should be in range(0, 1000).
|
||||
text_states: torch.Tensor = None,
|
||||
text_mask: torch.Tensor = None, # Now we don't use it.
|
||||
text_states_2: Optional[torch.Tensor] = None, # Text embedding for modulation.
|
||||
freqs_cos: Optional[torch.Tensor] = None,
|
||||
freqs_sin: Optional[torch.Tensor] = None,
|
||||
guidance: torch.Tensor = None, # Guidance for modulation, should be cfg_scale x 1000.
|
||||
return_dict: bool = True,
|
||||
) -> Union[torch.Tensor, Dict[str, torch.Tensor]]:
|
||||
out = {}
|
||||
img = x
|
||||
txt = text_states
|
||||
_, _, ot, oh, ow = x.shape
|
||||
tt, th, tw = (
|
||||
ot // self.patch_size[0],
|
||||
oh // self.patch_size[1],
|
||||
ow // self.patch_size[2],
|
||||
)
|
||||
|
||||
# Prepare modulation vectors.
|
||||
vec = self.time_in(t)
|
||||
|
||||
# text modulation
|
||||
vec = vec + self.vector_in(text_states_2)
|
||||
|
||||
# guidance modulation
|
||||
if self.guidance_embed:
|
||||
if guidance is None:
|
||||
raise ValueError(
|
||||
"Didn't get guidance strength for guidance distilled model."
|
||||
)
|
||||
|
||||
# our timestep_embedding is merged into guidance_in(TimestepEmbedder)
|
||||
vec = vec + self.guidance_in(guidance)
|
||||
|
||||
# Embed image and text.
|
||||
img = self.img_in(img)
|
||||
if self.text_projection == "linear":
|
||||
txt = self.txt_in(txt)
|
||||
elif self.text_projection == "single_refiner":
|
||||
txt = self.txt_in(txt, t, text_mask if self.use_attention_mask else None)
|
||||
else:
|
||||
raise NotImplementedError(
|
||||
f"Unsupported text_projection: {self.text_projection}"
|
||||
)
|
||||
|
||||
txt_seq_len = txt.shape[1]
|
||||
img_seq_len = img.shape[1]
|
||||
|
||||
# Compute cu_squlens and max_seqlen for flash attention
|
||||
cu_seqlens_q = get_cu_seqlens(text_mask, img_seq_len)
|
||||
cu_seqlens_kv = cu_seqlens_q
|
||||
max_seqlen_q = img_seq_len + txt_seq_len
|
||||
max_seqlen_kv = max_seqlen_q
|
||||
|
||||
freqs_cis = (freqs_cos, freqs_sin) if freqs_cos is not None else None
|
||||
# --------------------- Pass through DiT blocks ------------------------
|
||||
for _, block in enumerate(self.double_blocks):
|
||||
double_block_args = [
|
||||
img,
|
||||
txt,
|
||||
vec,
|
||||
cu_seqlens_q,
|
||||
cu_seqlens_kv,
|
||||
max_seqlen_q,
|
||||
max_seqlen_kv,
|
||||
freqs_cis,
|
||||
]
|
||||
|
||||
img, txt = block(*double_block_args)
|
||||
|
||||
# Merge txt and img to pass through single stream blocks.
|
||||
x = torch.cat((img, txt), 1)
|
||||
if len(self.single_blocks) > 0:
|
||||
for _, block in enumerate(self.single_blocks):
|
||||
single_block_args = [
|
||||
x,
|
||||
vec,
|
||||
txt_seq_len,
|
||||
cu_seqlens_q,
|
||||
cu_seqlens_kv,
|
||||
max_seqlen_q,
|
||||
max_seqlen_kv,
|
||||
(freqs_cos, freqs_sin),
|
||||
]
|
||||
|
||||
x = block(*single_block_args)
|
||||
|
||||
img = x[:, :img_seq_len, ...]
|
||||
|
||||
# ---------------------------- Final layer ------------------------------
|
||||
img = self.final_layer(img, vec) # (N, T, patch_size ** 2 * out_channels)
|
||||
|
||||
img = self.unpatchify(img, tt, th, tw)
|
||||
if return_dict:
|
||||
out["x"] = img
|
||||
return out
|
||||
return img
|
||||
|
||||
def unpatchify(self, x, t, h, w):
|
||||
"""
|
||||
x: (N, T, patch_size**2 * C)
|
||||
imgs: (N, H, W, C)
|
||||
"""
|
||||
c = self.unpatchify_channels
|
||||
pt, ph, pw = self.patch_size
|
||||
assert t * h * w == x.shape[1]
|
||||
|
||||
x = x.reshape(shape=(x.shape[0], t, h, w, c, pt, ph, pw))
|
||||
x = torch.einsum("nthwcopq->nctohpwq", x)
|
||||
imgs = x.reshape(shape=(x.shape[0], c, t * pt, h * ph, w * pw))
|
||||
|
||||
return imgs
|
||||
|
||||
def params_count(self):
|
||||
counts = {
|
||||
"double": sum(
|
||||
[
|
||||
sum(p.numel() for p in block.img_attn_qkv.parameters())
|
||||
+ sum(p.numel() for p in block.img_attn_proj.parameters())
|
||||
+ sum(p.numel() for p in block.img_mlp.parameters())
|
||||
+ sum(p.numel() for p in block.txt_attn_qkv.parameters())
|
||||
+ sum(p.numel() for p in block.txt_attn_proj.parameters())
|
||||
+ sum(p.numel() for p in block.txt_mlp.parameters())
|
||||
for block in self.double_blocks
|
||||
]
|
||||
),
|
||||
"single": sum(
|
||||
[
|
||||
sum(p.numel() for p in block.linear1.parameters())
|
||||
+ sum(p.numel() for p in block.linear2.parameters())
|
||||
for block in self.single_blocks
|
||||
]
|
||||
),
|
||||
"total": sum(p.numel() for p in self.parameters()),
|
||||
}
|
||||
counts["attn+mlp"] = counts["double"] + counts["single"]
|
||||
return counts
|
||||
|
||||
|
||||
#################################################################################
|
||||
# HunyuanVideo Configs #
|
||||
#################################################################################
|
||||
|
||||
HUNYUAN_VIDEO_CONFIG = {
|
||||
"HYVideo-T/2": {
|
||||
"mm_double_blocks_depth": 20,
|
||||
"mm_single_blocks_depth": 40,
|
||||
"rope_dim_list": [16, 56, 56],
|
||||
"hidden_size": 3072,
|
||||
"heads_num": 24,
|
||||
"mlp_width_ratio": 4,
|
||||
},
|
||||
"HYVideo-T/2-cfgdistill": {
|
||||
"mm_double_blocks_depth": 20,
|
||||
"mm_single_blocks_depth": 40,
|
||||
"rope_dim_list": [16, 56, 56],
|
||||
"hidden_size": 3072,
|
||||
"heads_num": 24,
|
||||
"mlp_width_ratio": 4,
|
||||
"guidance_embed": True,
|
||||
},
|
||||
}
|
||||
389
hyvideo/modules/placement.py
Normal file
389
hyvideo/modules/placement.py
Normal file
@@ -0,0 +1,389 @@
|
||||
import torch
|
||||
import triton
|
||||
import triton.language as tl
|
||||
|
||||
def hunyuan_token_reorder_to_token_major(tensor, fix_len, reorder_len, reorder_num_frame, frame_size):
|
||||
"""Reorder it from frame major to token major!"""
|
||||
assert reorder_len == reorder_num_frame * frame_size
|
||||
assert tensor.shape[2] == fix_len + reorder_len
|
||||
|
||||
tensor[:, :, :-fix_len, :] = tensor[:, :, :-fix_len:, :].reshape(tensor.shape[0], tensor.shape[1], reorder_num_frame, frame_size, tensor.shape[3]) \
|
||||
.transpose(2, 3).reshape(tensor.shape[0], tensor.shape[1], reorder_len, tensor.shape[3])
|
||||
return tensor
|
||||
|
||||
def hunyuan_token_reorder_to_frame_major(tensor, fix_len, reorder_len, reorder_num_frame, frame_size):
|
||||
"""Reorder it from token major to frame major!"""
|
||||
assert reorder_len == reorder_num_frame * frame_size
|
||||
assert tensor.shape[2] == fix_len + reorder_len
|
||||
|
||||
tensor[:, :, :-fix_len:, :] = tensor[:, :, :-fix_len:, :].reshape(tensor.shape[0], tensor.shape[1], frame_size, reorder_num_frame, tensor.shape[3]) \
|
||||
.transpose(2, 3).reshape(tensor.shape[0], tensor.shape[1], reorder_len, tensor.shape[3])
|
||||
return tensor
|
||||
|
||||
|
||||
@triton.jit
|
||||
def hunyuan_sparse_head_placement_kernel(
|
||||
query_ptr, key_ptr, value_ptr, # [cfg, num_heads, seq_len, head_dim] seq_len = context_length + num_frame * frame_size
|
||||
query_out_ptr, key_out_ptr, value_out_ptr, # [cfg, num_heads, seq_len, head_dim]
|
||||
best_mask_idx_ptr, # [cfg, num_heads]
|
||||
query_stride_b, query_stride_h, query_stride_s, query_stride_d,
|
||||
mask_idx_stride_b, mask_idx_stride_h,
|
||||
seq_len: tl.constexpr,
|
||||
head_dim: tl.constexpr,
|
||||
context_length: tl.constexpr,
|
||||
num_frame: tl.constexpr,
|
||||
frame_size: tl.constexpr,
|
||||
BLOCK_SIZE: tl.constexpr
|
||||
):
|
||||
# Copy query, key, value to output
|
||||
# range: [b, h, block_id * block_size: block_id * block_size + block_size, :]
|
||||
cfg = tl.program_id(0)
|
||||
head = tl.program_id(1)
|
||||
block_id = tl.program_id(2)
|
||||
|
||||
start_id = block_id * BLOCK_SIZE
|
||||
end_id = start_id + BLOCK_SIZE
|
||||
end_id = tl.where(end_id > seq_len, seq_len, end_id)
|
||||
|
||||
# Load best mask idx (0 is spatial, 1 is temporal)
|
||||
is_temporal = tl.load(best_mask_idx_ptr + cfg * mask_idx_stride_b + head * mask_idx_stride_h)
|
||||
|
||||
offset_token = tl.arange(0, BLOCK_SIZE) + start_id
|
||||
offset_mask = offset_token < seq_len
|
||||
offset_d = tl.arange(0, head_dim)
|
||||
|
||||
if is_temporal:
|
||||
frame_id = offset_token // frame_size
|
||||
patch_id = offset_token - frame_id * frame_size
|
||||
offset_store_token = tl.where(offset_token >= seq_len - context_length, offset_token, patch_id * num_frame + frame_id)
|
||||
|
||||
offset_load = (cfg * query_stride_b + head * query_stride_h + offset_token[:,None] * query_stride_s) + offset_d[None,:] * query_stride_d
|
||||
offset_query = query_ptr + offset_load
|
||||
offset_key = key_ptr + offset_load
|
||||
offset_value = value_ptr + offset_load
|
||||
|
||||
offset_store = (cfg * query_stride_b + head * query_stride_h + offset_store_token[:,None] * query_stride_s) + offset_d[None,:] * query_stride_d
|
||||
offset_query_out = query_out_ptr + offset_store
|
||||
offset_key_out = key_out_ptr + offset_store
|
||||
offset_value_out = value_out_ptr + offset_store
|
||||
|
||||
# Maybe tune the pipeline here
|
||||
query = tl.load(offset_query, mask=offset_mask[:,None])
|
||||
tl.store(offset_query_out, query, mask=offset_mask[:,None])
|
||||
key = tl.load(offset_key, mask=offset_mask[:,None])
|
||||
tl.store(offset_key_out, key, mask=offset_mask[:,None])
|
||||
value = tl.load(offset_value, mask=offset_mask[:,None])
|
||||
tl.store(offset_value_out, value, mask=offset_mask[:,None])
|
||||
|
||||
|
||||
else:
|
||||
offset_load = (cfg * query_stride_b + head * query_stride_h + offset_token[:,None] * query_stride_s) + offset_d[None,:] * query_stride_d
|
||||
offset_query = query_ptr + offset_load
|
||||
offset_key = key_ptr + offset_load
|
||||
offset_value = value_ptr + offset_load
|
||||
|
||||
offset_store = offset_load
|
||||
offset_query_out = query_out_ptr + offset_store
|
||||
offset_key_out = key_out_ptr + offset_store
|
||||
offset_value_out = value_out_ptr + offset_store
|
||||
|
||||
# Maybe tune the pipeline here
|
||||
query = tl.load(offset_query, mask=offset_mask[:,None])
|
||||
tl.store(offset_query_out, query, mask=offset_mask[:,None])
|
||||
key = tl.load(offset_key, mask=offset_mask[:,None])
|
||||
tl.store(offset_key_out, key, mask=offset_mask[:,None])
|
||||
value = tl.load(offset_value, mask=offset_mask[:,None])
|
||||
tl.store(offset_value_out, value, mask=offset_mask[:,None])
|
||||
|
||||
|
||||
def hunyuan_sparse_head_placement(query, key, value, query_out, key_out, value_out, best_mask_idx, context_length, num_frame, frame_size):
|
||||
cfg, num_heads, seq_len, head_dim = query.shape
|
||||
BLOCK_SIZE = 128
|
||||
assert seq_len == context_length + num_frame * frame_size
|
||||
|
||||
grid = (cfg, num_heads, (seq_len + BLOCK_SIZE - 1) // BLOCK_SIZE)
|
||||
|
||||
hunyuan_sparse_head_placement_kernel[grid](
|
||||
query, key, value,
|
||||
query_out, key_out, value_out,
|
||||
best_mask_idx,
|
||||
query.stride(0), query.stride(1), query.stride(2), query.stride(3),
|
||||
best_mask_idx.stride(0), best_mask_idx.stride(1),
|
||||
seq_len, head_dim, context_length, num_frame, frame_size,
|
||||
BLOCK_SIZE
|
||||
)
|
||||
|
||||
|
||||
def ref_hunyuan_sparse_head_placement(query, key, value, best_mask_idx, context_length, num_frame, frame_size):
|
||||
cfg, num_heads, seq_len, head_dim = query.shape
|
||||
assert seq_len == context_length + num_frame * frame_size
|
||||
|
||||
query_out = query.clone()
|
||||
key_out = key.clone()
|
||||
value_out = value.clone()
|
||||
|
||||
# Spatial
|
||||
query_out[best_mask_idx == 0], key_out[best_mask_idx == 0], value_out[best_mask_idx == 0] = \
|
||||
query[best_mask_idx == 0], key[best_mask_idx == 0], value[best_mask_idx == 0]
|
||||
|
||||
# Temporal
|
||||
query_out[best_mask_idx == 1], key_out[best_mask_idx == 1], value_out[best_mask_idx == 1] = \
|
||||
hunyuan_token_reorder_to_token_major(query[best_mask_idx == 1].unsqueeze(0), context_length, num_frame * frame_size, num_frame, frame_size).squeeze(0), \
|
||||
hunyuan_token_reorder_to_token_major(key[best_mask_idx == 1].unsqueeze(0), context_length, num_frame * frame_size, num_frame, frame_size).squeeze(0), \
|
||||
hunyuan_token_reorder_to_token_major(value[best_mask_idx == 1].unsqueeze(0), context_length, num_frame * frame_size, num_frame, frame_size).squeeze(0)
|
||||
|
||||
return query_out, key_out, value_out
|
||||
|
||||
|
||||
def test_hunyuan_sparse_head_placement():
|
||||
|
||||
context_length = 226
|
||||
num_frame = 11
|
||||
frame_size = 4080
|
||||
|
||||
cfg = 2
|
||||
num_heads = 48
|
||||
|
||||
seq_len = context_length + num_frame * frame_size
|
||||
head_dim = 64
|
||||
|
||||
dtype = torch.bfloat16
|
||||
device = torch.device("cuda")
|
||||
|
||||
query = torch.randn(cfg, num_heads, seq_len, head_dim, dtype=dtype, device=device)
|
||||
key = torch.randn(cfg, num_heads, seq_len, head_dim, dtype=dtype, device=device)
|
||||
value = torch.randn(cfg, num_heads, seq_len, head_dim, dtype=dtype, device=device)
|
||||
|
||||
best_mask_idx = torch.randint(0, 2, (cfg, num_heads), device=device)
|
||||
|
||||
query_out = torch.empty_like(query)
|
||||
key_out = torch.empty_like(key)
|
||||
value_out = torch.empty_like(value)
|
||||
|
||||
hunyuan_sparse_head_placement(query, key, value, query_out, key_out, value_out, best_mask_idx, context_length, num_frame, frame_size)
|
||||
ref_query_out, ref_key_out, ref_value_out = ref_hunyuan_sparse_head_placement(query, key, value, best_mask_idx, context_length, num_frame, frame_size)
|
||||
|
||||
torch.testing.assert_close(query_out, ref_query_out)
|
||||
torch.testing.assert_close(key_out, ref_key_out)
|
||||
torch.testing.assert_close(value_out, ref_value_out)
|
||||
|
||||
|
||||
def benchmark_hunyuan_sparse_head_placement():
|
||||
import time
|
||||
|
||||
context_length = 226
|
||||
num_frame = 11
|
||||
frame_size = 4080
|
||||
|
||||
cfg = 2
|
||||
num_heads = 48
|
||||
|
||||
seq_len = context_length + num_frame * frame_size
|
||||
head_dim = 64
|
||||
|
||||
dtype = torch.bfloat16
|
||||
device = torch.device("cuda")
|
||||
|
||||
query = torch.randn(cfg, num_heads, seq_len, head_dim, dtype=dtype, device=device)
|
||||
key = torch.randn(cfg, num_heads, seq_len, head_dim, dtype=dtype, device=device)
|
||||
value = torch.randn(cfg, num_heads, seq_len, head_dim, dtype=dtype, device=device)
|
||||
best_mask_idx = torch.randint(0, 2, (cfg, num_heads), device=device)
|
||||
|
||||
query_out = torch.empty_like(query)
|
||||
key_out = torch.empty_like(key)
|
||||
value_out = torch.empty_like(value)
|
||||
|
||||
warmup = 10
|
||||
all_iter = 1000
|
||||
|
||||
# warmup
|
||||
for _ in range(warmup):
|
||||
hunyuan_sparse_head_placement(query, key, value, query_out, key_out, value_out, best_mask_idx, context_length, num_frame, frame_size)
|
||||
|
||||
torch.cuda.synchronize()
|
||||
start = time.time()
|
||||
for _ in range(all_iter):
|
||||
hunyuan_sparse_head_placement(query, key, value, query_out, key_out, value_out, best_mask_idx, context_length, num_frame, frame_size)
|
||||
torch.cuda.synchronize()
|
||||
end = time.time()
|
||||
|
||||
print(f"Triton Elapsed Time: {(end - start) / all_iter * 1e3:.2f} ms")
|
||||
print(f"Triton Total Bandwidth: {query.nelement() * query.element_size() * 3 * 2 * all_iter / (end - start) / 1e9:.2f} GB/s")
|
||||
|
||||
torch.cuda.synchronize()
|
||||
start = time.time()
|
||||
for _ in range(all_iter):
|
||||
ref_hunyuan_sparse_head_placement(query, key, value, best_mask_idx, context_length, num_frame, frame_size)
|
||||
torch.cuda.synchronize()
|
||||
end = time.time()
|
||||
|
||||
print(f"Reference Elapsed Time: {(end - start) / all_iter * 1e3:.2f} ms")
|
||||
print(f"Reference Total Bandwidth: {query.nelement() * query.element_size() * 3 * 2 * all_iter / (end - start) / 1e9:.2f} GB/s")
|
||||
|
||||
|
||||
@triton.jit
|
||||
def hunyuan_hidden_states_placement_kernel(
|
||||
hidden_states_ptr, # [cfg, num_heads, seq_len, head_dim] seq_len = context_length + num_frame * frame_size
|
||||
hidden_states_out_ptr, # [cfg, num_heads, seq_len, head_dim]
|
||||
best_mask_idx_ptr, # [cfg, num_heads]
|
||||
hidden_states_stride_b, hidden_states_stride_h, hidden_states_stride_s, hidden_states_stride_d,
|
||||
mask_idx_stride_b, mask_idx_stride_h,
|
||||
seq_len: tl.constexpr,
|
||||
head_dim: tl.constexpr,
|
||||
context_length: tl.constexpr,
|
||||
num_frame: tl.constexpr,
|
||||
frame_size: tl.constexpr,
|
||||
BLOCK_SIZE: tl.constexpr
|
||||
):
|
||||
# Copy hidden_states to output
|
||||
# range: [b, h, block_id * block_size: block_id * block_size + block_size, :]
|
||||
cfg = tl.program_id(0)
|
||||
head = tl.program_id(1)
|
||||
block_id = tl.program_id(2)
|
||||
|
||||
start_id = block_id * BLOCK_SIZE
|
||||
end_id = start_id + BLOCK_SIZE
|
||||
end_id = tl.where(end_id > seq_len, seq_len, end_id)
|
||||
|
||||
# Load best mask idx (0 is spatial, 1 is temporal)
|
||||
is_temporal = tl.load(best_mask_idx_ptr + cfg * mask_idx_stride_b + head * mask_idx_stride_h)
|
||||
|
||||
offset_token = tl.arange(0, BLOCK_SIZE) + start_id
|
||||
offset_mask = offset_token < seq_len
|
||||
offset_d = tl.arange(0, head_dim)
|
||||
|
||||
if is_temporal:
|
||||
patch_id = offset_token // num_frame
|
||||
frame_id = offset_token - patch_id * num_frame
|
||||
offset_store_token = tl.where(offset_token >= seq_len - context_length, offset_token, frame_id * frame_size + patch_id)
|
||||
|
||||
offset_load = (cfg * hidden_states_stride_b + head * hidden_states_stride_h + offset_token[:,None] * hidden_states_stride_s) + offset_d[None,:] * hidden_states_stride_d
|
||||
offset_hidden_states = hidden_states_ptr + offset_load
|
||||
|
||||
offset_store = (cfg * hidden_states_stride_b + head * hidden_states_stride_h + offset_store_token[:,None] * hidden_states_stride_s) + offset_d[None,:] * hidden_states_stride_d
|
||||
offset_hidden_states_out = hidden_states_out_ptr + offset_store
|
||||
|
||||
# Maybe tune the pipeline here
|
||||
hidden_states = tl.load(offset_hidden_states, mask=offset_mask[:,None])
|
||||
tl.store(offset_hidden_states_out, hidden_states, mask=offset_mask[:,None])
|
||||
else:
|
||||
offset_load = (cfg * hidden_states_stride_b + head * hidden_states_stride_h + offset_token[:,None] * hidden_states_stride_s) + offset_d[None,:] * hidden_states_stride_d
|
||||
offset_hidden_states = hidden_states_ptr + offset_load
|
||||
|
||||
offset_store = offset_load
|
||||
offset_hidden_states_out = hidden_states_out_ptr + offset_store
|
||||
|
||||
# Maybe tune the pipeline here
|
||||
hidden_states = tl.load(offset_hidden_states, mask=offset_mask[:,None])
|
||||
tl.store(offset_hidden_states_out, hidden_states, mask=offset_mask[:,None])
|
||||
|
||||
|
||||
def hunyuan_hidden_states_placement(hidden_states, hidden_states_out, best_mask_idx, context_length, num_frame, frame_size):
|
||||
cfg, num_heads, seq_len, head_dim = hidden_states.shape
|
||||
BLOCK_SIZE = 128
|
||||
assert seq_len == context_length + num_frame * frame_size
|
||||
|
||||
grid = (cfg, num_heads, (seq_len + BLOCK_SIZE - 1) // BLOCK_SIZE)
|
||||
|
||||
|
||||
hunyuan_hidden_states_placement_kernel[grid](
|
||||
hidden_states,
|
||||
hidden_states_out,
|
||||
best_mask_idx,
|
||||
hidden_states.stride(0), hidden_states.stride(1), hidden_states.stride(2), hidden_states.stride(3),
|
||||
best_mask_idx.stride(0), best_mask_idx.stride(1),
|
||||
seq_len, head_dim, context_length, num_frame, frame_size,
|
||||
BLOCK_SIZE
|
||||
)
|
||||
|
||||
return hidden_states_out
|
||||
|
||||
def ref_hunyuan_hidden_states_placement(hidden_states, output_hidden_states, best_mask_idx, context_length, num_frame, frame_size):
|
||||
cfg, num_heads, seq_len, head_dim = hidden_states.shape
|
||||
assert seq_len == context_length + num_frame * frame_size
|
||||
|
||||
# Spatial
|
||||
output_hidden_states[best_mask_idx == 0] = hidden_states[best_mask_idx == 0]
|
||||
# Temporal
|
||||
output_hidden_states[best_mask_idx == 1] = hunyuan_token_reorder_to_frame_major(hidden_states[best_mask_idx == 1].unsqueeze(0), context_length, num_frame * frame_size, num_frame, frame_size).squeeze(0)
|
||||
|
||||
def test_hunyuan_hidden_states_placement():
|
||||
|
||||
context_length = 226
|
||||
num_frame = 11
|
||||
frame_size = 4080
|
||||
|
||||
cfg = 2
|
||||
num_heads = 48
|
||||
|
||||
seq_len = context_length + num_frame * frame_size
|
||||
head_dim = 64
|
||||
|
||||
dtype = torch.bfloat16
|
||||
device = torch.device("cuda")
|
||||
|
||||
hidden_states = torch.randn(cfg, num_heads, seq_len, head_dim, dtype=dtype, device=device)
|
||||
best_mask_idx = torch.randint(0, 2, (cfg, num_heads), device=device)
|
||||
|
||||
hidden_states_out1 = torch.empty_like(hidden_states)
|
||||
hidden_states_out2 = torch.empty_like(hidden_states)
|
||||
|
||||
hunyuan_hidden_states_placement(hidden_states, hidden_states_out1, best_mask_idx, context_length, num_frame, frame_size)
|
||||
ref_hunyuan_hidden_states_placement(hidden_states, hidden_states_out2, best_mask_idx, context_length, num_frame, frame_size)
|
||||
|
||||
torch.testing.assert_close(hidden_states_out1, hidden_states_out2)
|
||||
|
||||
def benchmark_hunyuan_hidden_states_placement():
|
||||
import time
|
||||
|
||||
context_length = 226
|
||||
num_frame = 11
|
||||
frame_size = 4080
|
||||
|
||||
cfg = 2
|
||||
num_heads = 48
|
||||
|
||||
seq_len = context_length + num_frame * frame_size
|
||||
head_dim = 64
|
||||
|
||||
dtype = torch.bfloat16
|
||||
device = torch.device("cuda")
|
||||
|
||||
hidden_states = torch.randn(cfg, num_heads, seq_len, head_dim, dtype=dtype, device=device)
|
||||
best_mask_idx = torch.randint(0, 2, (cfg, num_heads), device=device)
|
||||
|
||||
hidden_states_out = torch.empty_like(hidden_states)
|
||||
|
||||
warmup = 10
|
||||
all_iter = 1000
|
||||
|
||||
# warmup
|
||||
for _ in range(warmup):
|
||||
hunyuan_hidden_states_placement(hidden_states, hidden_states_out, best_mask_idx, context_length, num_frame, frame_size)
|
||||
|
||||
torch.cuda.synchronize()
|
||||
start = time.time()
|
||||
for _ in range(all_iter):
|
||||
hunyuan_hidden_states_placement(hidden_states, hidden_states_out, best_mask_idx, context_length, num_frame, frame_size)
|
||||
torch.cuda.synchronize()
|
||||
end = time.time()
|
||||
|
||||
print(f"Triton Elapsed Time: {(end - start) / all_iter * 1e3:.2f} ms")
|
||||
print(f"Triton Total Bandwidth: {hidden_states.nelement() * hidden_states.element_size() * 2 * all_iter / (end - start) / 1e9:.2f} GB/s")
|
||||
|
||||
torch.cuda.synchronize()
|
||||
start = time.time()
|
||||
for _ in range(all_iter):
|
||||
ref_hunyuan_hidden_states_placement(hidden_states, hidden_states.clone(), best_mask_idx, context_length, num_frame, frame_size)
|
||||
torch.cuda.synchronize()
|
||||
end = time.time()
|
||||
|
||||
print(f"Reference Elapsed Time: {(end - start) / all_iter * 1e3:.2f} ms")
|
||||
print(f"Reference Total Bandwidth: {hidden_states.nelement() * hidden_states.element_size() * 2 * all_iter / (end - start) / 1e9:.2f} GB/s")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
test_hunyuan_sparse_head_placement()
|
||||
benchmark_hunyuan_sparse_head_placement()
|
||||
test_hunyuan_hidden_states_placement()
|
||||
benchmark_hunyuan_hidden_states_placement()
|
||||
475
hyvideo/modules/posemb_layers.py
Normal file
475
hyvideo/modules/posemb_layers.py
Normal file
@@ -0,0 +1,475 @@
|
||||
import torch
|
||||
from typing import Union, Tuple, List, Optional
|
||||
import numpy as np
|
||||
|
||||
|
||||
###### Thanks to the RifleX project (https://github.com/thu-ml/RIFLEx/) for this alternative pos embed for long videos
|
||||
#
|
||||
def get_1d_rotary_pos_embed_riflex(
|
||||
dim: int,
|
||||
pos: Union[np.ndarray, int],
|
||||
theta: float = 10000.0,
|
||||
use_real=False,
|
||||
k: Optional[int] = None,
|
||||
L_test: Optional[int] = None,
|
||||
):
|
||||
"""
|
||||
RIFLEx: Precompute the frequency tensor for complex exponentials (cis) with given dimensions.
|
||||
|
||||
This function calculates a frequency tensor with complex exponentials using the given dimension 'dim' and the end
|
||||
index 'end'. The 'theta' parameter scales the frequencies. The returned tensor contains complex values in complex64
|
||||
data type.
|
||||
|
||||
Args:
|
||||
dim (`int`): Dimension of the frequency tensor.
|
||||
pos (`np.ndarray` or `int`): Position indices for the frequency tensor. [S] or scalar
|
||||
theta (`float`, *optional*, defaults to 10000.0):
|
||||
Scaling factor for frequency computation. Defaults to 10000.0.
|
||||
use_real (`bool`, *optional*):
|
||||
If True, return real part and imaginary part separately. Otherwise, return complex numbers.
|
||||
k (`int`, *optional*, defaults to None): the index for the intrinsic frequency in RoPE
|
||||
L_test (`int`, *optional*, defaults to None): the number of frames for inference
|
||||
Returns:
|
||||
`torch.Tensor`: Precomputed frequency tensor with complex exponentials. [S, D/2]
|
||||
"""
|
||||
assert dim % 2 == 0
|
||||
|
||||
if isinstance(pos, int):
|
||||
pos = torch.arange(pos)
|
||||
if isinstance(pos, np.ndarray):
|
||||
pos = torch.from_numpy(pos) # type: ignore # [S]
|
||||
|
||||
freqs = 1.0 / (
|
||||
theta ** (torch.arange(0, dim, 2, device=pos.device)[: (dim // 2)].float() / dim)
|
||||
) # [D/2]
|
||||
|
||||
# === Riflex modification start ===
|
||||
# Reduce the intrinsic frequency to stay within a single period after extrapolation (see Eq. (8)).
|
||||
# Empirical observations show that a few videos may exhibit repetition in the tail frames.
|
||||
# To be conservative, we multiply by 0.9 to keep the extrapolated length below 90% of a single period.
|
||||
if k is not None:
|
||||
freqs[k-1] = 0.9 * 2 * torch.pi / L_test
|
||||
# === Riflex modification end ===
|
||||
|
||||
freqs = torch.outer(pos, freqs) # type: ignore # [S, D/2]
|
||||
if use_real:
|
||||
freqs_cos = freqs.cos().repeat_interleave(2, dim=1).float() # [S, D]
|
||||
freqs_sin = freqs.sin().repeat_interleave(2, dim=1).float() # [S, D]
|
||||
return freqs_cos, freqs_sin
|
||||
else:
|
||||
# lumina
|
||||
freqs_cis = torch.polar(torch.ones_like(freqs), freqs) # complex64 # [S, D/2]
|
||||
return freqs_cis
|
||||
|
||||
def identify_k( b: float, d: int, N: int):
|
||||
"""
|
||||
This function identifies the index of the intrinsic frequency component in a RoPE-based pre-trained diffusion transformer.
|
||||
|
||||
Args:
|
||||
b (`float`): The base frequency for RoPE.
|
||||
d (`int`): Dimension of the frequency tensor
|
||||
N (`int`): the first observed repetition frame in latent space
|
||||
Returns:
|
||||
k (`int`): the index of intrinsic frequency component
|
||||
N_k (`int`): the period of intrinsic frequency component in latent space
|
||||
Example:
|
||||
In HunyuanVideo, b=256 and d=16, the repetition occurs approximately 8s (N=48 in latent space).
|
||||
k, N_k = identify_k(b=256, d=16, N=48)
|
||||
In this case, the intrinsic frequency index k is 4, and the period N_k is 50.
|
||||
"""
|
||||
|
||||
# Compute the period of each frequency in RoPE according to Eq.(4)
|
||||
periods = []
|
||||
for j in range(1, d // 2 + 1):
|
||||
theta_j = 1.0 / (b ** (2 * (j - 1) / d))
|
||||
N_j = round(2 * torch.pi / theta_j)
|
||||
periods.append(N_j)
|
||||
|
||||
# Identify the intrinsic frequency whose period is closed to N(see Eq.(7))
|
||||
diffs = [abs(N_j - N) for N_j in periods]
|
||||
k = diffs.index(min(diffs)) + 1
|
||||
N_k = periods[k-1]
|
||||
return k, N_k
|
||||
|
||||
def _to_tuple(x, dim=2):
|
||||
if isinstance(x, int):
|
||||
return (x,) * dim
|
||||
elif len(x) == dim:
|
||||
return x
|
||||
else:
|
||||
raise ValueError(f"Expected length {dim} or int, but got {x}")
|
||||
|
||||
|
||||
def get_meshgrid_nd(start, *args, dim=2):
|
||||
"""
|
||||
Get n-D meshgrid with start, stop and num.
|
||||
|
||||
Args:
|
||||
start (int or tuple): If len(args) == 0, start is num; If len(args) == 1, start is start, args[0] is stop,
|
||||
step is 1; If len(args) == 2, start is start, args[0] is stop, args[1] is num. For n-dim, start/stop/num
|
||||
should be int or n-tuple. If n-tuple is provided, the meshgrid will be stacked following the dim order in
|
||||
n-tuples.
|
||||
*args: See above.
|
||||
dim (int): Dimension of the meshgrid. Defaults to 2.
|
||||
|
||||
Returns:
|
||||
grid (np.ndarray): [dim, ...]
|
||||
"""
|
||||
if len(args) == 0:
|
||||
# start is grid_size
|
||||
num = _to_tuple(start, dim=dim)
|
||||
start = (0,) * dim
|
||||
stop = num
|
||||
elif len(args) == 1:
|
||||
# start is start, args[0] is stop, step is 1
|
||||
start = _to_tuple(start, dim=dim)
|
||||
stop = _to_tuple(args[0], dim=dim)
|
||||
num = [stop[i] - start[i] for i in range(dim)]
|
||||
elif len(args) == 2:
|
||||
# start is start, args[0] is stop, args[1] is num
|
||||
start = _to_tuple(start, dim=dim) # Left-Top eg: 12,0
|
||||
stop = _to_tuple(args[0], dim=dim) # Right-Bottom eg: 20,32
|
||||
num = _to_tuple(args[1], dim=dim) # Target Size eg: 32,124
|
||||
else:
|
||||
raise ValueError(f"len(args) should be 0, 1 or 2, but got {len(args)}")
|
||||
|
||||
# PyTorch implement of np.linspace(start[i], stop[i], num[i], endpoint=False)
|
||||
axis_grid = []
|
||||
for i in range(dim):
|
||||
a, b, n = start[i], stop[i], num[i]
|
||||
g = torch.linspace(a, b, n + 1, dtype=torch.float32)[:n]
|
||||
axis_grid.append(g)
|
||||
grid = torch.meshgrid(*axis_grid, indexing="ij") # dim x [W, H, D]
|
||||
grid = torch.stack(grid, dim=0) # [dim, W, H, D]
|
||||
|
||||
return grid
|
||||
|
||||
|
||||
#################################################################################
|
||||
# Rotary Positional Embedding Functions #
|
||||
#################################################################################
|
||||
# https://github.com/meta-llama/llama/blob/be327c427cc5e89cc1d3ab3d3fec4484df771245/llama/model.py#L80
|
||||
|
||||
|
||||
def reshape_for_broadcast(
|
||||
freqs_cis: Union[torch.Tensor, Tuple[torch.Tensor]],
|
||||
x: torch.Tensor,
|
||||
head_first=False,
|
||||
):
|
||||
"""
|
||||
Reshape frequency tensor for broadcasting it with another tensor.
|
||||
|
||||
This function reshapes the frequency tensor to have the same shape as the target tensor 'x'
|
||||
for the purpose of broadcasting the frequency tensor during element-wise operations.
|
||||
|
||||
Notes:
|
||||
When using FlashMHAModified, head_first should be False.
|
||||
When using Attention, head_first should be True.
|
||||
|
||||
Args:
|
||||
freqs_cis (Union[torch.Tensor, Tuple[torch.Tensor]]): Frequency tensor to be reshaped.
|
||||
x (torch.Tensor): Target tensor for broadcasting compatibility.
|
||||
head_first (bool): head dimension first (except batch dim) or not.
|
||||
|
||||
Returns:
|
||||
torch.Tensor: Reshaped frequency tensor.
|
||||
|
||||
Raises:
|
||||
AssertionError: If the frequency tensor doesn't match the expected shape.
|
||||
AssertionError: If the target tensor 'x' doesn't have the expected number of dimensions.
|
||||
"""
|
||||
ndim = x.ndim
|
||||
assert 0 <= 1 < ndim
|
||||
|
||||
if isinstance(freqs_cis, tuple):
|
||||
# freqs_cis: (cos, sin) in real space
|
||||
if head_first:
|
||||
assert freqs_cis[0].shape == (
|
||||
x.shape[-2],
|
||||
x.shape[-1],
|
||||
), f"freqs_cis shape {freqs_cis[0].shape} does not match x shape {x.shape}"
|
||||
shape = [
|
||||
d if i == ndim - 2 or i == ndim - 1 else 1
|
||||
for i, d in enumerate(x.shape)
|
||||
]
|
||||
else:
|
||||
assert freqs_cis[0].shape == (
|
||||
x.shape[1],
|
||||
x.shape[-1],
|
||||
), f"freqs_cis shape {freqs_cis[0].shape} does not match x shape {x.shape}"
|
||||
shape = [d if i == 1 or i == ndim - 1 else 1 for i, d in enumerate(x.shape)]
|
||||
return freqs_cis[0].view(*shape), freqs_cis[1].view(*shape)
|
||||
else:
|
||||
# freqs_cis: values in complex space
|
||||
if head_first:
|
||||
assert freqs_cis.shape == (
|
||||
x.shape[-2],
|
||||
x.shape[-1],
|
||||
), f"freqs_cis shape {freqs_cis.shape} does not match x shape {x.shape}"
|
||||
shape = [
|
||||
d if i == ndim - 2 or i == ndim - 1 else 1
|
||||
for i, d in enumerate(x.shape)
|
||||
]
|
||||
else:
|
||||
assert freqs_cis.shape == (
|
||||
x.shape[1],
|
||||
x.shape[-1],
|
||||
), f"freqs_cis shape {freqs_cis.shape} does not match x shape {x.shape}"
|
||||
shape = [d if i == 1 or i == ndim - 1 else 1 for i, d in enumerate(x.shape)]
|
||||
return freqs_cis.view(*shape)
|
||||
|
||||
|
||||
def rotate_half(x):
|
||||
x_real, x_imag = (
|
||||
x.float().reshape(*x.shape[:-1], -1, 2).unbind(-1)
|
||||
) # [B, S, H, D//2]
|
||||
return torch.stack([-x_imag, x_real], dim=-1).flatten(3)
|
||||
|
||||
|
||||
def apply_rotary_emb( qklist,
|
||||
freqs_cis: Union[torch.Tensor, Tuple[torch.Tensor, torch.Tensor]],
|
||||
head_first: bool = False,
|
||||
) -> Tuple[torch.Tensor, torch.Tensor]:
|
||||
"""
|
||||
Apply rotary embeddings to input tensors using the given frequency tensor.
|
||||
|
||||
This function applies rotary embeddings to the given query 'xq' and key 'xk' tensors using the provided
|
||||
frequency tensor 'freqs_cis'. The input tensors are reshaped as complex numbers, and the frequency tensor
|
||||
is reshaped for broadcasting compatibility. The resulting tensors contain rotary embeddings and are
|
||||
returned as real tensors.
|
||||
|
||||
Args:
|
||||
xq (torch.Tensor): Query tensor to apply rotary embeddings. [B, S, H, D]
|
||||
xk (torch.Tensor): Key tensor to apply rotary embeddings. [B, S, H, D]
|
||||
freqs_cis (torch.Tensor or tuple): Precomputed frequency tensor for complex exponential.
|
||||
head_first (bool): head dimension first (except batch dim) or not.
|
||||
|
||||
Returns:
|
||||
Tuple[torch.Tensor, torch.Tensor]: Tuple of modified query tensor and key tensor with rotary embeddings.
|
||||
|
||||
"""
|
||||
xq, xk = qklist
|
||||
qklist.clear()
|
||||
xk_out = None
|
||||
if isinstance(freqs_cis, tuple):
|
||||
cos, sin = reshape_for_broadcast(freqs_cis, xq, head_first) # [S, D]
|
||||
cos, sin = cos.to(xq.device), sin.to(xq.device)
|
||||
# real * cos - imag * sin
|
||||
# imag * cos + real * sin
|
||||
xq_dtype = xq.dtype
|
||||
xq_out = xq.to(torch.float)
|
||||
xq = None
|
||||
xq_rot = rotate_half(xq_out)
|
||||
xq_out *= cos
|
||||
xq_rot *= sin
|
||||
xq_out += xq_rot
|
||||
del xq_rot
|
||||
xq_out = xq_out.to(xq_dtype)
|
||||
|
||||
xk_out = xk.to(torch.float)
|
||||
xk = None
|
||||
xk_rot = rotate_half(xk_out)
|
||||
xk_out *= cos
|
||||
xk_rot *= sin
|
||||
xk_out += xk_rot
|
||||
del xk_rot
|
||||
xk_out = xk_out.to(xq_dtype)
|
||||
else:
|
||||
# view_as_complex will pack [..., D/2, 2](real) to [..., D/2](complex)
|
||||
xq_ = torch.view_as_complex(
|
||||
xq.float().reshape(*xq.shape[:-1], -1, 2)
|
||||
) # [B, S, H, D//2]
|
||||
freqs_cis = reshape_for_broadcast(freqs_cis, xq_, head_first).to(
|
||||
xq.device
|
||||
) # [S, D//2] --> [1, S, 1, D//2]
|
||||
# (real, imag) * (cos, sin) = (real * cos - imag * sin, imag * cos + real * sin)
|
||||
# view_as_real will expand [..., D/2](complex) to [..., D/2, 2](real)
|
||||
xq_out = torch.view_as_real(xq_ * freqs_cis).flatten(3).type_as(xq)
|
||||
xk_ = torch.view_as_complex(
|
||||
xk.float().reshape(*xk.shape[:-1], -1, 2)
|
||||
) # [B, S, H, D//2]
|
||||
xk_out = torch.view_as_real(xk_ * freqs_cis).flatten(3).type_as(xk)
|
||||
|
||||
return xq_out, xk_out
|
||||
|
||||
def get_nd_rotary_pos_embed_new(rope_dim_list, start, *args, theta=10000., use_real=False,
|
||||
theta_rescale_factor: Union[float, List[float]]=1.0,
|
||||
interpolation_factor: Union[float, List[float]]=1.0,
|
||||
concat_dict={}
|
||||
):
|
||||
|
||||
grid = get_meshgrid_nd(start, *args, dim=len(rope_dim_list)) # [3, W, H, D] / [2, W, H]
|
||||
if len(concat_dict)<1:
|
||||
pass
|
||||
else:
|
||||
if concat_dict['mode']=='timecat':
|
||||
bias = grid[:,:1].clone()
|
||||
bias[0] = concat_dict['bias']*torch.ones_like(bias[0])
|
||||
grid = torch.cat([bias, grid], dim=1)
|
||||
|
||||
elif concat_dict['mode']=='timecat-w':
|
||||
bias = grid[:,:1].clone()
|
||||
bias[0] = concat_dict['bias']*torch.ones_like(bias[0])
|
||||
bias[2] += start[-1] ## ref https://github.com/Yuanshi9815/OminiControl/blob/main/src/generate.py#L178
|
||||
grid = torch.cat([bias, grid], dim=1)
|
||||
if isinstance(theta_rescale_factor, int) or isinstance(theta_rescale_factor, float):
|
||||
theta_rescale_factor = [theta_rescale_factor] * len(rope_dim_list)
|
||||
elif isinstance(theta_rescale_factor, list) and len(theta_rescale_factor) == 1:
|
||||
theta_rescale_factor = [theta_rescale_factor[0]] * len(rope_dim_list)
|
||||
assert len(theta_rescale_factor) == len(rope_dim_list), "len(theta_rescale_factor) should equal to len(rope_dim_list)"
|
||||
|
||||
if isinstance(interpolation_factor, int) or isinstance(interpolation_factor, float):
|
||||
interpolation_factor = [interpolation_factor] * len(rope_dim_list)
|
||||
elif isinstance(interpolation_factor, list) and len(interpolation_factor) == 1:
|
||||
interpolation_factor = [interpolation_factor[0]] * len(rope_dim_list)
|
||||
assert len(interpolation_factor) == len(rope_dim_list), "len(interpolation_factor) should equal to len(rope_dim_list)"
|
||||
|
||||
# use 1/ndim of dimensions to encode grid_axis
|
||||
embs = []
|
||||
for i in range(len(rope_dim_list)):
|
||||
emb = get_1d_rotary_pos_embed(rope_dim_list[i], grid[i].reshape(-1), theta, use_real=use_real,
|
||||
theta_rescale_factor=theta_rescale_factor[i],
|
||||
interpolation_factor=interpolation_factor[i]) # 2 x [WHD, rope_dim_list[i]]
|
||||
|
||||
embs.append(emb)
|
||||
|
||||
if use_real:
|
||||
cos = torch.cat([emb[0] for emb in embs], dim=1) # (WHD, D/2)
|
||||
sin = torch.cat([emb[1] for emb in embs], dim=1) # (WHD, D/2)
|
||||
return cos, sin
|
||||
else:
|
||||
emb = torch.cat(embs, dim=1) # (WHD, D/2)
|
||||
return emb
|
||||
|
||||
def get_nd_rotary_pos_embed(
|
||||
rope_dim_list,
|
||||
start,
|
||||
*args,
|
||||
theta=10000.0,
|
||||
use_real=False,
|
||||
theta_rescale_factor: Union[float, List[float]] = 1.0,
|
||||
interpolation_factor: Union[float, List[float]] = 1.0,
|
||||
k = 4,
|
||||
L_test = 66,
|
||||
enable_riflex = True
|
||||
):
|
||||
"""
|
||||
This is a n-d version of precompute_freqs_cis, which is a RoPE for tokens with n-d structure.
|
||||
|
||||
Args:
|
||||
rope_dim_list (list of int): Dimension of each rope. len(rope_dim_list) should equal to n.
|
||||
sum(rope_dim_list) should equal to head_dim of attention layer.
|
||||
start (int | tuple of int | list of int): If len(args) == 0, start is num; If len(args) == 1, start is start,
|
||||
args[0] is stop, step is 1; If len(args) == 2, start is start, args[0] is stop, args[1] is num.
|
||||
*args: See above.
|
||||
theta (float): Scaling factor for frequency computation. Defaults to 10000.0.
|
||||
use_real (bool): If True, return real part and imaginary part separately. Otherwise, return complex numbers.
|
||||
Some libraries such as TensorRT does not support complex64 data type. So it is useful to provide a real
|
||||
part and an imaginary part separately.
|
||||
theta_rescale_factor (float): Rescale factor for theta. Defaults to 1.0.
|
||||
|
||||
Returns:
|
||||
pos_embed (torch.Tensor): [HW, D/2]
|
||||
"""
|
||||
|
||||
grid = get_meshgrid_nd(
|
||||
start, *args, dim=len(rope_dim_list)
|
||||
) # [3, W, H, D] / [2, W, H]
|
||||
|
||||
if isinstance(theta_rescale_factor, int) or isinstance(theta_rescale_factor, float):
|
||||
theta_rescale_factor = [theta_rescale_factor] * len(rope_dim_list)
|
||||
elif isinstance(theta_rescale_factor, list) and len(theta_rescale_factor) == 1:
|
||||
theta_rescale_factor = [theta_rescale_factor[0]] * len(rope_dim_list)
|
||||
assert len(theta_rescale_factor) == len(
|
||||
rope_dim_list
|
||||
), "len(theta_rescale_factor) should equal to len(rope_dim_list)"
|
||||
|
||||
if isinstance(interpolation_factor, int) or isinstance(interpolation_factor, float):
|
||||
interpolation_factor = [interpolation_factor] * len(rope_dim_list)
|
||||
elif isinstance(interpolation_factor, list) and len(interpolation_factor) == 1:
|
||||
interpolation_factor = [interpolation_factor[0]] * len(rope_dim_list)
|
||||
assert len(interpolation_factor) == len(
|
||||
rope_dim_list
|
||||
), "len(interpolation_factor) should equal to len(rope_dim_list)"
|
||||
|
||||
# use 1/ndim of dimensions to encode grid_axis
|
||||
embs = []
|
||||
for i in range(len(rope_dim_list)):
|
||||
# emb = get_1d_rotary_pos_embed(
|
||||
# rope_dim_list[i],
|
||||
# grid[i].reshape(-1),
|
||||
# theta,
|
||||
# use_real=use_real,
|
||||
# theta_rescale_factor=theta_rescale_factor[i],
|
||||
# interpolation_factor=interpolation_factor[i],
|
||||
# ) # 2 x [WHD, rope_dim_list[i]]
|
||||
|
||||
|
||||
# === RIFLEx modification start ===
|
||||
# apply RIFLEx for time dimension
|
||||
if i == 0 and enable_riflex:
|
||||
emb = get_1d_rotary_pos_embed_riflex(rope_dim_list[i], grid[i].reshape(-1), theta, use_real=True, k=k, L_test=L_test)
|
||||
# === RIFLEx modification end ===
|
||||
else:
|
||||
emb = get_1d_rotary_pos_embed(rope_dim_list[i], grid[i].reshape(-1), theta, use_real=True, theta_rescale_factor=theta_rescale_factor[i],interpolation_factor=interpolation_factor[i],)
|
||||
embs.append(emb)
|
||||
|
||||
if use_real:
|
||||
cos = torch.cat([emb[0] for emb in embs], dim=1) # (WHD, D/2)
|
||||
sin = torch.cat([emb[1] for emb in embs], dim=1) # (WHD, D/2)
|
||||
return cos, sin
|
||||
else:
|
||||
emb = torch.cat(embs, dim=1) # (WHD, D/2)
|
||||
return emb
|
||||
|
||||
|
||||
def get_1d_rotary_pos_embed(
|
||||
dim: int,
|
||||
pos: Union[torch.FloatTensor, int],
|
||||
theta: float = 10000.0,
|
||||
use_real: bool = False,
|
||||
theta_rescale_factor: float = 1.0,
|
||||
interpolation_factor: float = 1.0,
|
||||
) -> Union[torch.Tensor, Tuple[torch.Tensor, torch.Tensor]]:
|
||||
"""
|
||||
Precompute the frequency tensor for complex exponential (cis) with given dimensions.
|
||||
(Note: `cis` means `cos + i * sin`, where i is the imaginary unit.)
|
||||
|
||||
This function calculates a frequency tensor with complex exponential using the given dimension 'dim'
|
||||
and the end index 'end'. The 'theta' parameter scales the frequencies.
|
||||
The returned tensor contains complex values in complex64 data type.
|
||||
|
||||
Args:
|
||||
dim (int): Dimension of the frequency tensor.
|
||||
pos (int or torch.FloatTensor): Position indices for the frequency tensor. [S] or scalar
|
||||
theta (float, optional): Scaling factor for frequency computation. Defaults to 10000.0.
|
||||
use_real (bool, optional): If True, return real part and imaginary part separately.
|
||||
Otherwise, return complex numbers.
|
||||
theta_rescale_factor (float, optional): Rescale factor for theta. Defaults to 1.0.
|
||||
|
||||
Returns:
|
||||
freqs_cis: Precomputed frequency tensor with complex exponential. [S, D/2]
|
||||
freqs_cos, freqs_sin: Precomputed frequency tensor with real and imaginary parts separately. [S, D]
|
||||
"""
|
||||
if isinstance(pos, int):
|
||||
pos = torch.arange(pos).float()
|
||||
|
||||
# proposed by reddit user bloc97, to rescale rotary embeddings to longer sequence length without fine-tuning
|
||||
# has some connection to NTK literature
|
||||
if theta_rescale_factor != 1.0:
|
||||
theta *= theta_rescale_factor ** (dim / (dim - 2))
|
||||
|
||||
freqs = 1.0 / (
|
||||
theta ** (torch.arange(0, dim, 2)[: (dim // 2)].float() / dim)
|
||||
) # [D/2]
|
||||
# assert interpolation_factor == 1.0, f"interpolation_factor: {interpolation_factor}"
|
||||
freqs = torch.outer(pos * interpolation_factor, freqs) # [S, D/2]
|
||||
if use_real:
|
||||
freqs_cos = freqs.cos().repeat_interleave(2, dim=1) # [S, D]
|
||||
freqs_sin = freqs.sin().repeat_interleave(2, dim=1) # [S, D]
|
||||
return freqs_cos, freqs_sin
|
||||
else:
|
||||
freqs_cis = torch.polar(
|
||||
torch.ones_like(freqs), freqs
|
||||
) # complex64 # [S, D/2]
|
||||
return freqs_cis
|
||||
237
hyvideo/modules/token_refiner.py
Normal file
237
hyvideo/modules/token_refiner.py
Normal file
@@ -0,0 +1,237 @@
|
||||
from typing import Optional
|
||||
|
||||
from einops import rearrange
|
||||
import torch
|
||||
import torch.nn as nn
|
||||
|
||||
from .activation_layers import get_activation_layer
|
||||
from .attenion import attention
|
||||
from .norm_layers import get_norm_layer
|
||||
from .embed_layers import TimestepEmbedder, TextProjection
|
||||
from .attenion import attention
|
||||
from .mlp_layers import MLP
|
||||
from .modulate_layers import modulate, apply_gate
|
||||
|
||||
|
||||
class IndividualTokenRefinerBlock(nn.Module):
|
||||
def __init__(
|
||||
self,
|
||||
hidden_size,
|
||||
heads_num,
|
||||
mlp_width_ratio: str = 4.0,
|
||||
mlp_drop_rate: float = 0.0,
|
||||
act_type: str = "silu",
|
||||
qk_norm: bool = False,
|
||||
qk_norm_type: str = "layer",
|
||||
qkv_bias: bool = True,
|
||||
dtype: Optional[torch.dtype] = None,
|
||||
device: Optional[torch.device] = None,
|
||||
):
|
||||
factory_kwargs = {"device": device, "dtype": dtype}
|
||||
super().__init__()
|
||||
self.heads_num = heads_num
|
||||
head_dim = hidden_size // heads_num
|
||||
mlp_hidden_dim = int(hidden_size * mlp_width_ratio)
|
||||
|
||||
self.norm1 = nn.LayerNorm(
|
||||
hidden_size, elementwise_affine=True, eps=1e-6, **factory_kwargs
|
||||
)
|
||||
self.self_attn_qkv = nn.Linear(
|
||||
hidden_size, hidden_size * 3, bias=qkv_bias, **factory_kwargs
|
||||
)
|
||||
qk_norm_layer = get_norm_layer(qk_norm_type)
|
||||
self.self_attn_q_norm = (
|
||||
qk_norm_layer(head_dim, elementwise_affine=True, eps=1e-6, **factory_kwargs)
|
||||
if qk_norm
|
||||
else nn.Identity()
|
||||
)
|
||||
self.self_attn_k_norm = (
|
||||
qk_norm_layer(head_dim, elementwise_affine=True, eps=1e-6, **factory_kwargs)
|
||||
if qk_norm
|
||||
else nn.Identity()
|
||||
)
|
||||
self.self_attn_proj = nn.Linear(
|
||||
hidden_size, hidden_size, bias=qkv_bias, **factory_kwargs
|
||||
)
|
||||
|
||||
self.norm2 = nn.LayerNorm(
|
||||
hidden_size, elementwise_affine=True, eps=1e-6, **factory_kwargs
|
||||
)
|
||||
act_layer = get_activation_layer(act_type)
|
||||
self.mlp = MLP(
|
||||
in_channels=hidden_size,
|
||||
hidden_channels=mlp_hidden_dim,
|
||||
act_layer=act_layer,
|
||||
drop=mlp_drop_rate,
|
||||
**factory_kwargs,
|
||||
)
|
||||
|
||||
self.adaLN_modulation = nn.Sequential(
|
||||
act_layer(),
|
||||
nn.Linear(hidden_size, 2 * hidden_size, bias=True, **factory_kwargs),
|
||||
)
|
||||
# Zero-initialize the modulation
|
||||
nn.init.zeros_(self.adaLN_modulation[1].weight)
|
||||
nn.init.zeros_(self.adaLN_modulation[1].bias)
|
||||
|
||||
def forward(
|
||||
self,
|
||||
x: torch.Tensor,
|
||||
c: torch.Tensor, # timestep_aware_representations + context_aware_representations
|
||||
attn_mask: torch.Tensor = None,
|
||||
):
|
||||
gate_msa, gate_mlp = self.adaLN_modulation(c).chunk(2, dim=1)
|
||||
|
||||
norm_x = self.norm1(x)
|
||||
qkv = self.self_attn_qkv(norm_x)
|
||||
q, k, v = rearrange(qkv, "B L (K H D) -> K B L H D", K=3, H=self.heads_num)
|
||||
# Apply QK-Norm if needed
|
||||
q = self.self_attn_q_norm(q).to(v)
|
||||
k = self.self_attn_k_norm(k).to(v)
|
||||
qkv_list = [q, k, v]
|
||||
del q,k
|
||||
# Self-Attention
|
||||
attn = attention( qkv_list, mode="torch", attn_mask=attn_mask)
|
||||
|
||||
x = x + apply_gate(self.self_attn_proj(attn), gate_msa)
|
||||
|
||||
# FFN Layer
|
||||
x = x + apply_gate(self.mlp(self.norm2(x)), gate_mlp)
|
||||
|
||||
return x
|
||||
|
||||
|
||||
class IndividualTokenRefiner(nn.Module):
|
||||
def __init__(
|
||||
self,
|
||||
hidden_size,
|
||||
heads_num,
|
||||
depth,
|
||||
mlp_width_ratio: float = 4.0,
|
||||
mlp_drop_rate: float = 0.0,
|
||||
act_type: str = "silu",
|
||||
qk_norm: bool = False,
|
||||
qk_norm_type: str = "layer",
|
||||
qkv_bias: bool = True,
|
||||
dtype: Optional[torch.dtype] = None,
|
||||
device: Optional[torch.device] = None,
|
||||
):
|
||||
factory_kwargs = {"device": device, "dtype": dtype}
|
||||
super().__init__()
|
||||
self.blocks = nn.ModuleList(
|
||||
[
|
||||
IndividualTokenRefinerBlock(
|
||||
hidden_size=hidden_size,
|
||||
heads_num=heads_num,
|
||||
mlp_width_ratio=mlp_width_ratio,
|
||||
mlp_drop_rate=mlp_drop_rate,
|
||||
act_type=act_type,
|
||||
qk_norm=qk_norm,
|
||||
qk_norm_type=qk_norm_type,
|
||||
qkv_bias=qkv_bias,
|
||||
**factory_kwargs,
|
||||
)
|
||||
for _ in range(depth)
|
||||
]
|
||||
)
|
||||
|
||||
def forward(
|
||||
self,
|
||||
x: torch.Tensor,
|
||||
c: torch.LongTensor,
|
||||
mask: Optional[torch.Tensor] = None,
|
||||
):
|
||||
self_attn_mask = None
|
||||
if mask is not None:
|
||||
batch_size = mask.shape[0]
|
||||
seq_len = mask.shape[1]
|
||||
mask = mask.to(x.device)
|
||||
# batch_size x 1 x seq_len x seq_len
|
||||
self_attn_mask_1 = mask.view(batch_size, 1, 1, seq_len).repeat(
|
||||
1, 1, seq_len, 1
|
||||
)
|
||||
# batch_size x 1 x seq_len x seq_len
|
||||
self_attn_mask_2 = self_attn_mask_1.transpose(2, 3)
|
||||
# batch_size x 1 x seq_len x seq_len, 1 for broadcasting of heads_num
|
||||
self_attn_mask = (self_attn_mask_1 & self_attn_mask_2).bool()
|
||||
# avoids self-attention weight being NaN for padding tokens
|
||||
self_attn_mask[:, :, :, 0] = True
|
||||
|
||||
for block in self.blocks:
|
||||
x = block(x, c, self_attn_mask)
|
||||
return x
|
||||
|
||||
|
||||
class SingleTokenRefiner(nn.Module):
|
||||
"""
|
||||
A single token refiner block for llm text embedding refine.
|
||||
"""
|
||||
def __init__(
|
||||
self,
|
||||
in_channels,
|
||||
hidden_size,
|
||||
heads_num,
|
||||
depth,
|
||||
mlp_width_ratio: float = 4.0,
|
||||
mlp_drop_rate: float = 0.0,
|
||||
act_type: str = "silu",
|
||||
qk_norm: bool = False,
|
||||
qk_norm_type: str = "layer",
|
||||
qkv_bias: bool = True,
|
||||
attn_mode: str = "torch",
|
||||
dtype: Optional[torch.dtype] = None,
|
||||
device: Optional[torch.device] = None,
|
||||
):
|
||||
factory_kwargs = {"device": device, "dtype": dtype}
|
||||
super().__init__()
|
||||
self.attn_mode = attn_mode
|
||||
assert self.attn_mode == "torch", "Only support 'torch' mode for token refiner."
|
||||
|
||||
self.input_embedder = nn.Linear(
|
||||
in_channels, hidden_size, bias=True, **factory_kwargs
|
||||
)
|
||||
|
||||
act_layer = get_activation_layer(act_type)
|
||||
# Build timestep embedding layer
|
||||
self.t_embedder = TimestepEmbedder(hidden_size, act_layer, **factory_kwargs)
|
||||
# Build context embedding layer
|
||||
self.c_embedder = TextProjection(
|
||||
in_channels, hidden_size, act_layer, **factory_kwargs
|
||||
)
|
||||
|
||||
self.individual_token_refiner = IndividualTokenRefiner(
|
||||
hidden_size=hidden_size,
|
||||
heads_num=heads_num,
|
||||
depth=depth,
|
||||
mlp_width_ratio=mlp_width_ratio,
|
||||
mlp_drop_rate=mlp_drop_rate,
|
||||
act_type=act_type,
|
||||
qk_norm=qk_norm,
|
||||
qk_norm_type=qk_norm_type,
|
||||
qkv_bias=qkv_bias,
|
||||
**factory_kwargs,
|
||||
)
|
||||
|
||||
def forward(
|
||||
self,
|
||||
x: torch.Tensor,
|
||||
t: torch.LongTensor,
|
||||
mask: Optional[torch.LongTensor] = None,
|
||||
):
|
||||
timestep_aware_representations = self.t_embedder(t)
|
||||
|
||||
if mask is None:
|
||||
context_aware_representations = x.mean(dim=1)
|
||||
else:
|
||||
mask_float = mask.float().unsqueeze(-1) # [b, s1, 1]
|
||||
context_aware_representations = (x * mask_float).sum(
|
||||
dim=1
|
||||
) / mask_float.sum(dim=1)
|
||||
context_aware_representations = self.c_embedder(context_aware_representations.to(x.dtype))
|
||||
c = timestep_aware_representations + context_aware_representations
|
||||
|
||||
x = self.input_embedder(x)
|
||||
|
||||
x = self.individual_token_refiner(x, c, mask)
|
||||
|
||||
return x
|
||||
43
hyvideo/modules/utils.py
Normal file
43
hyvideo/modules/utils.py
Normal file
@@ -0,0 +1,43 @@
|
||||
"""Mask Mod for Image2Video"""
|
||||
|
||||
from math import floor
|
||||
import torch
|
||||
from torch import Tensor
|
||||
|
||||
|
||||
from functools import lru_cache
|
||||
from typing import Optional, List
|
||||
|
||||
import torch
|
||||
from torch.nn.attention.flex_attention import (
|
||||
create_block_mask,
|
||||
)
|
||||
|
||||
|
||||
@lru_cache
|
||||
def create_block_mask_cached(score_mod, B, H, M, N, device="cuda", _compile=False):
|
||||
block_mask = create_block_mask(score_mod, B, H, M, N, device=device, _compile=_compile)
|
||||
return block_mask
|
||||
|
||||
def generate_temporal_head_mask_mod(context_length: int = 226, prompt_length: int = 226, num_frames: int = 13, token_per_frame: int = 1350, mul: int = 2):
|
||||
|
||||
def round_to_multiple(idx):
|
||||
return floor(idx / 128) * 128
|
||||
|
||||
real_length = num_frames * token_per_frame + prompt_length
|
||||
def temporal_mask_mod(b, h, q_idx, kv_idx):
|
||||
real_mask = (kv_idx < real_length) & (q_idx < real_length)
|
||||
fake_mask = (kv_idx >= real_length) & (q_idx >= real_length)
|
||||
|
||||
two_frame = round_to_multiple(mul * token_per_frame)
|
||||
temporal_head_mask = (torch.abs(q_idx - kv_idx) < two_frame)
|
||||
|
||||
text_column_mask = (num_frames * token_per_frame <= kv_idx) & (kv_idx < real_length)
|
||||
text_row_mask = (num_frames * token_per_frame <= q_idx) & (q_idx < real_length)
|
||||
|
||||
video_mask = temporal_head_mask | text_column_mask | text_row_mask
|
||||
real_mask = real_mask & video_mask
|
||||
|
||||
return real_mask | fake_mask
|
||||
|
||||
return temporal_mask_mod
|
||||
51
hyvideo/prompt_rewrite.py
Normal file
51
hyvideo/prompt_rewrite.py
Normal file
@@ -0,0 +1,51 @@
|
||||
normal_mode_prompt = """Normal mode - Video Recaption Task:
|
||||
|
||||
You are a large language model specialized in rewriting video descriptions. Your task is to modify the input description.
|
||||
|
||||
0. Preserve ALL information, including style words and technical terms.
|
||||
|
||||
1. If the input is in Chinese, translate the entire description to English.
|
||||
|
||||
2. If the input is just one or two words describing an object or person, provide a brief, simple description focusing on basic visual characteristics. Limit the description to 1-2 short sentences.
|
||||
|
||||
3. If the input does not include style, lighting, atmosphere, you can make reasonable associations.
|
||||
|
||||
4. Output ALL must be in English.
|
||||
|
||||
Given Input:
|
||||
input: "{input}"
|
||||
"""
|
||||
|
||||
|
||||
master_mode_prompt = """Master mode - Video Recaption Task:
|
||||
|
||||
You are a large language model specialized in rewriting video descriptions. Your task is to modify the input description.
|
||||
|
||||
0. Preserve ALL information, including style words and technical terms.
|
||||
|
||||
1. If the input is in Chinese, translate the entire description to English.
|
||||
|
||||
2. If the input is just one or two words describing an object or person, provide a brief, simple description focusing on basic visual characteristics. Limit the description to 1-2 short sentences.
|
||||
|
||||
3. If the input does not include style, lighting, atmosphere, you can make reasonable associations.
|
||||
|
||||
4. Output ALL must be in English.
|
||||
|
||||
Given Input:
|
||||
input: "{input}"
|
||||
"""
|
||||
|
||||
def get_rewrite_prompt(ori_prompt, mode="Normal"):
|
||||
if mode == "Normal":
|
||||
prompt = normal_mode_prompt.format(input=ori_prompt)
|
||||
elif mode == "Master":
|
||||
prompt = master_mode_prompt.format(input=ori_prompt)
|
||||
else:
|
||||
raise Exception("Only supports Normal and Normal", mode)
|
||||
return prompt
|
||||
|
||||
ori_prompt = "一只小狗在草地上奔跑。"
|
||||
normal_prompt = get_rewrite_prompt(ori_prompt, mode="Normal")
|
||||
master_prompt = get_rewrite_prompt(ori_prompt, mode="Master")
|
||||
|
||||
# Then you can use the normal_prompt or master_prompt to access the hunyuan-large rewrite model to get the final prompt.
|
||||
552
hyvideo/text_encoder/__init__.py
Normal file
552
hyvideo/text_encoder/__init__.py
Normal file
@@ -0,0 +1,552 @@
|
||||
from dataclasses import dataclass
|
||||
from typing import Optional, Tuple
|
||||
from copy import deepcopy
|
||||
import torch
|
||||
import torch.nn as nn
|
||||
from transformers import (
|
||||
CLIPTextModel,
|
||||
CLIPTokenizer,
|
||||
AutoTokenizer,
|
||||
AutoModel,
|
||||
LlavaForConditionalGeneration,
|
||||
CLIPImageProcessor,
|
||||
)
|
||||
from transformers.utils import ModelOutput
|
||||
|
||||
from ..constants import TEXT_ENCODER_PATH, TOKENIZER_PATH
|
||||
from ..constants import PRECISION_TO_TYPE
|
||||
|
||||
|
||||
def use_default(value, default):
|
||||
return value if value is not None else default
|
||||
|
||||
|
||||
def load_text_encoder(
|
||||
text_encoder_type,
|
||||
text_encoder_precision=None,
|
||||
text_encoder_path=None,
|
||||
device=None,
|
||||
):
|
||||
if text_encoder_path is None:
|
||||
text_encoder_path = TEXT_ENCODER_PATH[text_encoder_type]
|
||||
|
||||
if text_encoder_type == "clipL":
|
||||
text_encoder = CLIPTextModel.from_pretrained(text_encoder_path)
|
||||
text_encoder.final_layer_norm = text_encoder.text_model.final_layer_norm
|
||||
elif text_encoder_type == "llm":
|
||||
text_encoder = AutoModel.from_pretrained(
|
||||
text_encoder_path, low_cpu_mem_usage=True
|
||||
)
|
||||
text_encoder.final_layer_norm = text_encoder.norm
|
||||
elif text_encoder_type == "llm-i2v":
|
||||
text_encoder = LlavaForConditionalGeneration.from_pretrained(
|
||||
text_encoder_path, low_cpu_mem_usage=True
|
||||
)
|
||||
else:
|
||||
raise ValueError(f"Unsupported text encoder type: {text_encoder_type}")
|
||||
# from_pretrained will ensure that the model is in eval mode.
|
||||
|
||||
if text_encoder_precision is not None:
|
||||
text_encoder = text_encoder.to(dtype=PRECISION_TO_TYPE[text_encoder_precision])
|
||||
|
||||
text_encoder.requires_grad_(False)
|
||||
|
||||
if device is not None:
|
||||
text_encoder = text_encoder.to(device)
|
||||
|
||||
return text_encoder, text_encoder_path
|
||||
|
||||
|
||||
def load_tokenizer(
|
||||
tokenizer_type, tokenizer_path=None, padding_side="right"
|
||||
):
|
||||
if tokenizer_path is None:
|
||||
tokenizer_path = TOKENIZER_PATH[tokenizer_type]
|
||||
|
||||
processor = None
|
||||
if tokenizer_type == "clipL":
|
||||
tokenizer = CLIPTokenizer.from_pretrained(tokenizer_path, max_length=77)
|
||||
elif tokenizer_type == "llm":
|
||||
tokenizer = AutoTokenizer.from_pretrained(
|
||||
tokenizer_path, padding_side=padding_side
|
||||
)
|
||||
elif tokenizer_type == "llm-i2v":
|
||||
tokenizer = AutoTokenizer.from_pretrained(
|
||||
tokenizer_path, padding_side=padding_side
|
||||
)
|
||||
processor = CLIPImageProcessor.from_pretrained(tokenizer_path)
|
||||
else:
|
||||
raise ValueError(f"Unsupported tokenizer type: {tokenizer_type}")
|
||||
|
||||
return tokenizer, tokenizer_path, processor
|
||||
|
||||
|
||||
@dataclass
|
||||
class TextEncoderModelOutput(ModelOutput):
|
||||
"""
|
||||
Base class for model's outputs that also contains a pooling of the last hidden states.
|
||||
|
||||
Args:
|
||||
hidden_state (`torch.FloatTensor` of shape `(batch_size, sequence_length, hidden_size)`):
|
||||
Sequence of hidden-states at the output of the last layer of the model.
|
||||
attention_mask (`torch.LongTensor` of shape `(batch_size, sequence_length)`, *optional*):
|
||||
Mask to avoid performing attention on padding token indices. Mask values selected in ``[0, 1]``:
|
||||
hidden_states_list (`tuple(torch.FloatTensor)`, *optional*, returned when `output_hidden_states=True` is passed):
|
||||
Tuple of `torch.FloatTensor` (one for the output of the embeddings, if the model has an embedding layer, +
|
||||
one for the output of each layer) of shape `(batch_size, sequence_length, hidden_size)`.
|
||||
Hidden-states of the model at the output of each layer plus the optional initial embedding outputs.
|
||||
text_outputs (`list`, *optional*, returned when `return_texts=True` is passed):
|
||||
List of decoded texts.
|
||||
"""
|
||||
|
||||
hidden_state: torch.FloatTensor = None
|
||||
attention_mask: Optional[torch.LongTensor] = None
|
||||
hidden_states_list: Optional[Tuple[torch.FloatTensor, ...]] = None
|
||||
text_outputs: Optional[list] = None
|
||||
|
||||
|
||||
class TextEncoder(nn.Module):
|
||||
def __init__(
|
||||
self,
|
||||
text_encoder_type: str,
|
||||
max_length: int,
|
||||
text_encoder_precision: Optional[str] = None,
|
||||
text_encoder_path: Optional[str] = None,
|
||||
tokenizer_type: Optional[str] = None,
|
||||
tokenizer_path: Optional[str] = None,
|
||||
output_key: Optional[str] = None,
|
||||
use_attention_mask: bool = True,
|
||||
i2v_mode: bool = False,
|
||||
input_max_length: Optional[int] = None,
|
||||
prompt_template: Optional[dict] = None,
|
||||
prompt_template_video: Optional[dict] = None,
|
||||
hidden_state_skip_layer: Optional[int] = None,
|
||||
apply_final_norm: bool = False,
|
||||
reproduce: bool = False,
|
||||
device=None,
|
||||
# image_embed_interleave (int): The number of times to interleave the image and text embeddings. Defaults to 2.
|
||||
image_embed_interleave=2,
|
||||
):
|
||||
super().__init__()
|
||||
self.text_encoder_type = text_encoder_type
|
||||
self.max_length = max_length
|
||||
self.precision = text_encoder_precision
|
||||
self.model_path = text_encoder_path
|
||||
self.tokenizer_type = (
|
||||
tokenizer_type if tokenizer_type is not None else text_encoder_type
|
||||
)
|
||||
self.tokenizer_path = (
|
||||
tokenizer_path if tokenizer_path is not None else None # text_encoder_path
|
||||
)
|
||||
self.use_attention_mask = use_attention_mask
|
||||
if prompt_template_video is not None:
|
||||
assert (
|
||||
use_attention_mask is True
|
||||
), "Attention mask is True required when training videos."
|
||||
self.input_max_length = (
|
||||
input_max_length if input_max_length is not None else max_length
|
||||
)
|
||||
self.prompt_template = prompt_template
|
||||
self.prompt_template_video = prompt_template_video
|
||||
self.hidden_state_skip_layer = hidden_state_skip_layer
|
||||
self.apply_final_norm = apply_final_norm
|
||||
self.i2v_mode = i2v_mode
|
||||
self.reproduce = reproduce
|
||||
self.image_embed_interleave = image_embed_interleave
|
||||
|
||||
self.use_template = self.prompt_template is not None
|
||||
if self.use_template:
|
||||
assert (
|
||||
isinstance(self.prompt_template, dict)
|
||||
and "template" in self.prompt_template
|
||||
), f"`prompt_template` must be a dictionary with a key 'template', got {self.prompt_template}"
|
||||
assert "{}" in str(self.prompt_template["template"]), (
|
||||
"`prompt_template['template']` must contain a placeholder `{}` for the input text, "
|
||||
f"got {self.prompt_template['template']}"
|
||||
)
|
||||
|
||||
self.use_video_template = self.prompt_template_video is not None
|
||||
if self.use_video_template:
|
||||
if self.prompt_template_video is not None:
|
||||
assert (
|
||||
isinstance(self.prompt_template_video, dict)
|
||||
and "template" in self.prompt_template_video
|
||||
), f"`prompt_template_video` must be a dictionary with a key 'template', got {self.prompt_template_video}"
|
||||
assert "{}" in str(self.prompt_template_video["template"]), (
|
||||
"`prompt_template_video['template']` must contain a placeholder `{}` for the input text, "
|
||||
f"got {self.prompt_template_video['template']}"
|
||||
)
|
||||
|
||||
if "t5" in text_encoder_type:
|
||||
self.output_key = output_key or "last_hidden_state"
|
||||
elif "clip" in text_encoder_type:
|
||||
self.output_key = output_key or "pooler_output"
|
||||
elif "llm" in text_encoder_type or "glm" in text_encoder_type:
|
||||
self.output_key = output_key or "last_hidden_state"
|
||||
else:
|
||||
raise ValueError(f"Unsupported text encoder type: {text_encoder_type}")
|
||||
|
||||
if "llm" in text_encoder_type:
|
||||
from mmgp import offload
|
||||
forcedConfigPath= None if "i2v" in text_encoder_type else "ckpts/llava-llama-3-8b/config.json"
|
||||
self.model= offload.fast_load_transformers_model(self.model_path, forcedConfigPath=forcedConfigPath, modelPrefix= "model" if forcedConfigPath !=None else None)
|
||||
if forcedConfigPath != None:
|
||||
self.model.final_layer_norm = self.model.norm
|
||||
|
||||
else:
|
||||
self.model, self.model_path = load_text_encoder(
|
||||
text_encoder_type=self.text_encoder_type,
|
||||
text_encoder_precision=self.precision,
|
||||
text_encoder_path=self.model_path,
|
||||
device=device,
|
||||
)
|
||||
|
||||
self.dtype = self.model.dtype
|
||||
self.device = self.model.device
|
||||
|
||||
self.tokenizer, self.tokenizer_path, self.processor = load_tokenizer(
|
||||
tokenizer_type=self.tokenizer_type,
|
||||
tokenizer_path=self.tokenizer_path,
|
||||
padding_side="right",
|
||||
)
|
||||
|
||||
def __repr__(self):
|
||||
return f"{self.text_encoder_type} ({self.precision} - {self.model_path})"
|
||||
|
||||
@staticmethod
|
||||
def apply_text_to_template(text, template, prevent_empty_text=True):
|
||||
"""
|
||||
Apply text to template.
|
||||
|
||||
Args:
|
||||
text (str): Input text.
|
||||
template (str or list): Template string or list of chat conversation.
|
||||
prevent_empty_text (bool): If Ture, we will prevent the user text from being empty
|
||||
by adding a space. Defaults to True.
|
||||
"""
|
||||
if isinstance(template, str):
|
||||
# Will send string to tokenizer. Used for llm
|
||||
return template.format(text)
|
||||
else:
|
||||
raise TypeError(f"Unsupported template type: {type(template)}")
|
||||
|
||||
def text2tokens(self, text, data_type="image", name = None):
|
||||
"""
|
||||
Tokenize the input text.
|
||||
|
||||
Args:
|
||||
text (str or list): Input text.
|
||||
"""
|
||||
tokenize_input_type = "str"
|
||||
if self.use_template:
|
||||
if data_type == "image":
|
||||
prompt_template = self.prompt_template["template"]
|
||||
elif data_type == "video":
|
||||
prompt_template = self.prompt_template_video["template"]
|
||||
else:
|
||||
raise ValueError(f"Unsupported data type: {data_type}")
|
||||
if isinstance(text, (list, tuple)):
|
||||
text = [
|
||||
self.apply_text_to_template(one_text, prompt_template)
|
||||
for one_text in text
|
||||
]
|
||||
if isinstance(text[0], list):
|
||||
tokenize_input_type = "list"
|
||||
elif isinstance(text, str):
|
||||
text = self.apply_text_to_template(text, prompt_template)
|
||||
if isinstance(text, list):
|
||||
tokenize_input_type = "list"
|
||||
else:
|
||||
raise TypeError(f"Unsupported text type: {type(text)}")
|
||||
|
||||
kwargs = dict(truncation=True, max_length=self.max_length, padding="max_length", return_tensors="pt")
|
||||
if self.text_encoder_type == "llm-i2v" and name != None: #llava-llama-3-8b
|
||||
if isinstance(text, list):
|
||||
for i in range(len(text)):
|
||||
text[i] = text[i] + '\nThe %s looks like<image>' % name
|
||||
elif isinstance(text, str):
|
||||
text = text + '\nThe %s looks like<image>' % name
|
||||
else:
|
||||
raise NotImplementedError
|
||||
|
||||
kwargs = dict(
|
||||
truncation=True,
|
||||
max_length=self.max_length,
|
||||
padding="max_length",
|
||||
return_tensors="pt",
|
||||
)
|
||||
if tokenize_input_type == "str":
|
||||
return self.tokenizer(
|
||||
text,
|
||||
return_length=False,
|
||||
return_overflowing_tokens=False,
|
||||
return_attention_mask=True,
|
||||
**kwargs,
|
||||
)
|
||||
elif tokenize_input_type == "list":
|
||||
return self.tokenizer.apply_chat_template(
|
||||
text,
|
||||
add_generation_prompt=True,
|
||||
tokenize=True,
|
||||
return_dict=True,
|
||||
**kwargs,
|
||||
)
|
||||
else:
|
||||
raise ValueError(f"Unsupported tokenize_input_type: {tokenize_input_type}")
|
||||
|
||||
def encode(
|
||||
self,
|
||||
batch_encoding,
|
||||
use_attention_mask=None,
|
||||
output_hidden_states=False,
|
||||
do_sample=None,
|
||||
hidden_state_skip_layer=None,
|
||||
return_texts=False,
|
||||
data_type="image",
|
||||
semantic_images=None,
|
||||
device=None,
|
||||
):
|
||||
"""
|
||||
Args:
|
||||
batch_encoding (dict): Batch encoding from tokenizer.
|
||||
use_attention_mask (bool): Whether to use attention mask. If None, use self.use_attention_mask.
|
||||
Defaults to None.
|
||||
output_hidden_states (bool): Whether to output hidden states. If False, return the value of
|
||||
self.output_key. If True, return the entire output. If set self.hidden_state_skip_layer,
|
||||
output_hidden_states will be set True. Defaults to False.
|
||||
do_sample (bool): Whether to sample from the model. Used for Decoder-Only LLMs. Defaults to None.
|
||||
When self.produce is False, do_sample is set to True by default.
|
||||
hidden_state_skip_layer (int): Number of hidden states to hidden_state_skip_layer. 0 means the last layer.
|
||||
If None, self.output_key will be used. Defaults to None.
|
||||
hidden_state_skip_layer (PIL.Image): The reference images for i2v models.
|
||||
image_embed_interleave (int): The number of times to interleave the image and text embeddings. Defaults to 2.
|
||||
return_texts (bool): Whether to return the decoded texts. Defaults to False.
|
||||
"""
|
||||
device = self.model.device if device is None else device
|
||||
use_attention_mask = use_default(use_attention_mask, self.use_attention_mask)
|
||||
hidden_state_skip_layer = use_default(
|
||||
hidden_state_skip_layer, self.hidden_state_skip_layer
|
||||
)
|
||||
do_sample = use_default(do_sample, not self.reproduce)
|
||||
if not self.i2v_mode:
|
||||
attention_mask = (
|
||||
batch_encoding["attention_mask"].to(device)
|
||||
if use_attention_mask
|
||||
else None
|
||||
)
|
||||
|
||||
if 'pixel_value_llava' in batch_encoding:
|
||||
outputs = self.model(
|
||||
input_ids=batch_encoding["input_ids"].to(self.model.device),
|
||||
attention_mask=attention_mask,
|
||||
pixel_values=batch_encoding["pixel_value_llava"].to(self.model.device),
|
||||
output_hidden_states=output_hidden_states or hidden_state_skip_layer is not None)
|
||||
else:
|
||||
outputs = self.model(
|
||||
input_ids=batch_encoding["input_ids"].to(self.model.device),
|
||||
attention_mask=attention_mask,
|
||||
output_hidden_states=output_hidden_states or hidden_state_skip_layer is not None,)
|
||||
|
||||
if hidden_state_skip_layer is not None:
|
||||
last_hidden_state = outputs.hidden_states[
|
||||
-(hidden_state_skip_layer + 1)
|
||||
]
|
||||
# Real last hidden state already has layer norm applied. So here we only apply it
|
||||
# for intermediate layers.
|
||||
if hidden_state_skip_layer > 0 and self.apply_final_norm:
|
||||
last_hidden_state = self.model.final_layer_norm(last_hidden_state)
|
||||
else:
|
||||
last_hidden_state = outputs[self.output_key]
|
||||
|
||||
# Remove hidden states of instruction tokens, only keep prompt tokens.
|
||||
if self.use_template:
|
||||
if data_type == "image":
|
||||
crop_start = self.prompt_template.get("crop_start", -1)
|
||||
elif data_type == "video":
|
||||
crop_start = self.prompt_template_video.get("crop_start", -1)
|
||||
else:
|
||||
raise ValueError(f"Unsupported data type: {data_type}")
|
||||
if crop_start > 0:
|
||||
last_hidden_state = last_hidden_state[:, crop_start:]
|
||||
attention_mask = (
|
||||
attention_mask[:, crop_start:] if use_attention_mask else None
|
||||
)
|
||||
|
||||
if output_hidden_states:
|
||||
return TextEncoderModelOutput(
|
||||
last_hidden_state, attention_mask, outputs.hidden_states
|
||||
)
|
||||
return TextEncoderModelOutput(last_hidden_state, attention_mask)
|
||||
else:
|
||||
image_outputs = self.processor(semantic_images, return_tensors="pt")[
|
||||
"pixel_values"
|
||||
].to(device)
|
||||
attention_mask = (
|
||||
batch_encoding["attention_mask"].to(device)
|
||||
if use_attention_mask
|
||||
else None
|
||||
)
|
||||
outputs = self.model(
|
||||
input_ids=batch_encoding["input_ids"].to(device),
|
||||
attention_mask=attention_mask,
|
||||
output_hidden_states=output_hidden_states
|
||||
or hidden_state_skip_layer is not None,
|
||||
pixel_values=image_outputs,
|
||||
)
|
||||
if hidden_state_skip_layer is not None:
|
||||
last_hidden_state = outputs.hidden_states[
|
||||
-(hidden_state_skip_layer + 1)
|
||||
]
|
||||
# Real last hidden state already has layer norm applied. So here we only apply it
|
||||
# for intermediate layers.
|
||||
if hidden_state_skip_layer > 0 and self.apply_final_norm:
|
||||
last_hidden_state = self.model.final_layer_norm(last_hidden_state)
|
||||
else:
|
||||
last_hidden_state = outputs[self.output_key]
|
||||
if self.use_template:
|
||||
if data_type == "video":
|
||||
crop_start = self.prompt_template_video.get("crop_start", -1)
|
||||
text_crop_start = (
|
||||
crop_start
|
||||
- 1
|
||||
+ self.prompt_template_video.get("image_emb_len", 576)
|
||||
)
|
||||
image_crop_start = self.prompt_template_video.get(
|
||||
"image_emb_start", 5
|
||||
)
|
||||
image_crop_end = self.prompt_template_video.get(
|
||||
"image_emb_end", 581
|
||||
)
|
||||
batch_indices, last_double_return_token_indices = torch.where(
|
||||
batch_encoding["input_ids"]
|
||||
== self.prompt_template_video.get("double_return_token_id", 271)
|
||||
)
|
||||
if last_double_return_token_indices.shape[0] == 3:
|
||||
# in case the prompt is too long
|
||||
last_double_return_token_indices = torch.cat(
|
||||
(
|
||||
last_double_return_token_indices,
|
||||
torch.tensor([batch_encoding["input_ids"].shape[-1]]),
|
||||
)
|
||||
)
|
||||
batch_indices = torch.cat((batch_indices, torch.tensor([0])))
|
||||
last_double_return_token_indices = (
|
||||
last_double_return_token_indices.reshape(
|
||||
batch_encoding["input_ids"].shape[0], -1
|
||||
)[:, -1]
|
||||
)
|
||||
batch_indices = batch_indices.reshape(
|
||||
batch_encoding["input_ids"].shape[0], -1
|
||||
)[:, -1]
|
||||
assistant_crop_start = (
|
||||
last_double_return_token_indices
|
||||
- 1
|
||||
+ self.prompt_template_video.get("image_emb_len", 576)
|
||||
- 4
|
||||
)
|
||||
assistant_crop_end = (
|
||||
last_double_return_token_indices
|
||||
- 1
|
||||
+ self.prompt_template_video.get("image_emb_len", 576)
|
||||
)
|
||||
attention_mask_assistant_crop_start = (
|
||||
last_double_return_token_indices - 4
|
||||
)
|
||||
attention_mask_assistant_crop_end = last_double_return_token_indices
|
||||
else:
|
||||
raise ValueError(f"Unsupported data type: {data_type}")
|
||||
text_last_hidden_state = []
|
||||
|
||||
text_attention_mask = []
|
||||
image_last_hidden_state = []
|
||||
image_attention_mask = []
|
||||
for i in range(batch_encoding["input_ids"].shape[0]):
|
||||
text_last_hidden_state.append(
|
||||
torch.cat(
|
||||
[
|
||||
last_hidden_state[
|
||||
i, text_crop_start : assistant_crop_start[i].item()
|
||||
],
|
||||
last_hidden_state[i, assistant_crop_end[i].item() :],
|
||||
]
|
||||
)
|
||||
)
|
||||
text_attention_mask.append(
|
||||
torch.cat(
|
||||
[
|
||||
attention_mask[
|
||||
i,
|
||||
crop_start : attention_mask_assistant_crop_start[
|
||||
i
|
||||
].item(),
|
||||
],
|
||||
attention_mask[
|
||||
i, attention_mask_assistant_crop_end[i].item() :
|
||||
],
|
||||
]
|
||||
)
|
||||
if use_attention_mask
|
||||
else None
|
||||
)
|
||||
image_last_hidden_state.append(
|
||||
last_hidden_state[i, image_crop_start:image_crop_end]
|
||||
)
|
||||
image_attention_mask.append(
|
||||
torch.ones(image_last_hidden_state[-1].shape[0])
|
||||
.to(last_hidden_state.device)
|
||||
.to(attention_mask.dtype)
|
||||
if use_attention_mask
|
||||
else None
|
||||
)
|
||||
|
||||
text_last_hidden_state = torch.stack(text_last_hidden_state)
|
||||
text_attention_mask = torch.stack(text_attention_mask)
|
||||
image_last_hidden_state = torch.stack(image_last_hidden_state)
|
||||
image_attention_mask = torch.stack(image_attention_mask)
|
||||
|
||||
if semantic_images is not None and 0 < self.image_embed_interleave < 6:
|
||||
image_last_hidden_state = image_last_hidden_state[
|
||||
:, ::self.image_embed_interleave, :
|
||||
]
|
||||
image_attention_mask = image_attention_mask[
|
||||
:, ::self.image_embed_interleave
|
||||
]
|
||||
|
||||
assert (
|
||||
text_last_hidden_state.shape[0] == text_attention_mask.shape[0]
|
||||
and image_last_hidden_state.shape[0]
|
||||
== image_attention_mask.shape[0]
|
||||
)
|
||||
|
||||
last_hidden_state = torch.cat(
|
||||
[image_last_hidden_state, text_last_hidden_state], dim=1
|
||||
)
|
||||
attention_mask = torch.cat(
|
||||
[image_attention_mask, text_attention_mask], dim=1
|
||||
)
|
||||
if output_hidden_states:
|
||||
return TextEncoderModelOutput(
|
||||
last_hidden_state,
|
||||
attention_mask,
|
||||
hidden_states_list=outputs.hidden_states,
|
||||
)
|
||||
return TextEncoderModelOutput(last_hidden_state, attention_mask)
|
||||
|
||||
def forward(
|
||||
self,
|
||||
text,
|
||||
use_attention_mask=None,
|
||||
output_hidden_states=False,
|
||||
do_sample=False,
|
||||
hidden_state_skip_layer=None,
|
||||
return_texts=False,
|
||||
):
|
||||
batch_encoding = self.text2tokens(text)
|
||||
return self.encode(
|
||||
batch_encoding,
|
||||
use_attention_mask=use_attention_mask,
|
||||
output_hidden_states=output_hidden_states,
|
||||
do_sample=do_sample,
|
||||
hidden_state_skip_layer=hidden_state_skip_layer,
|
||||
return_texts=return_texts,
|
||||
)
|
||||
0
hyvideo/utils/__init__.py
Normal file
0
hyvideo/utils/__init__.py
Normal file
90
hyvideo/utils/data_utils.py
Normal file
90
hyvideo/utils/data_utils.py
Normal file
@@ -0,0 +1,90 @@
|
||||
import numpy as np
|
||||
import math
|
||||
from PIL import Image
|
||||
import torch
|
||||
import copy
|
||||
import string
|
||||
import random
|
||||
|
||||
|
||||
def align_to(value, alignment):
|
||||
"""align hight, width according to alignment
|
||||
|
||||
Args:
|
||||
value (int): height or width
|
||||
alignment (int): target alignment factor
|
||||
|
||||
Returns:
|
||||
int: the aligned value
|
||||
"""
|
||||
return int(math.ceil(value / alignment) * alignment)
|
||||
|
||||
|
||||
def black_image(width, height):
|
||||
"""generate a black image
|
||||
|
||||
Args:
|
||||
width (int): image width
|
||||
height (int): image height
|
||||
|
||||
Returns:
|
||||
_type_: a black image
|
||||
"""
|
||||
black_image = Image.new("RGB", (width, height), (0, 0, 0))
|
||||
return black_image
|
||||
|
||||
|
||||
def get_closest_ratio(height: float, width: float, ratios: list, buckets: list):
|
||||
"""get the closest ratio in the buckets
|
||||
|
||||
Args:
|
||||
height (float): video height
|
||||
width (float): video width
|
||||
ratios (list): video aspect ratio
|
||||
buckets (list): buckets generate by `generate_crop_size_list`
|
||||
|
||||
Returns:
|
||||
the closest ratio in the buckets and the corresponding ratio
|
||||
"""
|
||||
aspect_ratio = float(height) / float(width)
|
||||
closest_ratio_id = np.abs(ratios - aspect_ratio).argmin()
|
||||
closest_ratio = min(ratios, key=lambda ratio: abs(float(ratio) - aspect_ratio))
|
||||
return buckets[closest_ratio_id], float(closest_ratio)
|
||||
|
||||
|
||||
def generate_crop_size_list(base_size=256, patch_size=32, max_ratio=4.0):
|
||||
"""generate crop size list
|
||||
|
||||
Args:
|
||||
base_size (int, optional): the base size for generate bucket. Defaults to 256.
|
||||
patch_size (int, optional): the stride to generate bucket. Defaults to 32.
|
||||
max_ratio (float, optional): th max ratio for h or w based on base_size . Defaults to 4.0.
|
||||
|
||||
Returns:
|
||||
list: generate crop size list
|
||||
"""
|
||||
num_patches = round((base_size / patch_size) ** 2)
|
||||
assert max_ratio >= 1.0
|
||||
crop_size_list = []
|
||||
wp, hp = num_patches, 1
|
||||
while wp > 0:
|
||||
if max(wp, hp) / min(wp, hp) <= max_ratio:
|
||||
crop_size_list.append((wp * patch_size, hp * patch_size))
|
||||
if (hp + 1) * wp <= num_patches:
|
||||
hp += 1
|
||||
else:
|
||||
wp -= 1
|
||||
return crop_size_list
|
||||
|
||||
|
||||
def align_floor_to(value, alignment):
|
||||
"""align hight, width according to alignment
|
||||
|
||||
Args:
|
||||
value (int): height or width
|
||||
alignment (int): target alignment factor
|
||||
|
||||
Returns:
|
||||
int: the aligned value
|
||||
"""
|
||||
return int(math.floor(value / alignment) * alignment)
|
||||
70
hyvideo/utils/file_utils.py
Normal file
70
hyvideo/utils/file_utils.py
Normal file
@@ -0,0 +1,70 @@
|
||||
import os
|
||||
from pathlib import Path
|
||||
from einops import rearrange
|
||||
|
||||
import torch
|
||||
import torchvision
|
||||
import numpy as np
|
||||
import imageio
|
||||
|
||||
CODE_SUFFIXES = {
|
||||
".py", # Python codes
|
||||
".sh", # Shell scripts
|
||||
".yaml",
|
||||
".yml", # Configuration files
|
||||
}
|
||||
|
||||
|
||||
def safe_dir(path):
|
||||
"""
|
||||
Create a directory (or the parent directory of a file) if it does not exist.
|
||||
|
||||
Args:
|
||||
path (str or Path): Path to the directory.
|
||||
|
||||
Returns:
|
||||
path (Path): Path object of the directory.
|
||||
"""
|
||||
path = Path(path)
|
||||
path.mkdir(exist_ok=True, parents=True)
|
||||
return path
|
||||
|
||||
|
||||
def safe_file(path):
|
||||
"""
|
||||
Create the parent directory of a file if it does not exist.
|
||||
|
||||
Args:
|
||||
path (str or Path): Path to the file.
|
||||
|
||||
Returns:
|
||||
path (Path): Path object of the file.
|
||||
"""
|
||||
path = Path(path)
|
||||
path.parent.mkdir(exist_ok=True, parents=True)
|
||||
return path
|
||||
|
||||
def save_videos_grid(videos: torch.Tensor, path: str, rescale=False, n_rows=1, fps=24):
|
||||
"""save videos by video tensor
|
||||
copy from https://github.com/guoyww/AnimateDiff/blob/e92bd5671ba62c0d774a32951453e328018b7c5b/animatediff/utils/util.py#L61
|
||||
|
||||
Args:
|
||||
videos (torch.Tensor): video tensor predicted by the model
|
||||
path (str): path to save video
|
||||
rescale (bool, optional): rescale the video tensor from [-1, 1] to . Defaults to False.
|
||||
n_rows (int, optional): Defaults to 1.
|
||||
fps (int, optional): video save fps. Defaults to 8.
|
||||
"""
|
||||
videos = rearrange(videos, "b c t h w -> t b c h w")
|
||||
outputs = []
|
||||
for x in videos:
|
||||
x = torchvision.utils.make_grid(x, nrow=n_rows)
|
||||
x = x.transpose(0, 1).transpose(1, 2).squeeze(-1)
|
||||
if rescale:
|
||||
x = (x + 1.0) / 2.0 # -1,1 -> 0,1
|
||||
x = torch.clamp(x, 0, 1)
|
||||
x = (x * 255).numpy().astype(np.uint8)
|
||||
outputs.append(x)
|
||||
|
||||
os.makedirs(os.path.dirname(path), exist_ok=True)
|
||||
imageio.mimsave(path, outputs, fps=fps)
|
||||
40
hyvideo/utils/helpers.py
Normal file
40
hyvideo/utils/helpers.py
Normal file
@@ -0,0 +1,40 @@
|
||||
import collections.abc
|
||||
|
||||
from itertools import repeat
|
||||
|
||||
|
||||
def _ntuple(n):
|
||||
def parse(x):
|
||||
if isinstance(x, collections.abc.Iterable) and not isinstance(x, str):
|
||||
x = tuple(x)
|
||||
if len(x) == 1:
|
||||
x = tuple(repeat(x[0], n))
|
||||
return x
|
||||
return tuple(repeat(x, n))
|
||||
return parse
|
||||
|
||||
|
||||
to_1tuple = _ntuple(1)
|
||||
to_2tuple = _ntuple(2)
|
||||
to_3tuple = _ntuple(3)
|
||||
to_4tuple = _ntuple(4)
|
||||
|
||||
|
||||
def as_tuple(x):
|
||||
if isinstance(x, collections.abc.Iterable) and not isinstance(x, str):
|
||||
return tuple(x)
|
||||
if x is None or isinstance(x, (int, float, str)):
|
||||
return (x,)
|
||||
else:
|
||||
raise ValueError(f"Unknown type {type(x)}")
|
||||
|
||||
|
||||
def as_list_of_2tuple(x):
|
||||
x = as_tuple(x)
|
||||
if len(x) == 1:
|
||||
x = (x[0], x[0])
|
||||
assert len(x) % 2 == 0, f"Expect even length, got {len(x)}."
|
||||
lst = []
|
||||
for i in range(0, len(x), 2):
|
||||
lst.append((x[i], x[i + 1]))
|
||||
return lst
|
||||
46
hyvideo/utils/preprocess_text_encoder_tokenizer_utils.py
Normal file
46
hyvideo/utils/preprocess_text_encoder_tokenizer_utils.py
Normal file
@@ -0,0 +1,46 @@
|
||||
import argparse
|
||||
import torch
|
||||
from transformers import (
|
||||
AutoProcessor,
|
||||
LlavaForConditionalGeneration,
|
||||
)
|
||||
|
||||
|
||||
def preprocess_text_encoder_tokenizer(args):
|
||||
|
||||
processor = AutoProcessor.from_pretrained(args.input_dir)
|
||||
model = LlavaForConditionalGeneration.from_pretrained(
|
||||
args.input_dir,
|
||||
torch_dtype=torch.float16,
|
||||
low_cpu_mem_usage=True,
|
||||
).to(0)
|
||||
|
||||
model.language_model.save_pretrained(
|
||||
f"{args.output_dir}"
|
||||
)
|
||||
processor.tokenizer.save_pretrained(
|
||||
f"{args.output_dir}"
|
||||
)
|
||||
|
||||
if __name__ == "__main__":
|
||||
|
||||
parser = argparse.ArgumentParser()
|
||||
parser.add_argument(
|
||||
"--input_dir",
|
||||
type=str,
|
||||
required=True,
|
||||
help="The path to the llava-llama-3-8b-v1_1-transformers.",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--output_dir",
|
||||
type=str,
|
||||
default="",
|
||||
help="The output path of the llava-llama-3-8b-text-encoder-tokenizer."
|
||||
"if '', the parent dir of output will be the same as input dir.",
|
||||
)
|
||||
args = parser.parse_args()
|
||||
|
||||
if len(args.output_dir) == 0:
|
||||
args.output_dir = "/".join(args.input_dir.split("/")[:-1])
|
||||
|
||||
preprocess_text_encoder_tokenizer(args)
|
||||
76
hyvideo/vae/__init__.py
Normal file
76
hyvideo/vae/__init__.py
Normal file
@@ -0,0 +1,76 @@
|
||||
from pathlib import Path
|
||||
|
||||
import torch
|
||||
|
||||
from .autoencoder_kl_causal_3d import AutoencoderKLCausal3D
|
||||
from ..constants import VAE_PATH, PRECISION_TO_TYPE
|
||||
|
||||
def load_vae(vae_type: str="884-16c-hy",
|
||||
vae_precision: str=None,
|
||||
sample_size: tuple=None,
|
||||
vae_path: str=None,
|
||||
vae_config_path: str=None,
|
||||
logger=None,
|
||||
device=None
|
||||
):
|
||||
"""the fucntion to load the 3D VAE model
|
||||
|
||||
Args:
|
||||
vae_type (str): the type of the 3D VAE model. Defaults to "884-16c-hy".
|
||||
vae_precision (str, optional): the precision to load vae. Defaults to None.
|
||||
sample_size (tuple, optional): the tiling size. Defaults to None.
|
||||
vae_path (str, optional): the path to vae. Defaults to None.
|
||||
logger (_type_, optional): logger. Defaults to None.
|
||||
device (_type_, optional): device to load vae. Defaults to None.
|
||||
"""
|
||||
if vae_path is None:
|
||||
vae_path = VAE_PATH[vae_type]
|
||||
|
||||
if logger is not None:
|
||||
logger.info(f"Loading 3D VAE model ({vae_type}) from: {vae_path}")
|
||||
|
||||
# config = AutoencoderKLCausal3D.load_config("ckpts/hunyuan_video_VAE_config.json")
|
||||
# config = AutoencoderKLCausal3D.load_config("c:/temp/hvae/config_vae.json")
|
||||
config = AutoencoderKLCausal3D.load_config(vae_config_path)
|
||||
if sample_size:
|
||||
vae = AutoencoderKLCausal3D.from_config(config, sample_size=sample_size)
|
||||
else:
|
||||
vae = AutoencoderKLCausal3D.from_config(config)
|
||||
|
||||
vae_ckpt = Path(vae_path)
|
||||
# vae_ckpt = Path("ckpts/hunyuan_video_VAE.pt")
|
||||
# vae_ckpt = Path("c:/temp/hvae/pytorch_model.pt")
|
||||
assert vae_ckpt.exists(), f"VAE checkpoint not found: {vae_ckpt}"
|
||||
|
||||
from mmgp import offload
|
||||
|
||||
# ckpt = torch.load(vae_ckpt, weights_only=True, map_location=vae.device)
|
||||
# if "state_dict" in ckpt:
|
||||
# ckpt = ckpt["state_dict"]
|
||||
# if any(k.startswith("vae.") for k in ckpt.keys()):
|
||||
# ckpt = {k.replace("vae.", ""): v for k, v in ckpt.items() if k.startswith("vae.")}
|
||||
# a,b = vae.load_state_dict(ckpt)
|
||||
|
||||
# offload.save_model(vae, "vae_32.safetensors")
|
||||
# vae.to(torch.bfloat16)
|
||||
# offload.save_model(vae, "vae_16.safetensors")
|
||||
offload.load_model_data(vae, vae_path )
|
||||
# ckpt = torch.load(vae_ckpt, weights_only=True, map_location=vae.device)
|
||||
|
||||
spatial_compression_ratio = vae.config.spatial_compression_ratio
|
||||
time_compression_ratio = vae.config.time_compression_ratio
|
||||
|
||||
if vae_precision is not None:
|
||||
vae = vae.to(dtype=PRECISION_TO_TYPE[vae_precision])
|
||||
|
||||
vae.requires_grad_(False)
|
||||
|
||||
if logger is not None:
|
||||
logger.info(f"VAE to dtype: {vae.dtype}")
|
||||
|
||||
if device is not None:
|
||||
vae = vae.to(device)
|
||||
|
||||
vae.eval()
|
||||
|
||||
return vae, vae_path, spatial_compression_ratio, time_compression_ratio
|
||||
927
hyvideo/vae/autoencoder_kl_causal_3d.py
Normal file
927
hyvideo/vae/autoencoder_kl_causal_3d.py
Normal file
@@ -0,0 +1,927 @@
|
||||
import os
|
||||
import math
|
||||
from typing import Dict, Optional, Tuple, Union
|
||||
from dataclasses import dataclass
|
||||
from torch import distributed as dist
|
||||
import loguru
|
||||
import torch
|
||||
import torch.nn as nn
|
||||
import torch.distributed
|
||||
|
||||
RECOMMENDED_DTYPE = torch.float16
|
||||
|
||||
def mpi_comm():
|
||||
from mpi4py import MPI
|
||||
return MPI.COMM_WORLD
|
||||
|
||||
from torch import distributed as dist
|
||||
def mpi_rank():
|
||||
return dist.get_rank()
|
||||
|
||||
def mpi_world_size():
|
||||
return dist.get_world_size()
|
||||
|
||||
|
||||
class TorchIGather:
|
||||
def __init__(self):
|
||||
if not torch.distributed.is_initialized():
|
||||
rank = mpi_rank()
|
||||
world_size = mpi_world_size()
|
||||
os.environ['RANK'] = str(rank)
|
||||
os.environ['WORLD_SIZE'] = str(world_size)
|
||||
os.environ['MASTER_ADDR'] = '127.0.0.1'
|
||||
os.environ['MASTER_PORT'] = str(29500)
|
||||
torch.cuda.set_device(rank)
|
||||
torch.distributed.init_process_group('nccl')
|
||||
|
||||
self.handles = []
|
||||
self.buffers = []
|
||||
|
||||
self.world_size = dist.get_world_size()
|
||||
self.rank = dist.get_rank()
|
||||
self.groups_ids = []
|
||||
self.group = {}
|
||||
|
||||
for i in range(self.world_size):
|
||||
self.groups_ids.append(tuple(range(i + 1)))
|
||||
|
||||
for group in self.groups_ids:
|
||||
new_group = dist.new_group(group)
|
||||
self.group[group[-1]] = new_group
|
||||
|
||||
|
||||
def gather(self, tensor, n_rank=None):
|
||||
if n_rank is not None:
|
||||
group = self.group[n_rank - 1]
|
||||
else:
|
||||
group = None
|
||||
rank = self.rank
|
||||
tensor = tensor.to(RECOMMENDED_DTYPE)
|
||||
if rank == 0:
|
||||
buffer = [torch.empty_like(tensor) for i in range(n_rank)]
|
||||
else:
|
||||
buffer = None
|
||||
self.buffers.append(buffer)
|
||||
handle = torch.distributed.gather(tensor, buffer, async_op=True, group=group)
|
||||
self.handles.append(handle)
|
||||
|
||||
def wait(self):
|
||||
for handle in self.handles:
|
||||
handle.wait()
|
||||
|
||||
def clear(self):
|
||||
self.buffers = []
|
||||
self.handles = []
|
||||
|
||||
|
||||
from diffusers.configuration_utils import ConfigMixin, register_to_config
|
||||
try:
|
||||
# This diffusers is modified and packed in the mirror.
|
||||
from diffusers.loaders import FromOriginalVAEMixin
|
||||
except ImportError:
|
||||
# Use this to be compatible with the original diffusers.
|
||||
from diffusers.loaders.single_file_model import FromOriginalModelMixin as FromOriginalVAEMixin
|
||||
from diffusers.utils.accelerate_utils import apply_forward_hook
|
||||
from diffusers.models.attention_processor import (
|
||||
ADDED_KV_ATTENTION_PROCESSORS,
|
||||
CROSS_ATTENTION_PROCESSORS,
|
||||
Attention,
|
||||
AttentionProcessor,
|
||||
AttnAddedKVProcessor,
|
||||
AttnProcessor,
|
||||
)
|
||||
from diffusers.models.modeling_outputs import AutoencoderKLOutput
|
||||
from diffusers.models.modeling_utils import ModelMixin
|
||||
from .vae import DecoderCausal3D, BaseOutput, DecoderOutput, DiagonalGaussianDistribution, EncoderCausal3D
|
||||
|
||||
# """
|
||||
# use trt need install polygraphy and onnx-graphsurgeon
|
||||
# python3 -m pip install --upgrade polygraphy>=0.47.0 onnx-graphsurgeon --extra-index-url https://pypi.ngc.nvidia.com
|
||||
# """
|
||||
# try:
|
||||
# from polygraphy.backend.trt import ( TrtRunner, EngineFromBytes)
|
||||
# from polygraphy.backend.common import BytesFromPath
|
||||
# except:
|
||||
# print("TrtRunner or EngineFromBytes is not available, you can not use trt engine")
|
||||
|
||||
@dataclass
|
||||
class DecoderOutput2(BaseOutput):
|
||||
sample: torch.FloatTensor
|
||||
posterior: Optional[DiagonalGaussianDistribution] = None
|
||||
|
||||
|
||||
MODEL_OUTPUT_PATH = os.environ.get('MODEL_OUTPUT_PATH')
|
||||
MODEL_BASE = os.environ.get('MODEL_BASE')
|
||||
|
||||
|
||||
class AutoencoderKLCausal3D(ModelMixin, ConfigMixin, FromOriginalVAEMixin):
|
||||
r"""
|
||||
A VAE model with KL loss for encoding images into latents and decoding latent representations into images.
|
||||
|
||||
This model inherits from [`ModelMixin`]. Check the superclass documentation for it's generic methods implemented
|
||||
for all models (such as downloading or saving).
|
||||
|
||||
Parameters:
|
||||
in_channels (int, *optional*, defaults to 3): Number of channels in the input image.
|
||||
out_channels (int, *optional*, defaults to 3): Number of channels in the output.
|
||||
down_block_types (`Tuple[str]`, *optional*, defaults to `("DownEncoderBlock2D",)`):
|
||||
Tuple of downsample block types.
|
||||
up_block_types (`Tuple[str]`, *optional*, defaults to `("UpDecoderBlock2D",)`):
|
||||
Tuple of upsample block types.
|
||||
block_out_channels (`Tuple[int]`, *optional*, defaults to `(64,)`):
|
||||
Tuple of block output channels.
|
||||
act_fn (`str`, *optional*, defaults to `"silu"`): The activation function to use.
|
||||
latent_channels (`int`, *optional*, defaults to 4): Number of channels in the latent space.
|
||||
sample_size (`int`, *optional*, defaults to `32`): Sample input size.
|
||||
scaling_factor (`float`, *optional*, defaults to 0.18215):
|
||||
The component-wise standard deviation of the trained latent space computed using the first batch of the
|
||||
training set. This is used to scale the latent space to have unit variance when training the diffusion
|
||||
model. The latents are scaled with the formula `z = z * scaling_factor` before being passed to the
|
||||
diffusion model. When decoding, the latents are scaled back to the original scale with the formula: `z = 1
|
||||
/ scaling_factor * z`. For more details, refer to sections 4.3.2 and D.1 of the [High-Resolution Image
|
||||
Synthesis with Latent Diffusion Models](https://arxiv.org/abs/2112.10752) paper.
|
||||
force_upcast (`bool`, *optional*, default to `True`):
|
||||
If enabled it will force the VAE to run in float32 for high image resolution pipelines, such as SD-XL. VAE
|
||||
can be fine-tuned / trained to a lower range without loosing too much precision in which case
|
||||
`force_upcast` can be set to `False` - see: https://huggingface.co/madebyollin/sdxl-vae-fp16-fix
|
||||
"""
|
||||
|
||||
def get_VAE_tile_size(self, vae_config, device_mem_capacity, mixed_precision):
|
||||
if mixed_precision:
|
||||
device_mem_capacity /= 1.5
|
||||
if vae_config == 0:
|
||||
if device_mem_capacity >= 24000:
|
||||
use_vae_config = 1
|
||||
elif device_mem_capacity >= 12000:
|
||||
use_vae_config = 2
|
||||
else:
|
||||
use_vae_config = 3
|
||||
else:
|
||||
use_vae_config = vae_config
|
||||
|
||||
if use_vae_config == 1:
|
||||
sample_tsize = 32
|
||||
sample_size = 256
|
||||
elif use_vae_config == 2:
|
||||
sample_tsize = 16
|
||||
sample_size = 256
|
||||
else:
|
||||
sample_tsize = 16
|
||||
sample_size = 192
|
||||
|
||||
VAE_tiling = {
|
||||
"tile_sample_min_tsize" : sample_tsize,
|
||||
"tile_latent_min_tsize" : sample_tsize // self.time_compression_ratio,
|
||||
"tile_sample_min_size" : sample_size,
|
||||
"tile_latent_min_size" : int(sample_size / (2 ** (len(self.config.block_out_channels) - 1))),
|
||||
"tile_overlap_factor" : 0.25
|
||||
}
|
||||
return VAE_tiling
|
||||
_supports_gradient_checkpointing = True
|
||||
|
||||
@register_to_config
|
||||
def __init__(
|
||||
self,
|
||||
in_channels: int = 3,
|
||||
out_channels: int = 3,
|
||||
down_block_types: Tuple[str] = ("DownEncoderBlockCausal3D",),
|
||||
up_block_types: Tuple[str] = ("UpDecoderBlockCausal3D",),
|
||||
block_out_channels: Tuple[int] = (64,),
|
||||
layers_per_block: int = 1,
|
||||
act_fn: str = "silu",
|
||||
latent_channels: int = 4,
|
||||
norm_num_groups: int = 32,
|
||||
sample_size: int = 32,
|
||||
sample_tsize: int = 64,
|
||||
scaling_factor: float = 0.18215,
|
||||
force_upcast: float = True,
|
||||
spatial_compression_ratio: int = 8,
|
||||
time_compression_ratio: int = 4,
|
||||
disable_causal_conv: bool = False,
|
||||
mid_block_add_attention: bool = True,
|
||||
mid_block_causal_attn: bool = False,
|
||||
use_trt_engine: bool = False,
|
||||
nccl_gather: bool = True,
|
||||
engine_path: str = f"{MODEL_BASE}/HYVAE_decoder+conv_256x256xT_fp16_H20.engine",
|
||||
):
|
||||
super().__init__()
|
||||
|
||||
self.disable_causal_conv = disable_causal_conv
|
||||
self.time_compression_ratio = time_compression_ratio
|
||||
|
||||
self.encoder = EncoderCausal3D(
|
||||
in_channels=in_channels,
|
||||
out_channels=latent_channels,
|
||||
down_block_types=down_block_types,
|
||||
block_out_channels=block_out_channels,
|
||||
layers_per_block=layers_per_block,
|
||||
act_fn=act_fn,
|
||||
norm_num_groups=norm_num_groups,
|
||||
double_z=True,
|
||||
time_compression_ratio=time_compression_ratio,
|
||||
spatial_compression_ratio=spatial_compression_ratio,
|
||||
disable_causal=disable_causal_conv,
|
||||
mid_block_add_attention=mid_block_add_attention,
|
||||
mid_block_causal_attn=mid_block_causal_attn,
|
||||
)
|
||||
|
||||
self.decoder = DecoderCausal3D(
|
||||
in_channels=latent_channels,
|
||||
out_channels=out_channels,
|
||||
up_block_types=up_block_types,
|
||||
block_out_channels=block_out_channels,
|
||||
layers_per_block=layers_per_block,
|
||||
norm_num_groups=norm_num_groups,
|
||||
act_fn=act_fn,
|
||||
time_compression_ratio=time_compression_ratio,
|
||||
spatial_compression_ratio=spatial_compression_ratio,
|
||||
disable_causal=disable_causal_conv,
|
||||
mid_block_add_attention=mid_block_add_attention,
|
||||
mid_block_causal_attn=mid_block_causal_attn,
|
||||
)
|
||||
|
||||
self.quant_conv = nn.Conv3d(2 * latent_channels, 2 * latent_channels, kernel_size=1)
|
||||
self.post_quant_conv = nn.Conv3d(latent_channels, latent_channels, kernel_size=1)
|
||||
|
||||
self.use_slicing = False
|
||||
self.use_spatial_tiling = False
|
||||
self.use_temporal_tiling = False
|
||||
|
||||
|
||||
# only relevant if vae tiling is enabled
|
||||
self.tile_sample_min_tsize = sample_tsize
|
||||
self.tile_latent_min_tsize = sample_tsize // time_compression_ratio
|
||||
|
||||
self.tile_sample_min_size = self.config.sample_size
|
||||
sample_size = (
|
||||
self.config.sample_size[0]
|
||||
if isinstance(self.config.sample_size, (list, tuple))
|
||||
else self.config.sample_size
|
||||
)
|
||||
self.tile_latent_min_size = int(sample_size / (2 ** (len(self.config.block_out_channels) - 1)))
|
||||
self.tile_overlap_factor = 0.25
|
||||
|
||||
use_trt_engine = False #if CPU_OFFLOAD else True
|
||||
# ============= parallism related code ===================
|
||||
self.parallel_decode = use_trt_engine
|
||||
self.nccl_gather = nccl_gather
|
||||
|
||||
# only relevant if parallel_decode is enabled
|
||||
self.gather_to_rank0 = self.parallel_decode
|
||||
|
||||
self.engine_path = engine_path
|
||||
|
||||
self.use_trt_decoder = use_trt_engine
|
||||
|
||||
@property
|
||||
def igather(self):
|
||||
assert self.nccl_gather and self.gather_to_rank0
|
||||
if hasattr(self, '_igather'):
|
||||
return self._igather
|
||||
else:
|
||||
self._igather = TorchIGather()
|
||||
return self._igather
|
||||
|
||||
@property
|
||||
def use_padding(self):
|
||||
return (
|
||||
self.use_trt_decoder
|
||||
# dist.gather demands all processes possess to have the same tile shape.
|
||||
or (self.nccl_gather and self.gather_to_rank0)
|
||||
)
|
||||
|
||||
def _set_gradient_checkpointing(self, module, value=False):
|
||||
if isinstance(module, (EncoderCausal3D, DecoderCausal3D)):
|
||||
module.gradient_checkpointing = value
|
||||
|
||||
def enable_temporal_tiling(self, use_tiling: bool = True):
|
||||
self.use_temporal_tiling = use_tiling
|
||||
|
||||
def disable_temporal_tiling(self):
|
||||
self.enable_temporal_tiling(False)
|
||||
|
||||
def enable_spatial_tiling(self, use_tiling: bool = True):
|
||||
self.use_spatial_tiling = use_tiling
|
||||
|
||||
def disable_spatial_tiling(self):
|
||||
self.enable_spatial_tiling(False)
|
||||
|
||||
def enable_tiling(self, use_tiling: bool = True):
|
||||
r"""
|
||||
Enable tiled VAE decoding. When this option is enabled, the VAE will split the input tensor into tiles to
|
||||
compute decoding and encoding in several steps. This is useful for saving a large amount of memory and to allow
|
||||
processing larger images.
|
||||
"""
|
||||
self.enable_spatial_tiling(use_tiling)
|
||||
self.enable_temporal_tiling(use_tiling)
|
||||
|
||||
def disable_tiling(self):
|
||||
r"""
|
||||
Disable tiled VAE decoding. If `enable_tiling` was previously enabled, this method will go back to computing
|
||||
decoding in one step.
|
||||
"""
|
||||
self.disable_spatial_tiling()
|
||||
self.disable_temporal_tiling()
|
||||
|
||||
def enable_slicing(self):
|
||||
r"""
|
||||
Enable sliced VAE decoding. When this option is enabled, the VAE will split the input tensor in slices to
|
||||
compute decoding in several steps. This is useful to save some memory and allow larger batch sizes.
|
||||
"""
|
||||
self.use_slicing = True
|
||||
|
||||
def disable_slicing(self):
|
||||
r"""
|
||||
Disable sliced VAE decoding. If `enable_slicing` was previously enabled, this method will go back to computing
|
||||
decoding in one step.
|
||||
"""
|
||||
self.use_slicing = False
|
||||
|
||||
|
||||
def load_trt_decoder(self):
|
||||
self.use_trt_decoder = True
|
||||
self.engine = EngineFromBytes(BytesFromPath(self.engine_path))
|
||||
|
||||
self.trt_decoder_runner = TrtRunner(self.engine)
|
||||
self.activate_trt_decoder()
|
||||
|
||||
def disable_trt_decoder(self):
|
||||
self.use_trt_decoder = False
|
||||
del self.engine
|
||||
|
||||
def activate_trt_decoder(self):
|
||||
self.trt_decoder_runner.activate()
|
||||
|
||||
def deactivate_trt_decoder(self):
|
||||
self.trt_decoder_runner.deactivate()
|
||||
|
||||
@property
|
||||
# Copied from diffusers.models.unet_2d_condition.UNet2DConditionModel.attn_processors
|
||||
def attn_processors(self) -> Dict[str, AttentionProcessor]:
|
||||
r"""
|
||||
Returns:
|
||||
`dict` of attention processors: A dictionary containing all attention processors used in the model with
|
||||
indexed by its weight name.
|
||||
"""
|
||||
# set recursively
|
||||
processors = {}
|
||||
|
||||
def fn_recursive_add_processors(name: str, module: torch.nn.Module, processors: Dict[str, AttentionProcessor]):
|
||||
if hasattr(module, "get_processor"):
|
||||
processors[f"{name}.processor"] = module.get_processor(return_deprecated_lora=True)
|
||||
|
||||
for sub_name, child in module.named_children():
|
||||
fn_recursive_add_processors(f"{name}.{sub_name}", child, processors)
|
||||
|
||||
return processors
|
||||
|
||||
for name, module in self.named_children():
|
||||
fn_recursive_add_processors(name, module, processors)
|
||||
|
||||
return processors
|
||||
|
||||
# Copied from diffusers.models.unet_2d_condition.UNet2DConditionModel.set_attn_processor
|
||||
def set_attn_processor(
|
||||
self, processor: Union[AttentionProcessor, Dict[str, AttentionProcessor]], _remove_lora=False
|
||||
):
|
||||
r"""
|
||||
Sets the attention processor to use to compute attention.
|
||||
|
||||
Parameters:
|
||||
processor (`dict` of `AttentionProcessor` or only `AttentionProcessor`):
|
||||
The instantiated processor class or a dictionary of processor classes that will be set as the processor
|
||||
for **all** `Attention` layers.
|
||||
|
||||
If `processor` is a dict, the key needs to define the path to the corresponding cross attention
|
||||
processor. This is strongly recommended when setting trainable attention processors.
|
||||
|
||||
"""
|
||||
count = len(self.attn_processors.keys())
|
||||
|
||||
if isinstance(processor, dict) and len(processor) != count:
|
||||
raise ValueError(
|
||||
f"A dict of processors was passed, but the number of processors {len(processor)} does not match the"
|
||||
f" number of attention layers: {count}. Please make sure to pass {count} processor classes."
|
||||
)
|
||||
|
||||
def fn_recursive_attn_processor(name: str, module: torch.nn.Module, processor):
|
||||
if hasattr(module, "set_processor"):
|
||||
if not isinstance(processor, dict):
|
||||
module.set_processor(processor, _remove_lora=_remove_lora)
|
||||
else:
|
||||
module.set_processor(processor.pop(f"{name}.processor"), _remove_lora=_remove_lora)
|
||||
|
||||
for sub_name, child in module.named_children():
|
||||
fn_recursive_attn_processor(f"{name}.{sub_name}", child, processor)
|
||||
|
||||
for name, module in self.named_children():
|
||||
fn_recursive_attn_processor(name, module, processor)
|
||||
|
||||
# Copied from diffusers.models.unet_2d_condition.UNet2DConditionModel.set_default_attn_processor
|
||||
def set_default_attn_processor(self):
|
||||
"""
|
||||
Disables custom attention processors and sets the default attention implementation.
|
||||
"""
|
||||
if all(proc.__class__ in ADDED_KV_ATTENTION_PROCESSORS for proc in self.attn_processors.values()):
|
||||
processor = AttnAddedKVProcessor()
|
||||
elif all(proc.__class__ in CROSS_ATTENTION_PROCESSORS for proc in self.attn_processors.values()):
|
||||
processor = AttnProcessor()
|
||||
else:
|
||||
raise ValueError(
|
||||
f"Cannot call `set_default_attn_processor` when attention processors are of type {next(iter(self.attn_processors.values()))}"
|
||||
)
|
||||
|
||||
self.set_attn_processor(processor, _remove_lora=True)
|
||||
|
||||
@apply_forward_hook
|
||||
def encode(
|
||||
self, x: torch.FloatTensor, return_dict: bool = True
|
||||
) -> Union[AutoencoderKLOutput, Tuple[DiagonalGaussianDistribution]]:
|
||||
"""
|
||||
Encode a batch of images into latents.
|
||||
|
||||
Args:
|
||||
x (`torch.FloatTensor`): Input batch of images.
|
||||
return_dict (`bool`, *optional*, defaults to `True`):
|
||||
Whether to return a [`~models.autoencoder_kl.AutoencoderKLOutput`] instead of a plain tuple.
|
||||
|
||||
Returns:
|
||||
The latent representations of the encoded images. If `return_dict` is True, a
|
||||
[`~models.autoencoder_kl.AutoencoderKLOutput`] is returned, otherwise a plain `tuple` is returned.
|
||||
"""
|
||||
assert len(x.shape) == 5, "The input tensor should have 5 dimensions"
|
||||
|
||||
if self.use_temporal_tiling and x.shape[2] > self.tile_sample_min_tsize:
|
||||
return self.temporal_tiled_encode(x, return_dict=return_dict)
|
||||
|
||||
if self.use_spatial_tiling and (x.shape[-1] > self.tile_sample_min_size or x.shape[-2] > self.tile_sample_min_size):
|
||||
return self.spatial_tiled_encode(x, return_dict=return_dict)
|
||||
|
||||
if self.use_slicing and x.shape[0] > 1:
|
||||
encoded_slices = [self.encoder(x_slice) for x_slice in x.split(1)]
|
||||
h = torch.cat(encoded_slices)
|
||||
else:
|
||||
h = self.encoder(x)
|
||||
|
||||
moments = self.quant_conv(h)
|
||||
posterior = DiagonalGaussianDistribution(moments)
|
||||
|
||||
if not return_dict:
|
||||
return (posterior,)
|
||||
|
||||
return AutoencoderKLOutput(latent_dist=posterior)
|
||||
|
||||
def _decode(self, z: torch.FloatTensor, return_dict: bool = True) -> Union[DecoderOutput, torch.FloatTensor]:
|
||||
assert len(z.shape) == 5, "The input tensor should have 5 dimensions"
|
||||
|
||||
if self.use_temporal_tiling and z.shape[2] > self.tile_latent_min_tsize:
|
||||
return self.temporal_tiled_decode(z, return_dict=return_dict)
|
||||
|
||||
if self.use_spatial_tiling and (z.shape[-1] > self.tile_latent_min_size or z.shape[-2] > self.tile_latent_min_size):
|
||||
return self.spatial_tiled_decode(z, return_dict=return_dict)
|
||||
|
||||
if self.use_trt_decoder:
|
||||
# For unknown reason, `copy_outputs_to_host` must be set to True
|
||||
dec = self.trt_decoder_runner.infer({"input": z.to(RECOMMENDED_DTYPE).contiguous()}, copy_outputs_to_host=True)["output"].to(device=z.device, dtype=z.dtype)
|
||||
else:
|
||||
z = self.post_quant_conv(z)
|
||||
dec = self.decoder(z)
|
||||
|
||||
if not return_dict:
|
||||
return (dec,)
|
||||
|
||||
return DecoderOutput(sample=dec)
|
||||
|
||||
@apply_forward_hook
|
||||
def decode(
|
||||
self, z: torch.FloatTensor, return_dict: bool = True, generator=None
|
||||
) -> Union[DecoderOutput, torch.FloatTensor]:
|
||||
"""
|
||||
Decode a batch of images.
|
||||
|
||||
Args:
|
||||
z (`torch.FloatTensor`): Input batch of latent vectors.
|
||||
return_dict (`bool`, *optional*, defaults to `True`):
|
||||
Whether to return a [`~models.vae.DecoderOutput`] instead of a plain tuple.
|
||||
|
||||
Returns:
|
||||
[`~models.vae.DecoderOutput`] or `tuple`:
|
||||
If return_dict is True, a [`~models.vae.DecoderOutput`] is returned, otherwise a plain `tuple` is
|
||||
returned.
|
||||
|
||||
"""
|
||||
|
||||
if self.parallel_decode:
|
||||
if z.dtype != RECOMMENDED_DTYPE:
|
||||
loguru.logger.warning(
|
||||
f'For better performance, using {RECOMMENDED_DTYPE} for both latent features and model parameters is recommended.'
|
||||
f'Current latent dtype {z.dtype}. '
|
||||
f'Please note that the input latent will be cast to {RECOMMENDED_DTYPE} internally when decoding.'
|
||||
)
|
||||
z = z.to(RECOMMENDED_DTYPE)
|
||||
|
||||
if self.use_slicing and z.shape[0] > 1:
|
||||
decoded_slices = [self._decode(z_slice).sample for z_slice in z.split(1)]
|
||||
decoded = torch.cat(decoded_slices)
|
||||
else:
|
||||
decoded = self._decode(z).sample
|
||||
|
||||
if not return_dict:
|
||||
return (decoded,)
|
||||
|
||||
return DecoderOutput(sample=decoded)
|
||||
|
||||
def blend_v(self, a: torch.Tensor, b: torch.Tensor, blend_extent: int) -> torch.Tensor:
|
||||
blend_extent = min(a.shape[-2], b.shape[-2], blend_extent)
|
||||
if blend_extent == 0:
|
||||
return b
|
||||
|
||||
a_region = a[..., -blend_extent:, :]
|
||||
b_region = b[..., :blend_extent, :]
|
||||
|
||||
weights = torch.arange(blend_extent, device=a.device, dtype=a.dtype) / blend_extent
|
||||
weights = weights.view(1, 1, 1, blend_extent, 1)
|
||||
|
||||
blended = a_region * (1 - weights) + b_region * weights
|
||||
|
||||
b[..., :blend_extent, :] = blended
|
||||
return b
|
||||
|
||||
def blend_h(self, a: torch.Tensor, b: torch.Tensor, blend_extent: int) -> torch.Tensor:
|
||||
blend_extent = min(a.shape[-1], b.shape[-1], blend_extent)
|
||||
if blend_extent == 0:
|
||||
return b
|
||||
|
||||
a_region = a[..., -blend_extent:]
|
||||
b_region = b[..., :blend_extent]
|
||||
|
||||
weights = torch.arange(blend_extent, device=a.device, dtype=a.dtype) / blend_extent
|
||||
weights = weights.view(1, 1, 1, 1, blend_extent)
|
||||
|
||||
blended = a_region * (1 - weights) + b_region * weights
|
||||
|
||||
b[..., :blend_extent] = blended
|
||||
return b
|
||||
def blend_t(self, a: torch.Tensor, b: torch.Tensor, blend_extent: int) -> torch.Tensor:
|
||||
blend_extent = min(a.shape[-3], b.shape[-3], blend_extent)
|
||||
if blend_extent == 0:
|
||||
return b
|
||||
|
||||
a_region = a[..., -blend_extent:, :, :]
|
||||
b_region = b[..., :blend_extent, :, :]
|
||||
|
||||
weights = torch.arange(blend_extent, device=a.device, dtype=a.dtype) / blend_extent
|
||||
weights = weights.view(1, 1, blend_extent, 1, 1)
|
||||
|
||||
blended = a_region * (1 - weights) + b_region * weights
|
||||
|
||||
b[..., :blend_extent, :, :] = blended
|
||||
return b
|
||||
|
||||
def spatial_tiled_encode(self, x: torch.FloatTensor, return_dict: bool = True, return_moments: bool = False) -> AutoencoderKLOutput:
|
||||
r"""Encode a batch of images using a tiled encoder.
|
||||
|
||||
When this option is enabled, the VAE will split the input tensor into tiles to compute encoding in several
|
||||
steps. This is useful to keep memory use constant regardless of image size. The end result of tiled encoding is
|
||||
different from non-tiled encoding because each tile uses a different encoder. To avoid tiling artifacts, the
|
||||
tiles overlap and are blended together to form a smooth output. You may still see tile-sized changes in the
|
||||
output, but they should be much less noticeable.
|
||||
|
||||
Args:
|
||||
x (`torch.FloatTensor`): Input batch of images.
|
||||
return_dict (`bool`, *optional*, defaults to `True`):
|
||||
Whether or not to return a [`~models.autoencoder_kl.AutoencoderKLOutput`] instead of a plain tuple.
|
||||
|
||||
Returns:
|
||||
[`~models.autoencoder_kl.AutoencoderKLOutput`] or `tuple`:
|
||||
If return_dict is True, a [`~models.autoencoder_kl.AutoencoderKLOutput`] is returned, otherwise a plain
|
||||
`tuple` is returned.
|
||||
"""
|
||||
overlap_size = int(self.tile_sample_min_size * (1 - self.tile_overlap_factor))
|
||||
blend_extent = int(self.tile_latent_min_size * self.tile_overlap_factor)
|
||||
row_limit = self.tile_latent_min_size - blend_extent
|
||||
|
||||
# Split video into tiles and encode them separately.
|
||||
rows = []
|
||||
for i in range(0, x.shape[-2], overlap_size):
|
||||
row = []
|
||||
for j in range(0, x.shape[-1], overlap_size):
|
||||
tile = x[:, :, :, i : i + self.tile_sample_min_size, j : j + self.tile_sample_min_size]
|
||||
tile = self.encoder(tile)
|
||||
tile = self.quant_conv(tile)
|
||||
row.append(tile)
|
||||
rows.append(row)
|
||||
result_rows = []
|
||||
for i, row in enumerate(rows):
|
||||
result_row = []
|
||||
for j, tile in enumerate(row):
|
||||
# blend the above tile and the left tile
|
||||
# to the current tile and add the current tile to the result row
|
||||
if i > 0:
|
||||
tile = self.blend_v(rows[i - 1][j], tile, blend_extent)
|
||||
if j > 0:
|
||||
tile = self.blend_h(row[j - 1], tile, blend_extent)
|
||||
result_row.append(tile[:, :, :, :row_limit, :row_limit])
|
||||
result_rows.append(torch.cat(result_row, dim=-1))
|
||||
|
||||
moments = torch.cat(result_rows, dim=-2)
|
||||
if return_moments:
|
||||
return moments
|
||||
|
||||
posterior = DiagonalGaussianDistribution(moments)
|
||||
if not return_dict:
|
||||
return (posterior,)
|
||||
|
||||
return AutoencoderKLOutput(latent_dist=posterior)
|
||||
|
||||
|
||||
def spatial_tiled_decode(self, z: torch.FloatTensor, return_dict: bool = True) -> Union[DecoderOutput, torch.FloatTensor]:
|
||||
r"""
|
||||
Decode a batch of images using a tiled decoder.
|
||||
|
||||
Args:
|
||||
z (`torch.FloatTensor`): Input batch of latent vectors.
|
||||
return_dict (`bool`, *optional*, defaults to `True`):
|
||||
Whether or not to return a [`~models.vae.DecoderOutput`] instead of a plain tuple.
|
||||
|
||||
Returns:
|
||||
[`~models.vae.DecoderOutput`] or `tuple`:
|
||||
If return_dict is True, a [`~models.vae.DecoderOutput`] is returned, otherwise a plain `tuple` is
|
||||
returned.
|
||||
"""
|
||||
overlap_size = int(self.tile_latent_min_size * (1 - self.tile_overlap_factor))
|
||||
blend_extent = int(self.tile_sample_min_size * self.tile_overlap_factor)
|
||||
row_limit = self.tile_sample_min_size - blend_extent
|
||||
|
||||
# Split z into overlapping tiles and decode them separately.
|
||||
# The tiles have an overlap to avoid seams between tiles.
|
||||
if self.parallel_decode:
|
||||
|
||||
rank = mpi_rank()
|
||||
torch.cuda.set_device(rank) # set device for trt_runner
|
||||
world_size = mpi_world_size()
|
||||
|
||||
tiles = []
|
||||
afters_if_padding = []
|
||||
for i in range(0, z.shape[-2], overlap_size):
|
||||
for j in range(0, z.shape[-1], overlap_size):
|
||||
tile = z[:, :, :, i : i + self.tile_latent_min_size, j : j + self.tile_latent_min_size]
|
||||
|
||||
if self.use_padding and (tile.shape[-2] < self.tile_latent_min_size or tile.shape[-1] < self.tile_latent_min_size):
|
||||
from torch.nn import functional as F
|
||||
after_h = tile.shape[-2] * 8
|
||||
after_w = tile.shape[-1] * 8
|
||||
padding = (0, self.tile_latent_min_size - tile.shape[-1], 0, self.tile_latent_min_size - tile.shape[-2], 0, 0)
|
||||
tile = F.pad(tile, padding, "replicate").to(device=tile.device, dtype=tile.dtype)
|
||||
afters_if_padding.append((after_h, after_w))
|
||||
else:
|
||||
afters_if_padding.append(None)
|
||||
|
||||
tiles.append(tile)
|
||||
|
||||
|
||||
# balance tasks
|
||||
ratio = math.ceil(len(tiles) / world_size)
|
||||
tiles_curr_rank = tiles[rank * ratio: None if rank == world_size - 1 else (rank + 1) * ratio]
|
||||
|
||||
decoded_results = []
|
||||
|
||||
|
||||
total = len(tiles)
|
||||
n_task = ([ratio] * (total // ratio) + ([total % ratio] if total % ratio else []))
|
||||
n_task = n_task + [0] * (8 - len(n_task))
|
||||
|
||||
for i, tile in enumerate(tiles_curr_rank):
|
||||
if self.use_trt_decoder:
|
||||
# For unknown reason, `copy_outputs_to_host` must be set to True
|
||||
decoded = self.trt_decoder_runner.infer(
|
||||
{"input": tile.to(RECOMMENDED_DTYPE).contiguous()},
|
||||
copy_outputs_to_host=True
|
||||
)["output"].to(device=z.device, dtype=z.dtype)
|
||||
decoded_results.append(decoded)
|
||||
else:
|
||||
decoded_results.append(self.decoder(self.post_quant_conv(tile)))
|
||||
|
||||
|
||||
def find(n):
|
||||
return next((i for i, task_n in enumerate(n_task) if task_n < n), len(n_task))
|
||||
|
||||
|
||||
if self.nccl_gather and self.gather_to_rank0:
|
||||
self.igather.gather(decoded, n_rank=find(i + 1))
|
||||
|
||||
if not self.nccl_gather:
|
||||
if self.gather_to_rank0:
|
||||
decoded_results = mpi_comm().gather(decoded_results, root=0)
|
||||
if rank != 0:
|
||||
return DecoderOutput(sample=None)
|
||||
else:
|
||||
decoded_results = mpi_comm().allgather(decoded_results)
|
||||
|
||||
decoded_results = sum(decoded_results, [])
|
||||
else:
|
||||
# [Kevin]:
|
||||
# We expect all tiles obtained from the same rank have the same shape.
|
||||
# Shapes among ranks can differ due to the imbalance of task assignment.
|
||||
if self.gather_to_rank0:
|
||||
if rank == 0:
|
||||
self.igather.wait()
|
||||
gather_results = self.igather.buffers
|
||||
self.igather.clear()
|
||||
else:
|
||||
raise NotImplementedError('The old `allgather` implementation is deprecated for nccl plan.')
|
||||
|
||||
if rank != 0 and self.gather_to_rank0:
|
||||
return DecoderOutput(sample=None)
|
||||
|
||||
decoded_results = [col[i] for i in range(max([len(k) for k in gather_results])) for col in gather_results if i < len(col)]
|
||||
|
||||
|
||||
# Crop the padding region in pixel level
|
||||
if self.use_padding:
|
||||
new_decoded_results = []
|
||||
for after, dec in zip(afters_if_padding, decoded_results):
|
||||
if after is not None:
|
||||
after_h, after_w = after
|
||||
new_decoded_results.append(dec[:, :, :, :after_h, :after_w])
|
||||
else:
|
||||
new_decoded_results.append(dec)
|
||||
decoded_results = new_decoded_results
|
||||
|
||||
rows = []
|
||||
decoded_results_iter = iter(decoded_results)
|
||||
for i in range(0, z.shape[-2], overlap_size):
|
||||
row = []
|
||||
for j in range(0, z.shape[-1], overlap_size):
|
||||
row.append(next(decoded_results_iter).to(rank))
|
||||
rows.append(row)
|
||||
else:
|
||||
rows = []
|
||||
for i in range(0, z.shape[-2], overlap_size):
|
||||
row = []
|
||||
for j in range(0, z.shape[-1], overlap_size):
|
||||
tile = z[:, :, :, i : i + self.tile_latent_min_size, j : j + self.tile_latent_min_size]
|
||||
tile = self.post_quant_conv(tile)
|
||||
decoded = self.decoder(tile)
|
||||
row.append(decoded)
|
||||
rows.append(row)
|
||||
|
||||
result_rows = []
|
||||
for i, row in enumerate(rows):
|
||||
result_row = []
|
||||
for j, tile in enumerate(row):
|
||||
# blend the above tile and the left tile
|
||||
# to the current tile and add the current tile to the result row
|
||||
if i > 0:
|
||||
tile = self.blend_v(rows[i - 1][j], tile, blend_extent)
|
||||
if j > 0:
|
||||
tile = self.blend_h(row[j - 1], tile, blend_extent)
|
||||
result_row.append(tile[:, :, :, :row_limit, :row_limit])
|
||||
result_rows.append(torch.cat(result_row, dim=-1))
|
||||
|
||||
dec = torch.cat(result_rows, dim=-2)
|
||||
if not return_dict:
|
||||
return (dec,)
|
||||
|
||||
return DecoderOutput(sample=dec)
|
||||
|
||||
def temporal_tiled_encode(self, x: torch.FloatTensor, return_dict: bool = True) -> AutoencoderKLOutput:
|
||||
assert not self.disable_causal_conv, "Temporal tiling is only compatible with causal convolutions."
|
||||
|
||||
B, C, T, H, W = x.shape
|
||||
overlap_size = int(self.tile_sample_min_tsize * (1 - self.tile_overlap_factor))
|
||||
blend_extent = int(self.tile_latent_min_tsize * self.tile_overlap_factor)
|
||||
t_limit = self.tile_latent_min_tsize - blend_extent
|
||||
|
||||
# Split the video into tiles and encode them separately.
|
||||
row = []
|
||||
for i in range(0, T, overlap_size):
|
||||
tile = x[:, :, i : i + self.tile_sample_min_tsize + 1, :, :]
|
||||
if self.use_spatial_tiling and (tile.shape[-1] > self.tile_sample_min_size or tile.shape[-2] > self.tile_sample_min_size):
|
||||
tile = self.spatial_tiled_encode(tile, return_moments=True)
|
||||
else:
|
||||
tile = self.encoder(tile)
|
||||
tile = self.quant_conv(tile)
|
||||
if i > 0:
|
||||
tile = tile[:, :, 1:, :, :]
|
||||
row.append(tile)
|
||||
result_row = []
|
||||
for i, tile in enumerate(row):
|
||||
if i > 0:
|
||||
tile = self.blend_t(row[i - 1], tile, blend_extent)
|
||||
result_row.append(tile[:, :, :t_limit, :, :])
|
||||
else:
|
||||
result_row.append(tile[:, :, :t_limit+1, :, :])
|
||||
|
||||
moments = torch.cat(result_row, dim=2)
|
||||
posterior = DiagonalGaussianDistribution(moments)
|
||||
|
||||
if not return_dict:
|
||||
return (posterior,)
|
||||
|
||||
return AutoencoderKLOutput(latent_dist=posterior)
|
||||
|
||||
def temporal_tiled_decode(self, z: torch.FloatTensor, return_dict: bool = True) -> Union[DecoderOutput, torch.FloatTensor]:
|
||||
# Split z into overlapping tiles and decode them separately.
|
||||
|
||||
B, C, T, H, W = z.shape
|
||||
overlap_size = int(self.tile_latent_min_tsize * (1 - self.tile_overlap_factor))
|
||||
blend_extent = int(self.tile_sample_min_tsize * self.tile_overlap_factor)
|
||||
t_limit = self.tile_sample_min_tsize - blend_extent
|
||||
|
||||
row = []
|
||||
for i in range(0, T, overlap_size):
|
||||
tile = z[:, :, i: i + self.tile_latent_min_tsize + 1, :, :]
|
||||
if self.use_spatial_tiling and (tile.shape[-1] > self.tile_latent_min_size or tile.shape[-2] > self.tile_latent_min_size):
|
||||
decoded = self.spatial_tiled_decode(tile, return_dict=True).sample
|
||||
else:
|
||||
tile = self.post_quant_conv(tile)
|
||||
decoded = self.decoder(tile)
|
||||
if i > 0:
|
||||
decoded = decoded[:, :, 1:, :, :]
|
||||
row.append(decoded)
|
||||
result_row = []
|
||||
for i, tile in enumerate(row):
|
||||
if i > 0:
|
||||
tile = self.blend_t(row[i - 1], tile, blend_extent)
|
||||
result_row.append(tile[:, :, :t_limit, :, :])
|
||||
else:
|
||||
result_row.append(tile[:, :, :t_limit + 1, :, :])
|
||||
|
||||
dec = torch.cat(result_row, dim=2)
|
||||
if not return_dict:
|
||||
return (dec,)
|
||||
|
||||
return DecoderOutput(sample=dec)
|
||||
|
||||
def forward(
|
||||
self,
|
||||
sample: torch.FloatTensor,
|
||||
sample_posterior: bool = False,
|
||||
return_dict: bool = True,
|
||||
return_posterior: bool = False,
|
||||
generator: Optional[torch.Generator] = None,
|
||||
) -> Union[DecoderOutput2, torch.FloatTensor]:
|
||||
r"""
|
||||
Args:
|
||||
sample (`torch.FloatTensor`): Input sample.
|
||||
sample_posterior (`bool`, *optional*, defaults to `False`):
|
||||
Whether to sample from the posterior.
|
||||
return_dict (`bool`, *optional*, defaults to `True`):
|
||||
Whether or not to return a [`DecoderOutput`] instead of a plain tuple.
|
||||
"""
|
||||
x = sample
|
||||
posterior = self.encode(x).latent_dist
|
||||
if sample_posterior:
|
||||
z = posterior.sample(generator=generator)
|
||||
else:
|
||||
z = posterior.mode()
|
||||
dec = self.decode(z).sample
|
||||
|
||||
if not return_dict:
|
||||
if return_posterior:
|
||||
return (dec, posterior)
|
||||
else:
|
||||
return (dec,)
|
||||
if return_posterior:
|
||||
return DecoderOutput2(sample=dec, posterior=posterior)
|
||||
else:
|
||||
return DecoderOutput2(sample=dec)
|
||||
|
||||
# Copied from diffusers.models.unet_2d_condition.UNet2DConditionModel.fuse_qkv_projections
|
||||
def fuse_qkv_projections(self):
|
||||
"""
|
||||
Enables fused QKV projections. For self-attention modules, all projection matrices (i.e., query,
|
||||
key, value) are fused. For cross-attention modules, key and value projection matrices are fused.
|
||||
|
||||
<Tip warning={true}>
|
||||
|
||||
This API is 🧪 experimental.
|
||||
|
||||
</Tip>
|
||||
"""
|
||||
self.original_attn_processors = None
|
||||
|
||||
for _, attn_processor in self.attn_processors.items():
|
||||
if "Added" in str(attn_processor.__class__.__name__):
|
||||
raise ValueError("`fuse_qkv_projections()` is not supported for models having added KV projections.")
|
||||
|
||||
self.original_attn_processors = self.attn_processors
|
||||
|
||||
for module in self.modules():
|
||||
if isinstance(module, Attention):
|
||||
module.fuse_projections(fuse=True)
|
||||
|
||||
# Copied from diffusers.models.unet_2d_condition.UNet2DConditionModel.unfuse_qkv_projections
|
||||
def unfuse_qkv_projections(self):
|
||||
"""Disables the fused QKV projection if enabled.
|
||||
|
||||
<Tip warning={true}>
|
||||
|
||||
This API is 🧪 experimental.
|
||||
|
||||
</Tip>
|
||||
|
||||
"""
|
||||
if self.original_attn_processors is not None:
|
||||
self.set_attn_processor(self.original_attn_processors)
|
||||
884
hyvideo/vae/unet_causal_3d_blocks.py
Normal file
884
hyvideo/vae/unet_causal_3d_blocks.py
Normal file
@@ -0,0 +1,884 @@
|
||||
# Copyright 2023 The HuggingFace Team. All rights reserved.
|
||||
#
|
||||
# Licensed under the Apache License, Version 2.0 (the "License");
|
||||
# you may not use this file except in compliance with the License.
|
||||
# You may obtain a copy of the License at
|
||||
#
|
||||
# http://www.apache.org/licenses/LICENSE-2.0
|
||||
#
|
||||
# Unless required by applicable law or agreed to in writing, software
|
||||
# distributed under the License is distributed on an "AS IS" BASIS,
|
||||
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
from typing import Any, Dict, Optional, Tuple, Union
|
||||
|
||||
import numpy as np
|
||||
import torch
|
||||
import torch.nn.functional as F
|
||||
from torch import nn
|
||||
from einops import rearrange
|
||||
|
||||
from diffusers.utils import is_torch_version, logging
|
||||
from diffusers.models.activations import get_activation
|
||||
from diffusers.models.attention_processor import SpatialNorm
|
||||
from diffusers.models.attention_processor import Attention
|
||||
from diffusers.models.normalization import AdaGroupNorm
|
||||
from diffusers.models.normalization import RMSNorm
|
||||
|
||||
|
||||
logger = logging.get_logger(__name__) # pylint: disable=invalid-name
|
||||
|
||||
|
||||
def prepare_causal_attention_mask(n_frame: int, n_hw: int, dtype, device, batch_size: int = None):
|
||||
seq_len = n_frame * n_hw
|
||||
mask = torch.full((seq_len, seq_len), float("-inf"), dtype=dtype, device=device)
|
||||
for i in range(seq_len):
|
||||
i_frame = i // n_hw
|
||||
mask[i, : (i_frame + 1) * n_hw] = 0
|
||||
if batch_size is not None:
|
||||
mask = mask.unsqueeze(0).expand(batch_size, -1, -1)
|
||||
return mask
|
||||
|
||||
|
||||
class CausalConv3d(nn.Module):
|
||||
def __init__(
|
||||
self,
|
||||
chan_in,
|
||||
chan_out,
|
||||
kernel_size: Union[int, Tuple[int, int, int]],
|
||||
stride: Union[int, Tuple[int, int, int]] = 1,
|
||||
dilation: Union[int, Tuple[int, int, int]] = 1,
|
||||
pad_mode = 'replicate',
|
||||
disable_causal=False,
|
||||
**kwargs
|
||||
):
|
||||
super().__init__()
|
||||
|
||||
self.pad_mode = pad_mode
|
||||
if disable_causal:
|
||||
padding = (kernel_size // 2, kernel_size // 2, kernel_size // 2, kernel_size // 2, kernel_size // 2, kernel_size // 2)
|
||||
else:
|
||||
padding = (kernel_size // 2, kernel_size // 2, kernel_size // 2, kernel_size // 2, kernel_size - 1, 0) # W, H, T
|
||||
self.time_causal_padding = padding
|
||||
|
||||
self.conv = nn.Conv3d(chan_in, chan_out, kernel_size, stride = stride, dilation = dilation, **kwargs)
|
||||
|
||||
def forward(self, x):
|
||||
x = F.pad(x, self.time_causal_padding, mode=self.pad_mode)
|
||||
return self.conv(x)
|
||||
|
||||
class CausalAvgPool3d(nn.Module):
|
||||
def __init__(
|
||||
self,
|
||||
kernel_size: Union[int, Tuple[int, int, int]],
|
||||
stride: Union[int, Tuple[int, int, int]],
|
||||
pad_mode = 'replicate',
|
||||
disable_causal=False,
|
||||
**kwargs
|
||||
):
|
||||
super().__init__()
|
||||
|
||||
self.pad_mode = pad_mode
|
||||
if disable_causal:
|
||||
padding = (0, 0, 0, 0, 0, 0)
|
||||
else:
|
||||
padding = (0, 0, 0, 0, stride - 1, 0) # W, H, T
|
||||
self.time_causal_padding = padding
|
||||
|
||||
self.conv = nn.AvgPool3d(kernel_size, stride=stride, ceil_mode=True, **kwargs)
|
||||
self.pad_mode = pad_mode
|
||||
|
||||
def forward(self, x):
|
||||
x = F.pad(x, self.time_causal_padding, mode=self.pad_mode)
|
||||
return self.conv(x)
|
||||
|
||||
class UpsampleCausal3D(nn.Module):
|
||||
"""A 3D upsampling layer with an optional convolution.
|
||||
|
||||
Parameters:
|
||||
channels (`int`):
|
||||
number of channels in the inputs and outputs.
|
||||
use_conv (`bool`, default `False`):
|
||||
option to use a convolution.
|
||||
use_conv_transpose (`bool`, default `False`):
|
||||
option to use a convolution transpose.
|
||||
out_channels (`int`, optional):
|
||||
number of output channels. Defaults to `channels`.
|
||||
name (`str`, default `conv`):
|
||||
name of the upsampling 3D layer.
|
||||
"""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
channels: int,
|
||||
use_conv: bool = False,
|
||||
use_conv_transpose: bool = False,
|
||||
out_channels: Optional[int] = None,
|
||||
name: str = "conv",
|
||||
kernel_size: Optional[int] = None,
|
||||
padding=1,
|
||||
norm_type=None,
|
||||
eps=None,
|
||||
elementwise_affine=None,
|
||||
bias=True,
|
||||
interpolate=True,
|
||||
upsample_factor=(2, 2, 2),
|
||||
disable_causal=False,
|
||||
):
|
||||
super().__init__()
|
||||
self.channels = channels
|
||||
self.out_channels = out_channels or channels
|
||||
self.use_conv = use_conv
|
||||
self.use_conv_transpose = use_conv_transpose
|
||||
self.name = name
|
||||
self.interpolate = interpolate
|
||||
self.upsample_factor = upsample_factor
|
||||
self.disable_causal = disable_causal
|
||||
|
||||
if norm_type == "ln_norm":
|
||||
self.norm = nn.LayerNorm(channels, eps, elementwise_affine)
|
||||
elif norm_type == "rms_norm":
|
||||
self.norm = RMSNorm(channels, eps, elementwise_affine)
|
||||
elif norm_type is None:
|
||||
self.norm = None
|
||||
else:
|
||||
raise ValueError(f"unknown norm_type: {norm_type}")
|
||||
|
||||
conv = None
|
||||
if use_conv_transpose:
|
||||
assert False, "Not Implement yet"
|
||||
if kernel_size is None:
|
||||
kernel_size = 4
|
||||
conv = nn.ConvTranspose2d(
|
||||
channels, self.out_channels, kernel_size=kernel_size, stride=2, padding=padding, bias=bias
|
||||
)
|
||||
elif use_conv:
|
||||
if kernel_size is None:
|
||||
kernel_size = 3
|
||||
conv = CausalConv3d(self.channels, self.out_channels, kernel_size=kernel_size, bias=bias, disable_causal=disable_causal)
|
||||
|
||||
if name == "conv":
|
||||
self.conv = conv
|
||||
else:
|
||||
self.Conv2d_0 = conv
|
||||
|
||||
def forward(
|
||||
self,
|
||||
hidden_states: torch.FloatTensor,
|
||||
output_size: Optional[int] = None,
|
||||
scale: float = 1.0,
|
||||
) -> torch.FloatTensor:
|
||||
assert hidden_states.shape[1] == self.channels
|
||||
|
||||
if self.norm is not None:
|
||||
assert False, "Not Implement yet"
|
||||
hidden_states = self.norm(hidden_states.permute(0, 2, 3, 1)).permute(0, 3, 1, 2)
|
||||
|
||||
if self.use_conv_transpose:
|
||||
return self.conv(hidden_states)
|
||||
|
||||
# Cast to float32 to as 'upsample_nearest2d_out_frame' op does not support bfloat16
|
||||
# https://github.com/pytorch/pytorch/issues/86679
|
||||
dtype = hidden_states.dtype
|
||||
if dtype == torch.bfloat16:
|
||||
hidden_states = hidden_states.to(torch.float32)
|
||||
|
||||
# upsample_nearest_nhwc fails with large batch sizes. see https://github.com/huggingface/diffusers/issues/984
|
||||
if hidden_states.shape[0] >= 64:
|
||||
hidden_states = hidden_states.contiguous()
|
||||
|
||||
# if `output_size` is passed we force the interpolation output
|
||||
# size and do not make use of `scale_factor=2`
|
||||
if self.interpolate:
|
||||
B, C, T, H, W = hidden_states.shape
|
||||
if not self.disable_causal:
|
||||
first_h, other_h = hidden_states.split((1, T-1), dim=2)
|
||||
if output_size is None:
|
||||
if T > 1:
|
||||
other_h = F.interpolate(other_h, scale_factor=self.upsample_factor, mode="nearest")
|
||||
|
||||
first_h = first_h.squeeze(2)
|
||||
first_h = F.interpolate(first_h, scale_factor=self.upsample_factor[1:], mode="nearest")
|
||||
first_h = first_h.unsqueeze(2)
|
||||
else:
|
||||
assert False, "Not Implement yet"
|
||||
other_h = F.interpolate(other_h, size=output_size, mode="nearest")
|
||||
|
||||
if T > 1:
|
||||
hidden_states = torch.cat((first_h, other_h), dim=2)
|
||||
else:
|
||||
hidden_states = first_h
|
||||
else:
|
||||
hidden_states = F.interpolate(hidden_states, scale_factor=self.upsample_factor, mode="nearest")
|
||||
|
||||
if dtype == torch.bfloat16:
|
||||
hidden_states = hidden_states.to(dtype)
|
||||
|
||||
if self.use_conv:
|
||||
if self.name == "conv":
|
||||
hidden_states = self.conv(hidden_states)
|
||||
else:
|
||||
hidden_states = self.Conv2d_0(hidden_states)
|
||||
|
||||
return hidden_states
|
||||
|
||||
class DownsampleCausal3D(nn.Module):
|
||||
"""A 3D downsampling layer with an optional convolution.
|
||||
|
||||
Parameters:
|
||||
channels (`int`):
|
||||
number of channels in the inputs and outputs.
|
||||
use_conv (`bool`, default `False`):
|
||||
option to use a convolution.
|
||||
out_channels (`int`, optional):
|
||||
number of output channels. Defaults to `channels`.
|
||||
padding (`int`, default `1`):
|
||||
padding for the convolution.
|
||||
name (`str`, default `conv`):
|
||||
name of the downsampling 3D layer.
|
||||
"""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
channels: int,
|
||||
use_conv: bool = False,
|
||||
out_channels: Optional[int] = None,
|
||||
padding: int = 1,
|
||||
name: str = "conv",
|
||||
kernel_size=3,
|
||||
norm_type=None,
|
||||
eps=None,
|
||||
elementwise_affine=None,
|
||||
bias=True,
|
||||
stride=2,
|
||||
disable_causal=False,
|
||||
):
|
||||
super().__init__()
|
||||
self.channels = channels
|
||||
self.out_channels = out_channels or channels
|
||||
self.use_conv = use_conv
|
||||
self.padding = padding
|
||||
stride = stride
|
||||
self.name = name
|
||||
|
||||
if norm_type == "ln_norm":
|
||||
self.norm = nn.LayerNorm(channels, eps, elementwise_affine)
|
||||
elif norm_type == "rms_norm":
|
||||
self.norm = RMSNorm(channels, eps, elementwise_affine)
|
||||
elif norm_type is None:
|
||||
self.norm = None
|
||||
else:
|
||||
raise ValueError(f"unknown norm_type: {norm_type}")
|
||||
|
||||
if use_conv:
|
||||
conv = CausalConv3d(
|
||||
self.channels, self.out_channels, kernel_size=kernel_size, stride=stride, disable_causal=disable_causal, bias=bias
|
||||
)
|
||||
else:
|
||||
raise NotImplementedError
|
||||
if name == "conv":
|
||||
self.Conv2d_0 = conv
|
||||
self.conv = conv
|
||||
elif name == "Conv2d_0":
|
||||
self.conv = conv
|
||||
else:
|
||||
self.conv = conv
|
||||
|
||||
def forward(self, hidden_states: torch.FloatTensor, scale: float = 1.0) -> torch.FloatTensor:
|
||||
assert hidden_states.shape[1] == self.channels
|
||||
|
||||
if self.norm is not None:
|
||||
hidden_states = self.norm(hidden_states.permute(0, 2, 3, 1)).permute(0, 3, 1, 2)
|
||||
|
||||
assert hidden_states.shape[1] == self.channels
|
||||
|
||||
hidden_states = self.conv(hidden_states)
|
||||
|
||||
return hidden_states
|
||||
|
||||
class ResnetBlockCausal3D(nn.Module):
|
||||
r"""
|
||||
A Resnet block.
|
||||
|
||||
Parameters:
|
||||
in_channels (`int`): The number of channels in the input.
|
||||
out_channels (`int`, *optional*, default to be `None`):
|
||||
The number of output channels for the first conv2d layer. If None, same as `in_channels`.
|
||||
dropout (`float`, *optional*, defaults to `0.0`): The dropout probability to use.
|
||||
temb_channels (`int`, *optional*, default to `512`): the number of channels in timestep embedding.
|
||||
groups (`int`, *optional*, default to `32`): The number of groups to use for the first normalization layer.
|
||||
groups_out (`int`, *optional*, default to None):
|
||||
The number of groups to use for the second normalization layer. if set to None, same as `groups`.
|
||||
eps (`float`, *optional*, defaults to `1e-6`): The epsilon to use for the normalization.
|
||||
non_linearity (`str`, *optional*, default to `"swish"`): the activation function to use.
|
||||
time_embedding_norm (`str`, *optional*, default to `"default"` ): Time scale shift config.
|
||||
By default, apply timestep embedding conditioning with a simple shift mechanism. Choose "scale_shift" or
|
||||
"ada_group" for a stronger conditioning with scale and shift.
|
||||
kernel (`torch.FloatTensor`, optional, default to None): FIR filter, see
|
||||
[`~models.resnet.FirUpsample2D`] and [`~models.resnet.FirDownsample2D`].
|
||||
output_scale_factor (`float`, *optional*, default to be `1.0`): the scale factor to use for the output.
|
||||
use_in_shortcut (`bool`, *optional*, default to `True`):
|
||||
If `True`, add a 1x1 nn.conv2d layer for skip-connection.
|
||||
up (`bool`, *optional*, default to `False`): If `True`, add an upsample layer.
|
||||
down (`bool`, *optional*, default to `False`): If `True`, add a downsample layer.
|
||||
conv_shortcut_bias (`bool`, *optional*, default to `True`): If `True`, adds a learnable bias to the
|
||||
`conv_shortcut` output.
|
||||
conv_3d_out_channels (`int`, *optional*, default to `None`): the number of channels in the output.
|
||||
If None, same as `out_channels`.
|
||||
"""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
*,
|
||||
in_channels: int,
|
||||
out_channels: Optional[int] = None,
|
||||
conv_shortcut: bool = False,
|
||||
dropout: float = 0.0,
|
||||
temb_channels: int = 512,
|
||||
groups: int = 32,
|
||||
groups_out: Optional[int] = None,
|
||||
pre_norm: bool = True,
|
||||
eps: float = 1e-6,
|
||||
non_linearity: str = "swish",
|
||||
skip_time_act: bool = False,
|
||||
time_embedding_norm: str = "default", # default, scale_shift, ada_group, spatial
|
||||
kernel: Optional[torch.FloatTensor] = None,
|
||||
output_scale_factor: float = 1.0,
|
||||
use_in_shortcut: Optional[bool] = None,
|
||||
up: bool = False,
|
||||
down: bool = False,
|
||||
conv_shortcut_bias: bool = True,
|
||||
conv_3d_out_channels: Optional[int] = None,
|
||||
disable_causal: bool = False,
|
||||
):
|
||||
super().__init__()
|
||||
self.pre_norm = pre_norm
|
||||
self.pre_norm = True
|
||||
self.in_channels = in_channels
|
||||
out_channels = in_channels if out_channels is None else out_channels
|
||||
self.out_channels = out_channels
|
||||
self.use_conv_shortcut = conv_shortcut
|
||||
self.up = up
|
||||
self.down = down
|
||||
self.output_scale_factor = output_scale_factor
|
||||
self.time_embedding_norm = time_embedding_norm
|
||||
self.skip_time_act = skip_time_act
|
||||
|
||||
linear_cls = nn.Linear
|
||||
|
||||
if groups_out is None:
|
||||
groups_out = groups
|
||||
|
||||
if self.time_embedding_norm == "ada_group":
|
||||
self.norm1 = AdaGroupNorm(temb_channels, in_channels, groups, eps=eps)
|
||||
elif self.time_embedding_norm == "spatial":
|
||||
self.norm1 = SpatialNorm(in_channels, temb_channels)
|
||||
else:
|
||||
self.norm1 = torch.nn.GroupNorm(num_groups=groups, num_channels=in_channels, eps=eps, affine=True)
|
||||
|
||||
self.conv1 = CausalConv3d(in_channels, out_channels, kernel_size=3, stride=1, disable_causal=disable_causal)
|
||||
|
||||
if temb_channels is not None:
|
||||
if self.time_embedding_norm == "default":
|
||||
self.time_emb_proj = linear_cls(temb_channels, out_channels)
|
||||
elif self.time_embedding_norm == "scale_shift":
|
||||
self.time_emb_proj = linear_cls(temb_channels, 2 * out_channels)
|
||||
elif self.time_embedding_norm == "ada_group" or self.time_embedding_norm == "spatial":
|
||||
self.time_emb_proj = None
|
||||
else:
|
||||
raise ValueError(f"unknown time_embedding_norm : {self.time_embedding_norm} ")
|
||||
else:
|
||||
self.time_emb_proj = None
|
||||
|
||||
if self.time_embedding_norm == "ada_group":
|
||||
self.norm2 = AdaGroupNorm(temb_channels, out_channels, groups_out, eps=eps)
|
||||
elif self.time_embedding_norm == "spatial":
|
||||
self.norm2 = SpatialNorm(out_channels, temb_channels)
|
||||
else:
|
||||
self.norm2 = torch.nn.GroupNorm(num_groups=groups_out, num_channels=out_channels, eps=eps, affine=True)
|
||||
|
||||
self.dropout = torch.nn.Dropout(dropout)
|
||||
conv_3d_out_channels = conv_3d_out_channels or out_channels
|
||||
self.conv2 = CausalConv3d(out_channels, conv_3d_out_channels, kernel_size=3, stride=1, disable_causal=disable_causal)
|
||||
|
||||
self.nonlinearity = get_activation(non_linearity)
|
||||
|
||||
self.upsample = self.downsample = None
|
||||
if self.up:
|
||||
self.upsample = UpsampleCausal3D(in_channels, use_conv=False, disable_causal=disable_causal)
|
||||
elif self.down:
|
||||
self.downsample = DownsampleCausal3D(in_channels, use_conv=False, disable_causal=disable_causal, name="op")
|
||||
|
||||
self.use_in_shortcut = self.in_channels != conv_3d_out_channels if use_in_shortcut is None else use_in_shortcut
|
||||
|
||||
self.conv_shortcut = None
|
||||
if self.use_in_shortcut:
|
||||
self.conv_shortcut = CausalConv3d(
|
||||
in_channels,
|
||||
conv_3d_out_channels,
|
||||
kernel_size=1,
|
||||
stride=1,
|
||||
disable_causal=disable_causal,
|
||||
bias=conv_shortcut_bias,
|
||||
)
|
||||
|
||||
def forward(
|
||||
self,
|
||||
input_tensor: torch.FloatTensor,
|
||||
temb: torch.FloatTensor,
|
||||
scale: float = 1.0,
|
||||
) -> torch.FloatTensor:
|
||||
hidden_states = input_tensor
|
||||
|
||||
if self.time_embedding_norm == "ada_group" or self.time_embedding_norm == "spatial":
|
||||
hidden_states = self.norm1(hidden_states, temb)
|
||||
else:
|
||||
hidden_states = self.norm1(hidden_states)
|
||||
|
||||
hidden_states = self.nonlinearity(hidden_states)
|
||||
|
||||
if self.upsample is not None:
|
||||
# upsample_nearest_nhwc fails with large batch sizes. see https://github.com/huggingface/diffusers/issues/984
|
||||
if hidden_states.shape[0] >= 64:
|
||||
input_tensor = input_tensor.contiguous()
|
||||
hidden_states = hidden_states.contiguous()
|
||||
input_tensor = (
|
||||
self.upsample(input_tensor, scale=scale)
|
||||
)
|
||||
hidden_states = (
|
||||
self.upsample(hidden_states, scale=scale)
|
||||
)
|
||||
elif self.downsample is not None:
|
||||
input_tensor = (
|
||||
self.downsample(input_tensor, scale=scale)
|
||||
)
|
||||
hidden_states = (
|
||||
self.downsample(hidden_states, scale=scale)
|
||||
)
|
||||
|
||||
hidden_states = self.conv1(hidden_states)
|
||||
|
||||
if self.time_emb_proj is not None:
|
||||
if not self.skip_time_act:
|
||||
temb = self.nonlinearity(temb)
|
||||
temb = (
|
||||
self.time_emb_proj(temb, scale)[:, :, None, None]
|
||||
)
|
||||
|
||||
if temb is not None and self.time_embedding_norm == "default":
|
||||
hidden_states = hidden_states + temb
|
||||
|
||||
if self.time_embedding_norm == "ada_group" or self.time_embedding_norm == "spatial":
|
||||
hidden_states = self.norm2(hidden_states, temb)
|
||||
else:
|
||||
hidden_states = self.norm2(hidden_states)
|
||||
|
||||
if temb is not None and self.time_embedding_norm == "scale_shift":
|
||||
scale, shift = torch.chunk(temb, 2, dim=1)
|
||||
hidden_states = hidden_states * (1 + scale) + shift
|
||||
|
||||
hidden_states = self.nonlinearity(hidden_states)
|
||||
|
||||
hidden_states = self.dropout(hidden_states)
|
||||
hidden_states = self.conv2(hidden_states)
|
||||
|
||||
if self.conv_shortcut is not None:
|
||||
input_tensor = (
|
||||
self.conv_shortcut(input_tensor)
|
||||
)
|
||||
|
||||
output_tensor = (input_tensor + hidden_states) / self.output_scale_factor
|
||||
|
||||
return output_tensor
|
||||
|
||||
def get_down_block3d(
|
||||
down_block_type: str,
|
||||
num_layers: int,
|
||||
in_channels: int,
|
||||
out_channels: int,
|
||||
temb_channels: int,
|
||||
add_downsample: bool,
|
||||
downsample_stride: int,
|
||||
resnet_eps: float,
|
||||
resnet_act_fn: str,
|
||||
transformer_layers_per_block: int = 1,
|
||||
num_attention_heads: Optional[int] = None,
|
||||
resnet_groups: Optional[int] = None,
|
||||
cross_attention_dim: Optional[int] = None,
|
||||
downsample_padding: Optional[int] = None,
|
||||
dual_cross_attention: bool = False,
|
||||
use_linear_projection: bool = False,
|
||||
only_cross_attention: bool = False,
|
||||
upcast_attention: bool = False,
|
||||
resnet_time_scale_shift: str = "default",
|
||||
attention_type: str = "default",
|
||||
resnet_skip_time_act: bool = False,
|
||||
resnet_out_scale_factor: float = 1.0,
|
||||
cross_attention_norm: Optional[str] = None,
|
||||
attention_head_dim: Optional[int] = None,
|
||||
downsample_type: Optional[str] = None,
|
||||
dropout: float = 0.0,
|
||||
disable_causal: bool = False,
|
||||
):
|
||||
# If attn head dim is not defined, we default it to the number of heads
|
||||
if attention_head_dim is None:
|
||||
logger.warn(
|
||||
f"It is recommended to provide `attention_head_dim` when calling `get_down_block`. Defaulting `attention_head_dim` to {num_attention_heads}."
|
||||
)
|
||||
attention_head_dim = num_attention_heads
|
||||
|
||||
down_block_type = down_block_type[7:] if down_block_type.startswith("UNetRes") else down_block_type
|
||||
if down_block_type == "DownEncoderBlockCausal3D":
|
||||
return DownEncoderBlockCausal3D(
|
||||
num_layers=num_layers,
|
||||
in_channels=in_channels,
|
||||
out_channels=out_channels,
|
||||
dropout=dropout,
|
||||
add_downsample=add_downsample,
|
||||
downsample_stride=downsample_stride,
|
||||
resnet_eps=resnet_eps,
|
||||
resnet_act_fn=resnet_act_fn,
|
||||
resnet_groups=resnet_groups,
|
||||
downsample_padding=downsample_padding,
|
||||
resnet_time_scale_shift=resnet_time_scale_shift,
|
||||
disable_causal=disable_causal,
|
||||
)
|
||||
raise ValueError(f"{down_block_type} does not exist.")
|
||||
|
||||
def get_up_block3d(
|
||||
up_block_type: str,
|
||||
num_layers: int,
|
||||
in_channels: int,
|
||||
out_channels: int,
|
||||
prev_output_channel: int,
|
||||
temb_channels: int,
|
||||
add_upsample: bool,
|
||||
upsample_scale_factor: Tuple,
|
||||
resnet_eps: float,
|
||||
resnet_act_fn: str,
|
||||
resolution_idx: Optional[int] = None,
|
||||
transformer_layers_per_block: int = 1,
|
||||
num_attention_heads: Optional[int] = None,
|
||||
resnet_groups: Optional[int] = None,
|
||||
cross_attention_dim: Optional[int] = None,
|
||||
dual_cross_attention: bool = False,
|
||||
use_linear_projection: bool = False,
|
||||
only_cross_attention: bool = False,
|
||||
upcast_attention: bool = False,
|
||||
resnet_time_scale_shift: str = "default",
|
||||
attention_type: str = "default",
|
||||
resnet_skip_time_act: bool = False,
|
||||
resnet_out_scale_factor: float = 1.0,
|
||||
cross_attention_norm: Optional[str] = None,
|
||||
attention_head_dim: Optional[int] = None,
|
||||
upsample_type: Optional[str] = None,
|
||||
dropout: float = 0.0,
|
||||
disable_causal: bool = False,
|
||||
) -> nn.Module:
|
||||
# If attn head dim is not defined, we default it to the number of heads
|
||||
if attention_head_dim is None:
|
||||
logger.warn(
|
||||
f"It is recommended to provide `attention_head_dim` when calling `get_up_block`. Defaulting `attention_head_dim` to {num_attention_heads}."
|
||||
)
|
||||
attention_head_dim = num_attention_heads
|
||||
|
||||
up_block_type = up_block_type[7:] if up_block_type.startswith("UNetRes") else up_block_type
|
||||
if up_block_type == "UpDecoderBlockCausal3D":
|
||||
return UpDecoderBlockCausal3D(
|
||||
num_layers=num_layers,
|
||||
in_channels=in_channels,
|
||||
out_channels=out_channels,
|
||||
resolution_idx=resolution_idx,
|
||||
dropout=dropout,
|
||||
add_upsample=add_upsample,
|
||||
upsample_scale_factor=upsample_scale_factor,
|
||||
resnet_eps=resnet_eps,
|
||||
resnet_act_fn=resnet_act_fn,
|
||||
resnet_groups=resnet_groups,
|
||||
resnet_time_scale_shift=resnet_time_scale_shift,
|
||||
temb_channels=temb_channels,
|
||||
disable_causal=disable_causal,
|
||||
)
|
||||
raise ValueError(f"{up_block_type} does not exist.")
|
||||
|
||||
|
||||
class UNetMidBlockCausal3D(nn.Module):
|
||||
"""
|
||||
A 3D UNet mid-block [`UNetMidBlockCausal3D`] with multiple residual blocks and optional attention blocks.
|
||||
|
||||
Args:
|
||||
in_channels (`int`): The number of input channels.
|
||||
temb_channels (`int`): The number of temporal embedding channels.
|
||||
dropout (`float`, *optional*, defaults to 0.0): The dropout rate.
|
||||
num_layers (`int`, *optional*, defaults to 1): The number of residual blocks.
|
||||
resnet_eps (`float`, *optional*, 1e-6 ): The epsilon value for the resnet blocks.
|
||||
resnet_time_scale_shift (`str`, *optional*, defaults to `default`):
|
||||
The type of normalization to apply to the time embeddings. This can help to improve the performance of the
|
||||
model on tasks with long-range temporal dependencies.
|
||||
resnet_act_fn (`str`, *optional*, defaults to `swish`): The activation function for the resnet blocks.
|
||||
resnet_groups (`int`, *optional*, defaults to 32):
|
||||
The number of groups to use in the group normalization layers of the resnet blocks.
|
||||
attn_groups (`Optional[int]`, *optional*, defaults to None): The number of groups for the attention blocks.
|
||||
resnet_pre_norm (`bool`, *optional*, defaults to `True`):
|
||||
Whether to use pre-normalization for the resnet blocks.
|
||||
add_attention (`bool`, *optional*, defaults to `True`): Whether to add attention blocks.
|
||||
attention_head_dim (`int`, *optional*, defaults to 1):
|
||||
Dimension of a single attention head. The number of attention heads is determined based on this value and
|
||||
the number of input channels.
|
||||
output_scale_factor (`float`, *optional*, defaults to 1.0): The output scale factor.
|
||||
|
||||
Returns:
|
||||
`torch.FloatTensor`: The output of the last residual block, which is a tensor of shape `(batch_size,
|
||||
in_channels, height, width)`.
|
||||
|
||||
"""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
in_channels: int,
|
||||
temb_channels: int,
|
||||
dropout: float = 0.0,
|
||||
num_layers: int = 1,
|
||||
resnet_eps: float = 1e-6,
|
||||
resnet_time_scale_shift: str = "default", # default, spatial
|
||||
resnet_act_fn: str = "swish",
|
||||
resnet_groups: int = 32,
|
||||
attn_groups: Optional[int] = None,
|
||||
resnet_pre_norm: bool = True,
|
||||
add_attention: bool = True,
|
||||
attention_head_dim: int = 1,
|
||||
output_scale_factor: float = 1.0,
|
||||
disable_causal: bool = False,
|
||||
causal_attention: bool = False,
|
||||
):
|
||||
super().__init__()
|
||||
resnet_groups = resnet_groups if resnet_groups is not None else min(in_channels // 4, 32)
|
||||
self.add_attention = add_attention
|
||||
self.causal_attention = causal_attention
|
||||
|
||||
if attn_groups is None:
|
||||
attn_groups = resnet_groups if resnet_time_scale_shift == "default" else None
|
||||
|
||||
# there is always at least one resnet
|
||||
resnets = [
|
||||
ResnetBlockCausal3D(
|
||||
in_channels=in_channels,
|
||||
out_channels=in_channels,
|
||||
temb_channels=temb_channels,
|
||||
eps=resnet_eps,
|
||||
groups=resnet_groups,
|
||||
dropout=dropout,
|
||||
time_embedding_norm=resnet_time_scale_shift,
|
||||
non_linearity=resnet_act_fn,
|
||||
output_scale_factor=output_scale_factor,
|
||||
pre_norm=resnet_pre_norm,
|
||||
disable_causal=disable_causal,
|
||||
)
|
||||
]
|
||||
attentions = []
|
||||
|
||||
if attention_head_dim is None:
|
||||
logger.warn(
|
||||
f"It is not recommend to pass `attention_head_dim=None`. Defaulting `attention_head_dim` to `in_channels`: {in_channels}."
|
||||
)
|
||||
attention_head_dim = in_channels
|
||||
|
||||
for _ in range(num_layers):
|
||||
if self.add_attention:
|
||||
#assert False, "Not implemented yet"
|
||||
attentions.append(
|
||||
Attention(
|
||||
in_channels,
|
||||
heads=in_channels // attention_head_dim,
|
||||
dim_head=attention_head_dim,
|
||||
rescale_output_factor=output_scale_factor,
|
||||
eps=resnet_eps,
|
||||
norm_num_groups=attn_groups,
|
||||
spatial_norm_dim=temb_channels if resnet_time_scale_shift == "spatial" else None,
|
||||
residual_connection=True,
|
||||
bias=True,
|
||||
upcast_softmax=True,
|
||||
_from_deprecated_attn_block=True,
|
||||
)
|
||||
)
|
||||
else:
|
||||
attentions.append(None)
|
||||
|
||||
resnets.append(
|
||||
ResnetBlockCausal3D(
|
||||
in_channels=in_channels,
|
||||
out_channels=in_channels,
|
||||
temb_channels=temb_channels,
|
||||
eps=resnet_eps,
|
||||
groups=resnet_groups,
|
||||
dropout=dropout,
|
||||
time_embedding_norm=resnet_time_scale_shift,
|
||||
non_linearity=resnet_act_fn,
|
||||
output_scale_factor=output_scale_factor,
|
||||
pre_norm=resnet_pre_norm,
|
||||
disable_causal=disable_causal,
|
||||
)
|
||||
)
|
||||
|
||||
self.attentions = nn.ModuleList(attentions)
|
||||
self.resnets = nn.ModuleList(resnets)
|
||||
|
||||
def forward(self, hidden_states: torch.FloatTensor, temb: Optional[torch.FloatTensor] = None) -> torch.FloatTensor:
|
||||
hidden_states = self.resnets[0](hidden_states, temb)
|
||||
for attn, resnet in zip(self.attentions, self.resnets[1:]):
|
||||
if attn is not None:
|
||||
B, C, T, H, W = hidden_states.shape
|
||||
hidden_states = rearrange(hidden_states, "b c f h w -> b (f h w) c")
|
||||
if self.causal_attention:
|
||||
attention_mask = prepare_causal_attention_mask(T, H * W, hidden_states.dtype, hidden_states.device, batch_size=B)
|
||||
else:
|
||||
attention_mask = None
|
||||
hidden_states = attn(hidden_states, temb=temb, attention_mask=attention_mask)
|
||||
hidden_states = rearrange(hidden_states, "b (f h w) c -> b c f h w", f=T, h=H, w=W)
|
||||
hidden_states = resnet(hidden_states, temb)
|
||||
|
||||
return hidden_states
|
||||
|
||||
|
||||
class DownEncoderBlockCausal3D(nn.Module):
|
||||
def __init__(
|
||||
self,
|
||||
in_channels: int,
|
||||
out_channels: int,
|
||||
dropout: float = 0.0,
|
||||
num_layers: int = 1,
|
||||
resnet_eps: float = 1e-6,
|
||||
resnet_time_scale_shift: str = "default",
|
||||
resnet_act_fn: str = "swish",
|
||||
resnet_groups: int = 32,
|
||||
resnet_pre_norm: bool = True,
|
||||
output_scale_factor: float = 1.0,
|
||||
add_downsample: bool = True,
|
||||
downsample_stride: int = 2,
|
||||
downsample_padding: int = 1,
|
||||
disable_causal: bool = False,
|
||||
):
|
||||
super().__init__()
|
||||
resnets = []
|
||||
|
||||
for i in range(num_layers):
|
||||
in_channels = in_channels if i == 0 else out_channels
|
||||
resnets.append(
|
||||
ResnetBlockCausal3D(
|
||||
in_channels=in_channels,
|
||||
out_channels=out_channels,
|
||||
temb_channels=None,
|
||||
eps=resnet_eps,
|
||||
groups=resnet_groups,
|
||||
dropout=dropout,
|
||||
time_embedding_norm=resnet_time_scale_shift,
|
||||
non_linearity=resnet_act_fn,
|
||||
output_scale_factor=output_scale_factor,
|
||||
pre_norm=resnet_pre_norm,
|
||||
disable_causal=disable_causal,
|
||||
)
|
||||
)
|
||||
|
||||
self.resnets = nn.ModuleList(resnets)
|
||||
|
||||
if add_downsample:
|
||||
self.downsamplers = nn.ModuleList(
|
||||
[
|
||||
DownsampleCausal3D(
|
||||
out_channels,
|
||||
use_conv=True,
|
||||
out_channels=out_channels,
|
||||
padding=downsample_padding,
|
||||
name="op",
|
||||
stride=downsample_stride,
|
||||
disable_causal=disable_causal,
|
||||
)
|
||||
]
|
||||
)
|
||||
else:
|
||||
self.downsamplers = None
|
||||
|
||||
def forward(self, hidden_states: torch.FloatTensor, scale: float = 1.0) -> torch.FloatTensor:
|
||||
for resnet in self.resnets:
|
||||
hidden_states = resnet(hidden_states, temb=None, scale=scale)
|
||||
|
||||
if self.downsamplers is not None:
|
||||
for downsampler in self.downsamplers:
|
||||
hidden_states = downsampler(hidden_states, scale)
|
||||
|
||||
return hidden_states
|
||||
|
||||
|
||||
class UpDecoderBlockCausal3D(nn.Module):
|
||||
def __init__(
|
||||
self,
|
||||
in_channels: int,
|
||||
out_channels: int,
|
||||
resolution_idx: Optional[int] = None,
|
||||
dropout: float = 0.0,
|
||||
num_layers: int = 1,
|
||||
resnet_eps: float = 1e-6,
|
||||
resnet_time_scale_shift: str = "default", # default, spatial
|
||||
resnet_act_fn: str = "swish",
|
||||
resnet_groups: int = 32,
|
||||
resnet_pre_norm: bool = True,
|
||||
output_scale_factor: float = 1.0,
|
||||
add_upsample: bool = True,
|
||||
upsample_scale_factor = (2, 2, 2),
|
||||
temb_channels: Optional[int] = None,
|
||||
disable_causal: bool = False,
|
||||
):
|
||||
super().__init__()
|
||||
resnets = []
|
||||
|
||||
for i in range(num_layers):
|
||||
input_channels = in_channels if i == 0 else out_channels
|
||||
|
||||
resnets.append(
|
||||
ResnetBlockCausal3D(
|
||||
in_channels=input_channels,
|
||||
out_channels=out_channels,
|
||||
temb_channels=temb_channels,
|
||||
eps=resnet_eps,
|
||||
groups=resnet_groups,
|
||||
dropout=dropout,
|
||||
time_embedding_norm=resnet_time_scale_shift,
|
||||
non_linearity=resnet_act_fn,
|
||||
output_scale_factor=output_scale_factor,
|
||||
pre_norm=resnet_pre_norm,
|
||||
disable_causal=disable_causal,
|
||||
)
|
||||
)
|
||||
|
||||
self.resnets = nn.ModuleList(resnets)
|
||||
|
||||
if add_upsample:
|
||||
self.upsamplers = nn.ModuleList(
|
||||
[
|
||||
UpsampleCausal3D(
|
||||
out_channels,
|
||||
use_conv=True,
|
||||
out_channels=out_channels,
|
||||
upsample_factor=upsample_scale_factor,
|
||||
disable_causal=disable_causal
|
||||
)
|
||||
]
|
||||
)
|
||||
else:
|
||||
self.upsamplers = None
|
||||
|
||||
self.resolution_idx = resolution_idx
|
||||
|
||||
def forward(
|
||||
self, hidden_states: torch.FloatTensor, temb: Optional[torch.FloatTensor] = None, scale: float = 1.0
|
||||
) -> torch.FloatTensor:
|
||||
for resnet in self.resnets:
|
||||
hidden_states = resnet(hidden_states, temb=temb, scale=scale)
|
||||
|
||||
if self.upsamplers is not None:
|
||||
for upsampler in self.upsamplers:
|
||||
hidden_states = upsampler(hidden_states)
|
||||
|
||||
return hidden_states
|
||||
|
||||
427
hyvideo/vae/vae.py
Normal file
427
hyvideo/vae/vae.py
Normal file
@@ -0,0 +1,427 @@
|
||||
from dataclasses import dataclass
|
||||
from typing import Optional, Tuple
|
||||
|
||||
import numpy as np
|
||||
import torch
|
||||
import torch.nn as nn
|
||||
|
||||
from diffusers.utils import BaseOutput, is_torch_version
|
||||
from diffusers.utils.torch_utils import randn_tensor
|
||||
from diffusers.models.attention_processor import SpatialNorm
|
||||
from .unet_causal_3d_blocks import (
|
||||
CausalConv3d,
|
||||
UNetMidBlockCausal3D,
|
||||
get_down_block3d,
|
||||
get_up_block3d,
|
||||
)
|
||||
|
||||
@dataclass
|
||||
class DecoderOutput(BaseOutput):
|
||||
r"""
|
||||
Output of decoding method.
|
||||
|
||||
Args:
|
||||
sample (`torch.FloatTensor` of shape `(batch_size, num_channels, height, width)`):
|
||||
The decoded output sample from the last layer of the model.
|
||||
"""
|
||||
|
||||
sample: torch.FloatTensor
|
||||
|
||||
|
||||
class EncoderCausal3D(nn.Module):
|
||||
r"""
|
||||
The `EncoderCausal3D` layer of a variational autoencoder that encodes its input into a latent representation.
|
||||
|
||||
Args:
|
||||
in_channels (`int`, *optional*, defaults to 3):
|
||||
The number of input channels.
|
||||
out_channels (`int`, *optional*, defaults to 3):
|
||||
The number of output channels.
|
||||
down_block_types (`Tuple[str, ...]`, *optional*, defaults to `("DownEncoderBlock2D",)`):
|
||||
The types of down blocks to use. See `~diffusers.models.unet_2d_blocks.get_down_block` for available
|
||||
options.
|
||||
block_out_channels (`Tuple[int, ...]`, *optional*, defaults to `(64,)`):
|
||||
The number of output channels for each block.
|
||||
layers_per_block (`int`, *optional*, defaults to 2):
|
||||
The number of layers per block.
|
||||
norm_num_groups (`int`, *optional*, defaults to 32):
|
||||
The number of groups for normalization.
|
||||
act_fn (`str`, *optional*, defaults to `"silu"`):
|
||||
The activation function to use. See `~diffusers.models.activations.get_activation` for available options.
|
||||
double_z (`bool`, *optional*, defaults to `True`):
|
||||
Whether to double the number of output channels for the last block.
|
||||
"""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
in_channels: int = 3,
|
||||
out_channels: int = 3,
|
||||
down_block_types: Tuple[str, ...] = ("DownEncoderBlockCausal3D",),
|
||||
block_out_channels: Tuple[int, ...] = (64,),
|
||||
layers_per_block: int = 2,
|
||||
norm_num_groups: int = 32,
|
||||
act_fn: str = "silu",
|
||||
double_z: bool = True,
|
||||
mid_block_add_attention=True,
|
||||
time_compression_ratio: int = 4,
|
||||
spatial_compression_ratio: int = 8,
|
||||
disable_causal: bool = False,
|
||||
mid_block_causal_attn: bool = False,
|
||||
):
|
||||
super().__init__()
|
||||
self.layers_per_block = layers_per_block
|
||||
|
||||
self.conv_in = CausalConv3d(in_channels, block_out_channels[0], kernel_size=3, stride=1, disable_causal=disable_causal)
|
||||
self.mid_block = None
|
||||
self.down_blocks = nn.ModuleList([])
|
||||
|
||||
# down
|
||||
output_channel = block_out_channels[0]
|
||||
for i, down_block_type in enumerate(down_block_types):
|
||||
input_channel = output_channel
|
||||
output_channel = block_out_channels[i]
|
||||
is_final_block = i == len(block_out_channels) - 1
|
||||
num_spatial_downsample_layers = int(np.log2(spatial_compression_ratio))
|
||||
num_time_downsample_layers = int(np.log2(time_compression_ratio))
|
||||
|
||||
if time_compression_ratio == 4:
|
||||
add_spatial_downsample = bool(i < num_spatial_downsample_layers)
|
||||
add_time_downsample = bool(i >= (len(block_out_channels) - 1 - num_time_downsample_layers) and not is_final_block)
|
||||
elif time_compression_ratio == 8:
|
||||
add_spatial_downsample = bool(i < num_spatial_downsample_layers)
|
||||
add_time_downsample = bool(i < num_time_downsample_layers)
|
||||
else:
|
||||
raise ValueError(f"Unsupported time_compression_ratio: {time_compression_ratio}")
|
||||
|
||||
downsample_stride_HW = (2, 2) if add_spatial_downsample else (1, 1)
|
||||
downsample_stride_T = (2, ) if add_time_downsample else (1, )
|
||||
downsample_stride = tuple(downsample_stride_T + downsample_stride_HW)
|
||||
down_block = get_down_block3d(
|
||||
down_block_type,
|
||||
num_layers=self.layers_per_block,
|
||||
in_channels=input_channel,
|
||||
out_channels=output_channel,
|
||||
add_downsample=bool(add_spatial_downsample or add_time_downsample),
|
||||
downsample_stride=downsample_stride,
|
||||
resnet_eps=1e-6,
|
||||
downsample_padding=0,
|
||||
resnet_act_fn=act_fn,
|
||||
resnet_groups=norm_num_groups,
|
||||
attention_head_dim=output_channel,
|
||||
temb_channels=None,
|
||||
disable_causal=disable_causal,
|
||||
)
|
||||
self.down_blocks.append(down_block)
|
||||
|
||||
# mid
|
||||
self.mid_block = UNetMidBlockCausal3D(
|
||||
in_channels=block_out_channels[-1],
|
||||
resnet_eps=1e-6,
|
||||
resnet_act_fn=act_fn,
|
||||
output_scale_factor=1,
|
||||
resnet_time_scale_shift="default",
|
||||
attention_head_dim=block_out_channels[-1],
|
||||
resnet_groups=norm_num_groups,
|
||||
temb_channels=None,
|
||||
add_attention=mid_block_add_attention,
|
||||
disable_causal=disable_causal,
|
||||
causal_attention=mid_block_causal_attn,
|
||||
)
|
||||
|
||||
# out
|
||||
self.conv_norm_out = nn.GroupNorm(num_channels=block_out_channels[-1], num_groups=norm_num_groups, eps=1e-6)
|
||||
self.conv_act = nn.SiLU()
|
||||
|
||||
conv_out_channels = 2 * out_channels if double_z else out_channels
|
||||
self.conv_out = CausalConv3d(block_out_channels[-1], conv_out_channels, kernel_size=3, disable_causal=disable_causal)
|
||||
|
||||
self.gradient_checkpointing = False
|
||||
|
||||
def forward(self, sample: torch.FloatTensor) -> torch.FloatTensor:
|
||||
r"""The forward method of the `EncoderCausal3D` class."""
|
||||
assert len(sample.shape) == 5, "The input tensor should have 5 dimensions"
|
||||
|
||||
sample = self.conv_in(sample)
|
||||
|
||||
if self.training and self.gradient_checkpointing:
|
||||
|
||||
def create_custom_forward(module):
|
||||
def custom_forward(*inputs):
|
||||
return module(*inputs)
|
||||
|
||||
return custom_forward
|
||||
|
||||
# down
|
||||
if is_torch_version(">=", "1.11.0"):
|
||||
for down_block in self.down_blocks:
|
||||
sample = torch.utils.checkpoint.checkpoint(
|
||||
create_custom_forward(down_block), sample, use_reentrant=False
|
||||
)
|
||||
# middle
|
||||
sample = torch.utils.checkpoint.checkpoint(
|
||||
create_custom_forward(self.mid_block), sample, use_reentrant=False
|
||||
)
|
||||
else:
|
||||
for down_block in self.down_blocks:
|
||||
sample = torch.utils.checkpoint.checkpoint(create_custom_forward(down_block), sample)
|
||||
# middle
|
||||
sample = torch.utils.checkpoint.checkpoint(create_custom_forward(self.mid_block), sample)
|
||||
|
||||
else:
|
||||
# down
|
||||
for down_block in self.down_blocks:
|
||||
sample = down_block(sample)
|
||||
|
||||
# middle
|
||||
sample = self.mid_block(sample)
|
||||
|
||||
# post-process
|
||||
sample = self.conv_norm_out(sample)
|
||||
sample = self.conv_act(sample)
|
||||
sample = self.conv_out(sample)
|
||||
|
||||
return sample
|
||||
|
||||
|
||||
class DecoderCausal3D(nn.Module):
|
||||
r"""
|
||||
The `DecoderCausal3D` layer of a variational autoencoder that decodes its latent representation into an output sample.
|
||||
|
||||
Args:
|
||||
in_channels (`int`, *optional*, defaults to 3):
|
||||
The number of input channels.
|
||||
out_channels (`int`, *optional*, defaults to 3):
|
||||
The number of output channels.
|
||||
up_block_types (`Tuple[str, ...]`, *optional*, defaults to `("UpDecoderBlock2D",)`):
|
||||
The types of up blocks to use. See `~diffusers.models.unet_2d_blocks.get_up_block` for available options.
|
||||
block_out_channels (`Tuple[int, ...]`, *optional*, defaults to `(64,)`):
|
||||
The number of output channels for each block.
|
||||
layers_per_block (`int`, *optional*, defaults to 2):
|
||||
The number of layers per block.
|
||||
norm_num_groups (`int`, *optional*, defaults to 32):
|
||||
The number of groups for normalization.
|
||||
act_fn (`str`, *optional*, defaults to `"silu"`):
|
||||
The activation function to use. See `~diffusers.models.activations.get_activation` for available options.
|
||||
norm_type (`str`, *optional*, defaults to `"group"`):
|
||||
The normalization type to use. Can be either `"group"` or `"spatial"`.
|
||||
"""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
in_channels: int = 3,
|
||||
out_channels: int = 3,
|
||||
up_block_types: Tuple[str, ...] = ("UpDecoderBlockCausal3D",),
|
||||
block_out_channels: Tuple[int, ...] = (64,),
|
||||
layers_per_block: int = 2,
|
||||
norm_num_groups: int = 32,
|
||||
act_fn: str = "silu",
|
||||
norm_type: str = "group", # group, spatial
|
||||
mid_block_add_attention=True,
|
||||
time_compression_ratio: int = 4,
|
||||
spatial_compression_ratio: int = 8,
|
||||
disable_causal: bool = False,
|
||||
mid_block_causal_attn: bool = False,
|
||||
):
|
||||
super().__init__()
|
||||
self.layers_per_block = layers_per_block
|
||||
|
||||
self.conv_in = CausalConv3d(in_channels, block_out_channels[-1], kernel_size=3, stride=1, disable_causal=disable_causal)
|
||||
self.mid_block = None
|
||||
self.up_blocks = nn.ModuleList([])
|
||||
|
||||
temb_channels = in_channels if norm_type == "spatial" else None
|
||||
|
||||
# mid
|
||||
self.mid_block = UNetMidBlockCausal3D(
|
||||
in_channels=block_out_channels[-1],
|
||||
resnet_eps=1e-6,
|
||||
resnet_act_fn=act_fn,
|
||||
output_scale_factor=1,
|
||||
resnet_time_scale_shift="default" if norm_type == "group" else norm_type,
|
||||
attention_head_dim=block_out_channels[-1],
|
||||
resnet_groups=norm_num_groups,
|
||||
temb_channels=temb_channels,
|
||||
add_attention=mid_block_add_attention,
|
||||
disable_causal=disable_causal,
|
||||
causal_attention=mid_block_causal_attn,
|
||||
)
|
||||
|
||||
# up
|
||||
reversed_block_out_channels = list(reversed(block_out_channels))
|
||||
output_channel = reversed_block_out_channels[0]
|
||||
for i, up_block_type in enumerate(up_block_types):
|
||||
prev_output_channel = output_channel
|
||||
output_channel = reversed_block_out_channels[i]
|
||||
is_final_block = i == len(block_out_channels) - 1
|
||||
num_spatial_upsample_layers = int(np.log2(spatial_compression_ratio))
|
||||
num_time_upsample_layers = int(np.log2(time_compression_ratio))
|
||||
|
||||
if time_compression_ratio == 4:
|
||||
add_spatial_upsample = bool(i < num_spatial_upsample_layers)
|
||||
add_time_upsample = bool(i >= len(block_out_channels) - 1 - num_time_upsample_layers and not is_final_block)
|
||||
elif time_compression_ratio == 8:
|
||||
add_spatial_upsample = bool(i >= len(block_out_channels) - num_spatial_upsample_layers)
|
||||
add_time_upsample = bool(i >= len(block_out_channels) - num_time_upsample_layers)
|
||||
else:
|
||||
raise ValueError(f"Unsupported time_compression_ratio: {time_compression_ratio}")
|
||||
|
||||
upsample_scale_factor_HW = (2, 2) if add_spatial_upsample else (1, 1)
|
||||
upsample_scale_factor_T = (2, ) if add_time_upsample else (1, )
|
||||
upsample_scale_factor = tuple(upsample_scale_factor_T + upsample_scale_factor_HW)
|
||||
up_block = get_up_block3d(
|
||||
up_block_type,
|
||||
num_layers=self.layers_per_block + 1,
|
||||
in_channels=prev_output_channel,
|
||||
out_channels=output_channel,
|
||||
prev_output_channel=None,
|
||||
add_upsample=bool(add_spatial_upsample or add_time_upsample),
|
||||
upsample_scale_factor=upsample_scale_factor,
|
||||
resnet_eps=1e-6,
|
||||
resnet_act_fn=act_fn,
|
||||
resnet_groups=norm_num_groups,
|
||||
attention_head_dim=output_channel,
|
||||
temb_channels=temb_channels,
|
||||
resnet_time_scale_shift=norm_type,
|
||||
disable_causal=disable_causal,
|
||||
)
|
||||
self.up_blocks.append(up_block)
|
||||
prev_output_channel = output_channel
|
||||
|
||||
# out
|
||||
if norm_type == "spatial":
|
||||
self.conv_norm_out = SpatialNorm(block_out_channels[0], temb_channels)
|
||||
else:
|
||||
self.conv_norm_out = nn.GroupNorm(num_channels=block_out_channels[0], num_groups=norm_num_groups, eps=1e-6)
|
||||
self.conv_act = nn.SiLU()
|
||||
self.conv_out = CausalConv3d(block_out_channels[0], out_channels, kernel_size=3, disable_causal=disable_causal)
|
||||
|
||||
self.gradient_checkpointing = False
|
||||
|
||||
def forward(
|
||||
self,
|
||||
sample: torch.FloatTensor,
|
||||
latent_embeds: Optional[torch.FloatTensor] = None,
|
||||
) -> torch.FloatTensor:
|
||||
r"""The forward method of the `DecoderCausal3D` class."""
|
||||
assert len(sample.shape) == 5, "The input tensor should have 5 dimensions"
|
||||
|
||||
sample = self.conv_in(sample)
|
||||
|
||||
upscale_dtype = next(iter(self.up_blocks.parameters())).dtype
|
||||
if self.training and self.gradient_checkpointing:
|
||||
|
||||
def create_custom_forward(module):
|
||||
def custom_forward(*inputs):
|
||||
return module(*inputs)
|
||||
|
||||
return custom_forward
|
||||
|
||||
if is_torch_version(">=", "1.11.0"):
|
||||
# middle
|
||||
sample = torch.utils.checkpoint.checkpoint(
|
||||
create_custom_forward(self.mid_block),
|
||||
sample,
|
||||
latent_embeds,
|
||||
use_reentrant=False,
|
||||
)
|
||||
sample = sample.to(upscale_dtype)
|
||||
|
||||
# up
|
||||
for up_block in self.up_blocks:
|
||||
sample = torch.utils.checkpoint.checkpoint(
|
||||
create_custom_forward(up_block),
|
||||
sample,
|
||||
latent_embeds,
|
||||
use_reentrant=False,
|
||||
)
|
||||
else:
|
||||
# middle
|
||||
sample = torch.utils.checkpoint.checkpoint(
|
||||
create_custom_forward(self.mid_block), sample, latent_embeds
|
||||
)
|
||||
sample = sample.to(upscale_dtype)
|
||||
|
||||
# up
|
||||
for up_block in self.up_blocks:
|
||||
sample = torch.utils.checkpoint.checkpoint(create_custom_forward(up_block), sample, latent_embeds)
|
||||
else:
|
||||
# middle
|
||||
sample = self.mid_block(sample, latent_embeds)
|
||||
sample = sample.to(upscale_dtype)
|
||||
|
||||
# up
|
||||
for up_block in self.up_blocks:
|
||||
sample = up_block(sample, latent_embeds)
|
||||
|
||||
# post-process
|
||||
if latent_embeds is None:
|
||||
sample = self.conv_norm_out(sample)
|
||||
else:
|
||||
sample = self.conv_norm_out(sample, latent_embeds)
|
||||
sample = self.conv_act(sample)
|
||||
sample = self.conv_out(sample)
|
||||
|
||||
return sample
|
||||
|
||||
|
||||
class DiagonalGaussianDistribution(object):
|
||||
def __init__(self, parameters: torch.Tensor, deterministic: bool = False):
|
||||
if parameters.ndim == 3:
|
||||
dim = 2 # (B, L, C)
|
||||
elif parameters.ndim == 5 or parameters.ndim == 4:
|
||||
dim = 1 # (B, C, T, H ,W) / (B, C, H, W)
|
||||
else:
|
||||
raise NotImplementedError
|
||||
self.parameters = parameters
|
||||
self.mean, self.logvar = torch.chunk(parameters, 2, dim=dim)
|
||||
self.logvar = torch.clamp(self.logvar, -30.0, 20.0)
|
||||
self.deterministic = deterministic
|
||||
self.std = torch.exp(0.5 * self.logvar)
|
||||
self.var = torch.exp(self.logvar)
|
||||
if self.deterministic:
|
||||
self.var = self.std = torch.zeros_like(
|
||||
self.mean, device=self.parameters.device, dtype=self.parameters.dtype
|
||||
)
|
||||
|
||||
def sample(self, generator: Optional[torch.Generator] = None) -> torch.FloatTensor:
|
||||
# make sure sample is on the same device as the parameters and has same dtype
|
||||
sample = randn_tensor(
|
||||
self.mean.shape,
|
||||
generator=generator,
|
||||
device=self.parameters.device,
|
||||
dtype=self.parameters.dtype,
|
||||
)
|
||||
x = self.mean + self.std * sample
|
||||
return x
|
||||
|
||||
def kl(self, other: "DiagonalGaussianDistribution" = None) -> torch.Tensor:
|
||||
if self.deterministic:
|
||||
return torch.Tensor([0.0])
|
||||
else:
|
||||
reduce_dim = list(range(1, self.mean.ndim))
|
||||
if other is None:
|
||||
return 0.5 * torch.sum(
|
||||
torch.pow(self.mean, 2) + self.var - 1.0 - self.logvar,
|
||||
dim=reduce_dim,
|
||||
)
|
||||
else:
|
||||
return 0.5 * torch.sum(
|
||||
torch.pow(self.mean - other.mean, 2) / other.var
|
||||
+ self.var / other.var
|
||||
- 1.0
|
||||
- self.logvar
|
||||
+ other.logvar,
|
||||
dim=reduce_dim,
|
||||
)
|
||||
|
||||
def nll(self, sample: torch.Tensor, dims: Tuple[int, ...] = [1, 2, 3]) -> torch.Tensor:
|
||||
if self.deterministic:
|
||||
return torch.Tensor([0.0])
|
||||
logtwopi = np.log(2.0 * np.pi)
|
||||
return 0.5 * torch.sum(
|
||||
logtwopi + self.logvar + torch.pow(sample - self.mean, 2) / self.var,
|
||||
dim=dims,
|
||||
)
|
||||
|
||||
def mode(self) -> torch.Tensor:
|
||||
return self.mean
|
||||
1
loras_hunyuan/Readme.txt
Normal file
1
loras_hunyuan/Readme.txt
Normal file
@@ -0,0 +1 @@
|
||||
loras for hunyuan t2v
|
||||
1
loras_hunyuan_i2v/Readme.txt
Normal file
1
loras_hunyuan_i2v/Readme.txt
Normal file
@@ -0,0 +1 @@
|
||||
loras for hunyuan i2v
|
||||
0
ltx_video/__init__.py
Normal file
0
ltx_video/__init__.py
Normal file
41
ltx_video/configs/ltxv-13b-0.9.7-dev.original.yaml
Normal file
41
ltx_video/configs/ltxv-13b-0.9.7-dev.original.yaml
Normal file
@@ -0,0 +1,41 @@
|
||||
|
||||
pipeline_type: multi-scale
|
||||
checkpoint_path: "ltxv-13b-0.9.7-dev.safetensors"
|
||||
downscale_factor: 0.6666666
|
||||
spatial_upscaler_model_path: "ltxv-spatial-upscaler-0.9.7.safetensors"
|
||||
stg_mode: "attention_values" # options: "attention_values", "attention_skip", "residual", "transformer_block"
|
||||
decode_timestep: 0.05
|
||||
decode_noise_scale: 0.025
|
||||
text_encoder_model_name_or_path: "PixArt-alpha/PixArt-XL-2-1024-MS"
|
||||
sampler: "from_checkpoint" # options: "uniform", "linear-quadratic", "from_checkpoint"
|
||||
prompt_enhancement_words_threshold: 120
|
||||
prompt_enhancer_image_caption_model_name_or_path: "MiaoshouAI/Florence-2-large-PromptGen-v2.0"
|
||||
prompt_enhancer_llm_model_name_or_path: "unsloth/Llama-3.2-3B-Instruct"
|
||||
stochastic_sampling: false
|
||||
|
||||
|
||||
first_pass:
|
||||
#13b Dynamic
|
||||
guidance_scale: [1, 6, 8, 6, 1, 1]
|
||||
stg_scale: [0, 4, 4, 4, 2, 1]
|
||||
rescaling_scale: [1, 0.5, 0.5, 1, 1, 1]
|
||||
guidance_timesteps: [1.0, 0.9933, 0.9850, 0.9767, 0.9008, 0.6180]
|
||||
skip_block_list: [[11, 25, 35, 39], [22, 35, 39], [28], [28], [28], [28]]
|
||||
num_inference_steps: 30 #default
|
||||
|
||||
|
||||
second_pass:
|
||||
#13b Dynamic
|
||||
guidance_scale: [1, 6, 8, 6, 1, 1]
|
||||
stg_scale: [0, 4, 4, 4, 2, 1]
|
||||
rescaling_scale: [1, 0.5, 0.5, 1, 1, 1]
|
||||
guidance_timesteps: [1.0, 0.9933, 0.9850, 0.9767, 0.9008, 0.6180]
|
||||
skip_block_list: [[11, 25, 35, 39], [22, 35, 39], [28], [28], [28], [28]]
|
||||
#13b Upscale
|
||||
# guidance_scale: [1, 1, 1, 1, 1, 1]
|
||||
# stg_scale: [1, 1, 1, 1, 1, 1]
|
||||
# rescaling_scale: [1, 1, 1, 1, 1, 1]
|
||||
# guidance_timesteps: [1.0, 0.9933, 0.9850, 0.9767, 0.9008, 0.6180]
|
||||
# skip_block_list: [[42], [42], [42], [42], [42], [42]]
|
||||
num_inference_steps: 30 #default
|
||||
strength: 0.85
|
||||
34
ltx_video/configs/ltxv-13b-0.9.7-dev.yaml
Normal file
34
ltx_video/configs/ltxv-13b-0.9.7-dev.yaml
Normal file
@@ -0,0 +1,34 @@
|
||||
pipeline_type: multi-scale
|
||||
checkpoint_path: "ltxv-13b-0.9.7-dev.safetensors"
|
||||
downscale_factor: 0.6666666
|
||||
spatial_upscaler_model_path: "ltxv-spatial-upscaler-0.9.7.safetensors"
|
||||
stg_mode: "attention_values" # options: "attention_values", "attention_skip", "residual", "transformer_block"
|
||||
decode_timestep: 0.05
|
||||
decode_noise_scale: 0.025
|
||||
text_encoder_model_name_or_path: "PixArt-alpha/PixArt-XL-2-1024-MS"
|
||||
precision: "bfloat16"
|
||||
sampler: "from_checkpoint" # options: "uniform", "linear-quadratic", "from_checkpoint"
|
||||
prompt_enhancement_words_threshold: 120
|
||||
prompt_enhancer_image_caption_model_name_or_path: "MiaoshouAI/Florence-2-large-PromptGen-v2.0"
|
||||
prompt_enhancer_llm_model_name_or_path: "unsloth/Llama-3.2-3B-Instruct"
|
||||
stochastic_sampling: false
|
||||
|
||||
first_pass:
|
||||
guidance_scale: [1, 1, 6, 8, 6, 1, 1]
|
||||
stg_scale: [0, 0, 4, 4, 4, 2, 1]
|
||||
rescaling_scale: [1, 1, 0.5, 0.5, 1, 1, 1]
|
||||
guidance_timesteps: [1.0, 0.996, 0.9933, 0.9850, 0.9767, 0.9008, 0.6180]
|
||||
skip_block_list: [[], [11, 25, 35, 39], [22, 35, 39], [28], [28], [28], [28]]
|
||||
num_inference_steps: 30
|
||||
skip_final_inference_steps: 3
|
||||
cfg_star_rescale: true
|
||||
|
||||
second_pass:
|
||||
guidance_scale: [1]
|
||||
stg_scale: [1]
|
||||
rescaling_scale: [1]
|
||||
guidance_timesteps: [1.0]
|
||||
skip_block_list: [27]
|
||||
num_inference_steps: 30
|
||||
skip_initial_inference_steps: 17
|
||||
cfg_star_rescale: true
|
||||
28
ltx_video/configs/ltxv-13b-0.9.7-distilled.yaml
Normal file
28
ltx_video/configs/ltxv-13b-0.9.7-distilled.yaml
Normal file
@@ -0,0 +1,28 @@
|
||||
pipeline_type: multi-scale
|
||||
checkpoint_path: "ltxv-13b-0.9.7-distilled.safetensors"
|
||||
downscale_factor: 0.6666666
|
||||
spatial_upscaler_model_path: "ltxv-spatial-upscaler-0.9.7.safetensors"
|
||||
stg_mode: "attention_values" # options: "attention_values", "attention_skip", "residual", "transformer_block"
|
||||
decode_timestep: 0.05
|
||||
decode_noise_scale: 0.025
|
||||
text_encoder_model_name_or_path: "PixArt-alpha/PixArt-XL-2-1024-MS"
|
||||
precision: "bfloat16"
|
||||
sampler: "from_checkpoint" # options: "uniform", "linear-quadratic", "from_checkpoint"
|
||||
prompt_enhancement_words_threshold: 120
|
||||
prompt_enhancer_image_caption_model_name_or_path: "MiaoshouAI/Florence-2-large-PromptGen-v2.0"
|
||||
prompt_enhancer_llm_model_name_or_path: "unsloth/Llama-3.2-3B-Instruct"
|
||||
stochastic_sampling: false
|
||||
|
||||
first_pass:
|
||||
timesteps: [1.0000, 0.9937, 0.9875, 0.9812, 0.9750, 0.9094, 0.7250]
|
||||
guidance_scale: 1
|
||||
stg_scale: 0
|
||||
rescaling_scale: 1
|
||||
skip_block_list: [42]
|
||||
|
||||
second_pass:
|
||||
timesteps: [0.9094, 0.7250, 0.4219]
|
||||
guidance_scale: 1
|
||||
stg_scale: 0
|
||||
rescaling_scale: 1
|
||||
skip_block_list: [42]
|
||||
17
ltx_video/configs/ltxv-2b-0.9.6-dev.yaml
Normal file
17
ltx_video/configs/ltxv-2b-0.9.6-dev.yaml
Normal file
@@ -0,0 +1,17 @@
|
||||
pipeline_type: base
|
||||
checkpoint_path: "ltxv-2b-0.9.6-dev-04-25.safetensors"
|
||||
guidance_scale: 3
|
||||
stg_scale: 1
|
||||
rescaling_scale: 0.7
|
||||
skip_block_list: [19]
|
||||
num_inference_steps: 40
|
||||
stg_mode: "attention_values" # options: "attention_values", "attention_skip", "residual", "transformer_block"
|
||||
decode_timestep: 0.05
|
||||
decode_noise_scale: 0.025
|
||||
text_encoder_model_name_or_path: "PixArt-alpha/PixArt-XL-2-1024-MS"
|
||||
precision: "bfloat16"
|
||||
sampler: "from_checkpoint" # options: "uniform", "linear-quadratic", "from_checkpoint"
|
||||
prompt_enhancement_words_threshold: 120
|
||||
prompt_enhancer_image_caption_model_name_or_path: "MiaoshouAI/Florence-2-large-PromptGen-v2.0"
|
||||
prompt_enhancer_llm_model_name_or_path: "unsloth/Llama-3.2-3B-Instruct"
|
||||
stochastic_sampling: false
|
||||
17
ltx_video/configs/ltxv-2b-0.9.6-distilled.yaml
Normal file
17
ltx_video/configs/ltxv-2b-0.9.6-distilled.yaml
Normal file
@@ -0,0 +1,17 @@
|
||||
pipeline_type: base
|
||||
checkpoint_path: "ltxv-2b-0.9.6-distilled-04-25.safetensors"
|
||||
guidance_scale: 3
|
||||
stg_scale: 1
|
||||
rescaling_scale: 0.7
|
||||
skip_block_list: [19]
|
||||
num_inference_steps: 8
|
||||
stg_mode: "attention_values" # options: "attention_values", "attention_skip", "residual", "transformer_block"
|
||||
decode_timestep: 0.05
|
||||
decode_noise_scale: 0.025
|
||||
text_encoder_model_name_or_path: "PixArt-alpha/PixArt-XL-2-1024-MS"
|
||||
precision: "bfloat16"
|
||||
sampler: "from_checkpoint" # options: "uniform", "linear-quadratic", "from_checkpoint"
|
||||
prompt_enhancement_words_threshold: 120
|
||||
prompt_enhancer_image_caption_model_name_or_path: "MiaoshouAI/Florence-2-large-PromptGen-v2.0"
|
||||
prompt_enhancer_llm_model_name_or_path: "unsloth/Llama-3.2-3B-Instruct"
|
||||
stochastic_sampling: true
|
||||
562
ltx_video/ltxv.py
Normal file
562
ltx_video/ltxv.py
Normal file
@@ -0,0 +1,562 @@
|
||||
from mmgp import offload
|
||||
import argparse
|
||||
import os
|
||||
import random
|
||||
from datetime import datetime
|
||||
from pathlib import Path
|
||||
from diffusers.utils import logging
|
||||
from typing import Optional, List, Union
|
||||
import yaml
|
||||
from wan.utils.utils import calculate_new_dimensions
|
||||
import imageio
|
||||
import json
|
||||
import numpy as np
|
||||
import torch
|
||||
from safetensors import safe_open
|
||||
from PIL import Image
|
||||
from transformers import (
|
||||
T5EncoderModel,
|
||||
T5Tokenizer,
|
||||
AutoModelForCausalLM,
|
||||
AutoProcessor,
|
||||
AutoTokenizer,
|
||||
)
|
||||
from huggingface_hub import hf_hub_download
|
||||
|
||||
from .models.autoencoders.causal_video_autoencoder import (
|
||||
CausalVideoAutoencoder,
|
||||
)
|
||||
from .models.transformers.symmetric_patchifier import SymmetricPatchifier
|
||||
from .models.transformers.transformer3d import Transformer3DModel
|
||||
from .pipelines.pipeline_ltx_video import (
|
||||
ConditioningItem,
|
||||
LTXVideoPipeline,
|
||||
LTXMultiScalePipeline,
|
||||
)
|
||||
from .schedulers.rf import RectifiedFlowScheduler
|
||||
from .utils.skip_layer_strategy import SkipLayerStrategy
|
||||
from .models.autoencoders.latent_upsampler import LatentUpsampler
|
||||
from .pipelines import crf_compressor
|
||||
import cv2
|
||||
|
||||
MAX_HEIGHT = 720
|
||||
MAX_WIDTH = 1280
|
||||
MAX_NUM_FRAMES = 257
|
||||
|
||||
logger = logging.get_logger("LTX-Video")
|
||||
|
||||
|
||||
def get_total_gpu_memory():
|
||||
if torch.cuda.is_available():
|
||||
total_memory = torch.cuda.get_device_properties(0).total_memory / (1024**3)
|
||||
return total_memory
|
||||
return 0
|
||||
|
||||
|
||||
def get_device():
|
||||
if torch.cuda.is_available():
|
||||
return "cuda"
|
||||
elif torch.backends.mps.is_available():
|
||||
return "mps"
|
||||
return "cpu"
|
||||
|
||||
|
||||
def load_image_to_tensor_with_resize_and_crop(
|
||||
image_input: Union[str, Image.Image],
|
||||
target_height: int = 512,
|
||||
target_width: int = 768,
|
||||
just_crop: bool = False,
|
||||
) -> torch.Tensor:
|
||||
"""Load and process an image into a tensor.
|
||||
|
||||
Args:
|
||||
image_input: Either a file path (str) or a PIL Image object
|
||||
target_height: Desired height of output tensor
|
||||
target_width: Desired width of output tensor
|
||||
just_crop: If True, only crop the image to the target size without resizing
|
||||
"""
|
||||
if isinstance(image_input, str):
|
||||
image = Image.open(image_input).convert("RGB")
|
||||
elif isinstance(image_input, Image.Image):
|
||||
image = image_input
|
||||
else:
|
||||
raise ValueError("image_input must be either a file path or a PIL Image object")
|
||||
|
||||
input_width, input_height = image.size
|
||||
aspect_ratio_target = target_width / target_height
|
||||
aspect_ratio_frame = input_width / input_height
|
||||
if aspect_ratio_frame > aspect_ratio_target:
|
||||
new_width = int(input_height * aspect_ratio_target)
|
||||
new_height = input_height
|
||||
x_start = (input_width - new_width) // 2
|
||||
y_start = 0
|
||||
else:
|
||||
new_width = input_width
|
||||
new_height = int(input_width / aspect_ratio_target)
|
||||
x_start = 0
|
||||
y_start = (input_height - new_height) // 2
|
||||
|
||||
image = image.crop((x_start, y_start, x_start + new_width, y_start + new_height))
|
||||
if not just_crop:
|
||||
image = image.resize((target_width, target_height))
|
||||
|
||||
image = np.array(image)
|
||||
image = cv2.GaussianBlur(image, (3, 3), 0)
|
||||
frame_tensor = torch.from_numpy(image).float()
|
||||
frame_tensor = crf_compressor.compress(frame_tensor / 255.0) * 255.0
|
||||
frame_tensor = frame_tensor.permute(2, 0, 1)
|
||||
frame_tensor = (frame_tensor / 127.5) - 1.0
|
||||
# Create 5D tensor: (batch_size=1, channels=3, num_frames=1, height, width)
|
||||
return frame_tensor.unsqueeze(0).unsqueeze(2)
|
||||
|
||||
|
||||
|
||||
def calculate_padding(
|
||||
source_height: int, source_width: int, target_height: int, target_width: int
|
||||
) -> tuple[int, int, int, int]:
|
||||
|
||||
# Calculate total padding needed
|
||||
pad_height = target_height - source_height
|
||||
pad_width = target_width - source_width
|
||||
|
||||
# Calculate padding for each side
|
||||
pad_top = pad_height // 2
|
||||
pad_bottom = pad_height - pad_top # Handles odd padding
|
||||
pad_left = pad_width // 2
|
||||
pad_right = pad_width - pad_left # Handles odd padding
|
||||
|
||||
# Return padded tensor
|
||||
# Padding format is (left, right, top, bottom)
|
||||
padding = (pad_left, pad_right, pad_top, pad_bottom)
|
||||
return padding
|
||||
|
||||
|
||||
|
||||
|
||||
def seed_everething(seed: int):
|
||||
random.seed(seed)
|
||||
np.random.seed(seed)
|
||||
torch.manual_seed(seed)
|
||||
if torch.cuda.is_available():
|
||||
torch.cuda.manual_seed(seed)
|
||||
if torch.backends.mps.is_available():
|
||||
torch.mps.manual_seed(seed)
|
||||
|
||||
|
||||
class LTXV:
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
model_filepath: str,
|
||||
text_encoder_filepath: str,
|
||||
dtype = torch.bfloat16,
|
||||
VAE_dtype = torch.bfloat16,
|
||||
mixed_precision_transformer = False
|
||||
):
|
||||
|
||||
self.mixed_precision_transformer = mixed_precision_transformer
|
||||
# ckpt_path = Path(ckpt_path)
|
||||
# with safe_open(ckpt_path, framework="pt") as f:
|
||||
# metadata = f.metadata()
|
||||
# config_str = metadata.get("config")
|
||||
# configs = json.loads(config_str)
|
||||
# allowed_inference_steps = configs.get("allowed_inference_steps", None)
|
||||
# transformer = Transformer3DModel.from_pretrained(ckpt_path)
|
||||
# offload.save_model(transformer, "ckpts/ltxv_0.9.7_13B_dev_bf16.safetensors", config_file_path="config_transformer.json")
|
||||
|
||||
# vae = CausalVideoAutoencoder.from_pretrained(ckpt_path)
|
||||
vae = offload.fast_load_transformers_model("ckpts/ltxv_0.9.7_VAE.safetensors", modelClass=CausalVideoAutoencoder)
|
||||
if VAE_dtype == torch.float16:
|
||||
VAE_dtype = torch.bfloat16
|
||||
|
||||
vae = vae.to(VAE_dtype)
|
||||
vae._model_dtype = VAE_dtype
|
||||
# vae = offload.fast_load_transformers_model("vae.safetensors", modelClass=CausalVideoAutoencoder, modelPrefix= "vae", forcedConfigPath="config_vae.json")
|
||||
# offload.save_model(vae, "vae.safetensors", config_file_path="config_vae.json")
|
||||
|
||||
|
||||
transformer = offload.fast_load_transformers_model(model_filepath, modelClass=Transformer3DModel)
|
||||
transformer._model_dtype = dtype
|
||||
if mixed_precision_transformer:
|
||||
transformer._lock_dtype = torch.float
|
||||
|
||||
|
||||
scheduler = RectifiedFlowScheduler.from_pretrained("ckpts/ltxv_scheduler.json")
|
||||
# transformer = offload.fast_load_transformers_model("ltx_13B_quanto_bf16_int8.safetensors", modelClass=Transformer3DModel, modelPrefix= "model.diffusion_model", forcedConfigPath="config_transformer.json")
|
||||
# offload.save_model(transformer, "ltx_13B_quanto_bf16_int8.safetensors", do_quantize= True, config_file_path="config_transformer.json")
|
||||
|
||||
latent_upsampler = LatentUpsampler.from_pretrained("ckpts/ltxv_0.9.7_spatial_upscaler.safetensors").to("cpu").eval()
|
||||
latent_upsampler.to(VAE_dtype)
|
||||
latent_upsampler._model_dtype = VAE_dtype
|
||||
|
||||
allowed_inference_steps = None
|
||||
|
||||
# text_encoder = T5EncoderModel.from_pretrained(
|
||||
# "PixArt-alpha/PixArt-XL-2-1024-MS", subfolder="text_encoder"
|
||||
# )
|
||||
# text_encoder.to(torch.bfloat16)
|
||||
# offload.save_model(text_encoder, "T5_xxl_1.1_enc_bf16.safetensors", config_file_path="T5_config.json")
|
||||
# offload.save_model(text_encoder, "T5_xxl_1.1_enc_quanto_bf16_int8.safetensors", do_quantize= True, config_file_path="T5_config.json")
|
||||
|
||||
text_encoder = offload.fast_load_transformers_model(text_encoder_filepath)
|
||||
patchifier = SymmetricPatchifier(patch_size=1)
|
||||
tokenizer = T5Tokenizer.from_pretrained( "ckpts/T5_xxl_1.1")
|
||||
|
||||
enhance_prompt = False
|
||||
if enhance_prompt:
|
||||
prompt_enhancer_image_caption_model = AutoModelForCausalLM.from_pretrained( "ckpts/Florence2", trust_remote_code=True)
|
||||
prompt_enhancer_image_caption_processor = AutoProcessor.from_pretrained( "ckpts/Florence2", trust_remote_code=True)
|
||||
prompt_enhancer_llm_model = offload.fast_load_transformers_model("ckpts/Llama3_2_quanto_bf16_int8.safetensors")
|
||||
prompt_enhancer_llm_tokenizer = AutoTokenizer.from_pretrained("ckpts/Llama3_2")
|
||||
else:
|
||||
prompt_enhancer_image_caption_model = None
|
||||
prompt_enhancer_image_caption_processor = None
|
||||
prompt_enhancer_llm_model = None
|
||||
prompt_enhancer_llm_tokenizer = None
|
||||
|
||||
if prompt_enhancer_image_caption_model != None:
|
||||
pipe["prompt_enhancer_image_caption_model"] = prompt_enhancer_image_caption_model
|
||||
prompt_enhancer_image_caption_model._model_dtype = torch.float
|
||||
|
||||
pipe["prompt_enhancer_llm_model"] = prompt_enhancer_llm_model
|
||||
|
||||
# offload.profile(pipe, profile_no=5, extraModelsToQuantize = None, quantizeTransformer = False, budgets = { "prompt_enhancer_llm_model" : 10000, "prompt_enhancer_image_caption_model" : 10000, "vae" : 3000, "*" : 100 }, verboseLevel=2)
|
||||
|
||||
|
||||
# Use submodels for the pipeline
|
||||
submodel_dict = {
|
||||
"transformer": transformer,
|
||||
"patchifier": patchifier,
|
||||
"text_encoder": text_encoder,
|
||||
"tokenizer": tokenizer,
|
||||
"scheduler": scheduler,
|
||||
"vae": vae,
|
||||
"prompt_enhancer_image_caption_model": prompt_enhancer_image_caption_model,
|
||||
"prompt_enhancer_image_caption_processor": prompt_enhancer_image_caption_processor,
|
||||
"prompt_enhancer_llm_model": prompt_enhancer_llm_model,
|
||||
"prompt_enhancer_llm_tokenizer": prompt_enhancer_llm_tokenizer,
|
||||
"allowed_inference_steps": allowed_inference_steps,
|
||||
}
|
||||
pipeline = LTXVideoPipeline(**submodel_dict)
|
||||
pipeline = LTXMultiScalePipeline(pipeline, latent_upsampler=latent_upsampler)
|
||||
|
||||
self.pipeline = pipeline
|
||||
self.model = transformer
|
||||
self.vae = vae
|
||||
# return pipeline, pipe
|
||||
|
||||
def generate(
|
||||
self,
|
||||
input_prompt: str,
|
||||
n_prompt: str,
|
||||
image_start = None,
|
||||
image_end = None,
|
||||
input_video = None,
|
||||
sampling_steps = 50,
|
||||
image_cond_noise_scale: float = 0.15,
|
||||
input_media_path: Optional[str] = None,
|
||||
strength: Optional[float] = 1.0,
|
||||
seed: int = 42,
|
||||
height: Optional[int] = 704,
|
||||
width: Optional[int] = 1216,
|
||||
frame_num: int = 81,
|
||||
frame_rate: int = 30,
|
||||
fit_into_canvas = True,
|
||||
callback=None,
|
||||
device: Optional[str] = None,
|
||||
VAE_tile_size = None,
|
||||
**kwargs,
|
||||
):
|
||||
|
||||
num_inference_steps1 = sampling_steps
|
||||
num_inference_steps2 = sampling_steps #10
|
||||
conditioning_strengths = None
|
||||
conditioning_media_paths = []
|
||||
conditioning_start_frames = []
|
||||
|
||||
|
||||
if input_video != None:
|
||||
conditioning_media_paths.append(input_video)
|
||||
conditioning_start_frames.append(0)
|
||||
height, width = input_video.shape[-2:]
|
||||
else:
|
||||
if image_start != None:
|
||||
image_start = image_start[0]
|
||||
frame_width, frame_height = image_start.size
|
||||
height, width = calculate_new_dimensions(height, width, frame_height, frame_width, fit_into_canvas, 32)
|
||||
conditioning_media_paths.append(image_start)
|
||||
conditioning_start_frames.append(0)
|
||||
if image_end != None:
|
||||
image_end = image_end[0]
|
||||
conditioning_media_paths.append(image_end)
|
||||
conditioning_start_frames.append(frame_num-1)
|
||||
|
||||
if len(conditioning_media_paths) == 0:
|
||||
conditioning_media_paths = None
|
||||
conditioning_start_frames = None
|
||||
|
||||
pipeline_config = "ltx_video/configs/ltxv-13b-0.9.7-dev.yaml"
|
||||
# check if pipeline_config is a file
|
||||
if not os.path.isfile(pipeline_config):
|
||||
raise ValueError(f"Pipeline config file {pipeline_config} does not exist")
|
||||
with open(pipeline_config, "r") as f:
|
||||
pipeline_config = yaml.safe_load(f)
|
||||
|
||||
|
||||
# Validate conditioning arguments
|
||||
if conditioning_media_paths:
|
||||
# Use default strengths of 1.0
|
||||
if not conditioning_strengths:
|
||||
conditioning_strengths = [1.0] * len(conditioning_media_paths)
|
||||
if not conditioning_start_frames:
|
||||
raise ValueError(
|
||||
"If `conditioning_media_paths` is provided, "
|
||||
"`conditioning_start_frames` must also be provided"
|
||||
)
|
||||
if len(conditioning_media_paths) != len(conditioning_strengths) or len(
|
||||
conditioning_media_paths
|
||||
) != len(conditioning_start_frames):
|
||||
raise ValueError(
|
||||
"`conditioning_media_paths`, `conditioning_strengths`, "
|
||||
"and `conditioning_start_frames` must have the same length"
|
||||
)
|
||||
if any(s < 0 or s > 1 for s in conditioning_strengths):
|
||||
raise ValueError("All conditioning strengths must be between 0 and 1")
|
||||
if any(f < 0 or f >= frame_num for f in conditioning_start_frames):
|
||||
raise ValueError(
|
||||
f"All conditioning start frames must be between 0 and {frame_num-1}"
|
||||
)
|
||||
|
||||
# Adjust dimensions to be divisible by 32 and num_frames to be (N * 8 + 1)
|
||||
height_padded = ((height - 1) // 32 + 1) * 32
|
||||
width_padded = ((width - 1) // 32 + 1) * 32
|
||||
num_frames_padded = ((frame_num - 2) // 8 + 1) * 8 + 1
|
||||
|
||||
padding = calculate_padding(height, width, height_padded, width_padded)
|
||||
|
||||
logger.warning(
|
||||
f"Padded dimensions: {height_padded}x{width_padded}x{num_frames_padded}"
|
||||
)
|
||||
|
||||
|
||||
# prompt_enhancement_words_threshold = pipeline_config[
|
||||
# "prompt_enhancement_words_threshold"
|
||||
# ]
|
||||
|
||||
# prompt_word_count = len(prompt.split())
|
||||
# enhance_prompt = (
|
||||
# prompt_enhancement_words_threshold > 0
|
||||
# and prompt_word_count < prompt_enhancement_words_threshold
|
||||
# )
|
||||
|
||||
# # enhance_prompt = False
|
||||
|
||||
# if prompt_enhancement_words_threshold > 0 and not enhance_prompt:
|
||||
# logger.info(
|
||||
# f"Prompt has {prompt_word_count} words, which exceeds the threshold of {prompt_enhancement_words_threshold}. Prompt enhancement disabled."
|
||||
# )
|
||||
|
||||
|
||||
seed_everething(seed)
|
||||
device = device or get_device()
|
||||
generator = torch.Generator(device=device).manual_seed(seed)
|
||||
|
||||
media_item = None
|
||||
if input_media_path:
|
||||
media_item = load_media_file(
|
||||
media_path=input_media_path,
|
||||
height=height,
|
||||
width=width,
|
||||
max_frames=num_frames_padded,
|
||||
padding=padding,
|
||||
)
|
||||
|
||||
conditioning_items = (
|
||||
prepare_conditioning(
|
||||
conditioning_media_paths=conditioning_media_paths,
|
||||
conditioning_strengths=conditioning_strengths,
|
||||
conditioning_start_frames=conditioning_start_frames,
|
||||
height=height,
|
||||
width=width,
|
||||
num_frames=frame_num,
|
||||
padding=padding,
|
||||
pipeline=self.pipeline,
|
||||
)
|
||||
if conditioning_media_paths
|
||||
else None
|
||||
)
|
||||
|
||||
stg_mode = pipeline_config.get("stg_mode", "attention_values")
|
||||
del pipeline_config["stg_mode"]
|
||||
if stg_mode.lower() == "stg_av" or stg_mode.lower() == "attention_values":
|
||||
skip_layer_strategy = SkipLayerStrategy.AttentionValues
|
||||
elif stg_mode.lower() == "stg_as" or stg_mode.lower() == "attention_skip":
|
||||
skip_layer_strategy = SkipLayerStrategy.AttentionSkip
|
||||
elif stg_mode.lower() == "stg_r" or stg_mode.lower() == "residual":
|
||||
skip_layer_strategy = SkipLayerStrategy.Residual
|
||||
elif stg_mode.lower() == "stg_t" or stg_mode.lower() == "transformer_block":
|
||||
skip_layer_strategy = SkipLayerStrategy.TransformerBlock
|
||||
else:
|
||||
raise ValueError(f"Invalid spatiotemporal guidance mode: {stg_mode}")
|
||||
|
||||
# Prepare input for the pipeline
|
||||
sample = {
|
||||
"prompt": input_prompt,
|
||||
"prompt_attention_mask": None,
|
||||
"negative_prompt": n_prompt,
|
||||
"negative_prompt_attention_mask": None,
|
||||
}
|
||||
|
||||
|
||||
images = self.pipeline(
|
||||
**pipeline_config,
|
||||
ltxv_model = self,
|
||||
num_inference_steps1 = num_inference_steps1,
|
||||
num_inference_steps2 = num_inference_steps2,
|
||||
skip_layer_strategy=skip_layer_strategy,
|
||||
generator=generator,
|
||||
output_type="pt",
|
||||
callback_on_step_end=None,
|
||||
height=height_padded,
|
||||
width=width_padded,
|
||||
num_frames=num_frames_padded,
|
||||
frame_rate=frame_rate,
|
||||
**sample,
|
||||
media_items=media_item,
|
||||
strength=strength,
|
||||
conditioning_items=conditioning_items,
|
||||
is_video=True,
|
||||
vae_per_channel_normalize=True,
|
||||
image_cond_noise_scale=image_cond_noise_scale,
|
||||
mixed_precision=pipeline_config.get("mixed", self.mixed_precision_transformer),
|
||||
callback=callback,
|
||||
VAE_tile_size = VAE_tile_size,
|
||||
device=device,
|
||||
# enhance_prompt=enhance_prompt,
|
||||
)
|
||||
if images == None:
|
||||
return None
|
||||
|
||||
# Crop the padded images to the desired resolution and number of frames
|
||||
(pad_left, pad_right, pad_top, pad_bottom) = padding
|
||||
pad_bottom = -pad_bottom
|
||||
pad_right = -pad_right
|
||||
if pad_bottom == 0:
|
||||
pad_bottom = images.shape[3]
|
||||
if pad_right == 0:
|
||||
pad_right = images.shape[4]
|
||||
images = images[:, :, :frame_num, pad_top:pad_bottom, pad_left:pad_right]
|
||||
images = images.sub_(0.5).mul_(2).squeeze(0)
|
||||
return images
|
||||
|
||||
|
||||
def prepare_conditioning(
|
||||
conditioning_media_paths: List[str],
|
||||
conditioning_strengths: List[float],
|
||||
conditioning_start_frames: List[int],
|
||||
height: int,
|
||||
width: int,
|
||||
num_frames: int,
|
||||
padding: tuple[int, int, int, int],
|
||||
pipeline: LTXVideoPipeline,
|
||||
) -> Optional[List[ConditioningItem]]:
|
||||
"""Prepare conditioning items based on input media paths and their parameters.
|
||||
|
||||
Args:
|
||||
conditioning_media_paths: List of paths to conditioning media (images or videos)
|
||||
conditioning_strengths: List of conditioning strengths for each media item
|
||||
conditioning_start_frames: List of frame indices where each item should be applied
|
||||
height: Height of the output frames
|
||||
width: Width of the output frames
|
||||
num_frames: Number of frames in the output video
|
||||
padding: Padding to apply to the frames
|
||||
pipeline: LTXVideoPipeline object used for condition video trimming
|
||||
|
||||
Returns:
|
||||
A list of ConditioningItem objects.
|
||||
"""
|
||||
conditioning_items = []
|
||||
for path, strength, start_frame in zip(
|
||||
conditioning_media_paths, conditioning_strengths, conditioning_start_frames
|
||||
):
|
||||
if isinstance(path, Image.Image):
|
||||
num_input_frames = orig_num_input_frames = 1
|
||||
else:
|
||||
num_input_frames = orig_num_input_frames = get_media_num_frames(path)
|
||||
if hasattr(pipeline, "trim_conditioning_sequence") and callable(
|
||||
getattr(pipeline, "trim_conditioning_sequence")
|
||||
):
|
||||
num_input_frames = pipeline.trim_conditioning_sequence(
|
||||
start_frame, orig_num_input_frames, num_frames
|
||||
)
|
||||
if num_input_frames < orig_num_input_frames:
|
||||
logger.warning(
|
||||
f"Trimming conditioning video {path} from {orig_num_input_frames} to {num_input_frames} frames."
|
||||
)
|
||||
|
||||
media_tensor = load_media_file(
|
||||
media_path=path,
|
||||
height=height,
|
||||
width=width,
|
||||
max_frames=num_input_frames,
|
||||
padding=padding,
|
||||
just_crop=True,
|
||||
)
|
||||
conditioning_items.append(ConditioningItem(media_tensor, start_frame, strength))
|
||||
return conditioning_items
|
||||
|
||||
|
||||
def get_media_num_frames(media_path: str) -> int:
|
||||
if isinstance(media_path, Image.Image):
|
||||
return 1
|
||||
elif torch.is_tensor(media_path):
|
||||
return media_path.shape[1]
|
||||
elif isinstance(media_path, str) and any( media_path.lower().endswith(ext) for ext in [".mp4", ".avi", ".mov", ".mkv"]):
|
||||
reader = imageio.get_reader(media_path)
|
||||
return min(reader.count_frames(), max_frames)
|
||||
else:
|
||||
raise Exception("video format not supported")
|
||||
|
||||
|
||||
def load_media_file(
|
||||
media_path: str,
|
||||
height: int,
|
||||
width: int,
|
||||
max_frames: int,
|
||||
padding: tuple[int, int, int, int],
|
||||
just_crop: bool = False,
|
||||
) -> torch.Tensor:
|
||||
if isinstance(media_path, Image.Image):
|
||||
# Input image
|
||||
media_tensor = load_image_to_tensor_with_resize_and_crop(
|
||||
media_path, height, width, just_crop=just_crop
|
||||
)
|
||||
media_tensor = torch.nn.functional.pad(media_tensor, padding)
|
||||
|
||||
elif torch.is_tensor(media_path):
|
||||
media_tensor = media_path.unsqueeze(0)
|
||||
num_input_frames = media_tensor.shape[2]
|
||||
elif isinstance(media_path, str) and any( media_path.lower().endswith(ext) for ext in [".mp4", ".avi", ".mov", ".mkv"]):
|
||||
reader = imageio.get_reader(media_path)
|
||||
num_input_frames = min(reader.count_frames(), max_frames)
|
||||
|
||||
# Read and preprocess the relevant frames from the video file.
|
||||
frames = []
|
||||
for i in range(num_input_frames):
|
||||
frame = Image.fromarray(reader.get_data(i))
|
||||
frame_tensor = load_image_to_tensor_with_resize_and_crop(
|
||||
frame, height, width, just_crop=just_crop
|
||||
)
|
||||
frame_tensor = torch.nn.functional.pad(frame_tensor, padding)
|
||||
frames.append(frame_tensor)
|
||||
reader.close()
|
||||
|
||||
# Stack frames along the temporal dimension
|
||||
media_tensor = torch.cat(frames, dim=2)
|
||||
else:
|
||||
raise Exception("video format not supported")
|
||||
return media_tensor
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
0
ltx_video/models/__init__.py
Normal file
0
ltx_video/models/__init__.py
Normal file
0
ltx_video/models/autoencoders/__init__.py
Normal file
0
ltx_video/models/autoencoders/__init__.py
Normal file
63
ltx_video/models/autoencoders/causal_conv3d.py
Normal file
63
ltx_video/models/autoencoders/causal_conv3d.py
Normal file
@@ -0,0 +1,63 @@
|
||||
from typing import Tuple, Union
|
||||
|
||||
import torch
|
||||
import torch.nn as nn
|
||||
|
||||
|
||||
class CausalConv3d(nn.Module):
|
||||
def __init__(
|
||||
self,
|
||||
in_channels,
|
||||
out_channels,
|
||||
kernel_size: int = 3,
|
||||
stride: Union[int, Tuple[int]] = 1,
|
||||
dilation: int = 1,
|
||||
groups: int = 1,
|
||||
spatial_padding_mode: str = "zeros",
|
||||
**kwargs,
|
||||
):
|
||||
super().__init__()
|
||||
|
||||
self.in_channels = in_channels
|
||||
self.out_channels = out_channels
|
||||
|
||||
kernel_size = (kernel_size, kernel_size, kernel_size)
|
||||
self.time_kernel_size = kernel_size[0]
|
||||
|
||||
dilation = (dilation, 1, 1)
|
||||
|
||||
height_pad = kernel_size[1] // 2
|
||||
width_pad = kernel_size[2] // 2
|
||||
padding = (0, height_pad, width_pad)
|
||||
|
||||
self.conv = nn.Conv3d(
|
||||
in_channels,
|
||||
out_channels,
|
||||
kernel_size,
|
||||
stride=stride,
|
||||
dilation=dilation,
|
||||
padding=padding,
|
||||
padding_mode=spatial_padding_mode,
|
||||
groups=groups,
|
||||
)
|
||||
|
||||
def forward(self, x, causal: bool = True):
|
||||
if causal:
|
||||
first_frame_pad = x[:, :, :1, :, :].repeat(
|
||||
(1, 1, self.time_kernel_size - 1, 1, 1)
|
||||
)
|
||||
x = torch.concatenate((first_frame_pad, x), dim=2)
|
||||
else:
|
||||
first_frame_pad = x[:, :, :1, :, :].repeat(
|
||||
(1, 1, (self.time_kernel_size - 1) // 2, 1, 1)
|
||||
)
|
||||
last_frame_pad = x[:, :, -1:, :, :].repeat(
|
||||
(1, 1, (self.time_kernel_size - 1) // 2, 1, 1)
|
||||
)
|
||||
x = torch.concatenate((first_frame_pad, x, last_frame_pad), dim=2)
|
||||
x = self.conv(x)
|
||||
return x
|
||||
|
||||
@property
|
||||
def weight(self):
|
||||
return self.conv.weight
|
||||
1405
ltx_video/models/autoencoders/causal_video_autoencoder.py
Normal file
1405
ltx_video/models/autoencoders/causal_video_autoencoder.py
Normal file
File diff suppressed because it is too large
Load Diff
90
ltx_video/models/autoencoders/conv_nd_factory.py
Normal file
90
ltx_video/models/autoencoders/conv_nd_factory.py
Normal file
@@ -0,0 +1,90 @@
|
||||
from typing import Tuple, Union
|
||||
|
||||
import torch
|
||||
|
||||
from ltx_video.models.autoencoders.dual_conv3d import DualConv3d
|
||||
from ltx_video.models.autoencoders.causal_conv3d import CausalConv3d
|
||||
|
||||
|
||||
def make_conv_nd(
|
||||
dims: Union[int, Tuple[int, int]],
|
||||
in_channels: int,
|
||||
out_channels: int,
|
||||
kernel_size: int,
|
||||
stride=1,
|
||||
padding=0,
|
||||
dilation=1,
|
||||
groups=1,
|
||||
bias=True,
|
||||
causal=False,
|
||||
spatial_padding_mode="zeros",
|
||||
temporal_padding_mode="zeros",
|
||||
):
|
||||
if not (spatial_padding_mode == temporal_padding_mode or causal):
|
||||
raise NotImplementedError("spatial and temporal padding modes must be equal")
|
||||
if dims == 2:
|
||||
return torch.nn.Conv2d(
|
||||
in_channels=in_channels,
|
||||
out_channels=out_channels,
|
||||
kernel_size=kernel_size,
|
||||
stride=stride,
|
||||
padding=padding,
|
||||
dilation=dilation,
|
||||
groups=groups,
|
||||
bias=bias,
|
||||
padding_mode=spatial_padding_mode,
|
||||
)
|
||||
elif dims == 3:
|
||||
if causal:
|
||||
return CausalConv3d(
|
||||
in_channels=in_channels,
|
||||
out_channels=out_channels,
|
||||
kernel_size=kernel_size,
|
||||
stride=stride,
|
||||
padding=padding,
|
||||
dilation=dilation,
|
||||
groups=groups,
|
||||
bias=bias,
|
||||
spatial_padding_mode=spatial_padding_mode,
|
||||
)
|
||||
return torch.nn.Conv3d(
|
||||
in_channels=in_channels,
|
||||
out_channels=out_channels,
|
||||
kernel_size=kernel_size,
|
||||
stride=stride,
|
||||
padding=padding,
|
||||
dilation=dilation,
|
||||
groups=groups,
|
||||
bias=bias,
|
||||
padding_mode=spatial_padding_mode,
|
||||
)
|
||||
elif dims == (2, 1):
|
||||
return DualConv3d(
|
||||
in_channels=in_channels,
|
||||
out_channels=out_channels,
|
||||
kernel_size=kernel_size,
|
||||
stride=stride,
|
||||
padding=padding,
|
||||
bias=bias,
|
||||
padding_mode=spatial_padding_mode,
|
||||
)
|
||||
else:
|
||||
raise ValueError(f"unsupported dimensions: {dims}")
|
||||
|
||||
|
||||
def make_linear_nd(
|
||||
dims: int,
|
||||
in_channels: int,
|
||||
out_channels: int,
|
||||
bias=True,
|
||||
):
|
||||
if dims == 2:
|
||||
return torch.nn.Conv2d(
|
||||
in_channels=in_channels, out_channels=out_channels, kernel_size=1, bias=bias
|
||||
)
|
||||
elif dims == 3 or dims == (2, 1):
|
||||
return torch.nn.Conv3d(
|
||||
in_channels=in_channels, out_channels=out_channels, kernel_size=1, bias=bias
|
||||
)
|
||||
else:
|
||||
raise ValueError(f"unsupported dimensions: {dims}")
|
||||
217
ltx_video/models/autoencoders/dual_conv3d.py
Normal file
217
ltx_video/models/autoencoders/dual_conv3d.py
Normal file
@@ -0,0 +1,217 @@
|
||||
import math
|
||||
from typing import Tuple, Union
|
||||
|
||||
import torch
|
||||
import torch.nn as nn
|
||||
import torch.nn.functional as F
|
||||
from einops import rearrange
|
||||
|
||||
|
||||
class DualConv3d(nn.Module):
|
||||
def __init__(
|
||||
self,
|
||||
in_channels,
|
||||
out_channels,
|
||||
kernel_size,
|
||||
stride: Union[int, Tuple[int, int, int]] = 1,
|
||||
padding: Union[int, Tuple[int, int, int]] = 0,
|
||||
dilation: Union[int, Tuple[int, int, int]] = 1,
|
||||
groups=1,
|
||||
bias=True,
|
||||
padding_mode="zeros",
|
||||
):
|
||||
super(DualConv3d, self).__init__()
|
||||
|
||||
self.in_channels = in_channels
|
||||
self.out_channels = out_channels
|
||||
self.padding_mode = padding_mode
|
||||
# Ensure kernel_size, stride, padding, and dilation are tuples of length 3
|
||||
if isinstance(kernel_size, int):
|
||||
kernel_size = (kernel_size, kernel_size, kernel_size)
|
||||
if kernel_size == (1, 1, 1):
|
||||
raise ValueError(
|
||||
"kernel_size must be greater than 1. Use make_linear_nd instead."
|
||||
)
|
||||
if isinstance(stride, int):
|
||||
stride = (stride, stride, stride)
|
||||
if isinstance(padding, int):
|
||||
padding = (padding, padding, padding)
|
||||
if isinstance(dilation, int):
|
||||
dilation = (dilation, dilation, dilation)
|
||||
|
||||
# Set parameters for convolutions
|
||||
self.groups = groups
|
||||
self.bias = bias
|
||||
|
||||
# Define the size of the channels after the first convolution
|
||||
intermediate_channels = (
|
||||
out_channels if in_channels < out_channels else in_channels
|
||||
)
|
||||
|
||||
# Define parameters for the first convolution
|
||||
self.weight1 = nn.Parameter(
|
||||
torch.Tensor(
|
||||
intermediate_channels,
|
||||
in_channels // groups,
|
||||
1,
|
||||
kernel_size[1],
|
||||
kernel_size[2],
|
||||
)
|
||||
)
|
||||
self.stride1 = (1, stride[1], stride[2])
|
||||
self.padding1 = (0, padding[1], padding[2])
|
||||
self.dilation1 = (1, dilation[1], dilation[2])
|
||||
if bias:
|
||||
self.bias1 = nn.Parameter(torch.Tensor(intermediate_channels))
|
||||
else:
|
||||
self.register_parameter("bias1", None)
|
||||
|
||||
# Define parameters for the second convolution
|
||||
self.weight2 = nn.Parameter(
|
||||
torch.Tensor(
|
||||
out_channels, intermediate_channels // groups, kernel_size[0], 1, 1
|
||||
)
|
||||
)
|
||||
self.stride2 = (stride[0], 1, 1)
|
||||
self.padding2 = (padding[0], 0, 0)
|
||||
self.dilation2 = (dilation[0], 1, 1)
|
||||
if bias:
|
||||
self.bias2 = nn.Parameter(torch.Tensor(out_channels))
|
||||
else:
|
||||
self.register_parameter("bias2", None)
|
||||
|
||||
# Initialize weights and biases
|
||||
self.reset_parameters()
|
||||
|
||||
def reset_parameters(self):
|
||||
nn.init.kaiming_uniform_(self.weight1, a=math.sqrt(5))
|
||||
nn.init.kaiming_uniform_(self.weight2, a=math.sqrt(5))
|
||||
if self.bias:
|
||||
fan_in1, _ = nn.init._calculate_fan_in_and_fan_out(self.weight1)
|
||||
bound1 = 1 / math.sqrt(fan_in1)
|
||||
nn.init.uniform_(self.bias1, -bound1, bound1)
|
||||
fan_in2, _ = nn.init._calculate_fan_in_and_fan_out(self.weight2)
|
||||
bound2 = 1 / math.sqrt(fan_in2)
|
||||
nn.init.uniform_(self.bias2, -bound2, bound2)
|
||||
|
||||
def forward(self, x, use_conv3d=False, skip_time_conv=False):
|
||||
if use_conv3d:
|
||||
return self.forward_with_3d(x=x, skip_time_conv=skip_time_conv)
|
||||
else:
|
||||
return self.forward_with_2d(x=x, skip_time_conv=skip_time_conv)
|
||||
|
||||
def forward_with_3d(self, x, skip_time_conv):
|
||||
# First convolution
|
||||
x = F.conv3d(
|
||||
x,
|
||||
self.weight1,
|
||||
self.bias1,
|
||||
self.stride1,
|
||||
self.padding1,
|
||||
self.dilation1,
|
||||
self.groups,
|
||||
padding_mode=self.padding_mode,
|
||||
)
|
||||
|
||||
if skip_time_conv:
|
||||
return x
|
||||
|
||||
# Second convolution
|
||||
x = F.conv3d(
|
||||
x,
|
||||
self.weight2,
|
||||
self.bias2,
|
||||
self.stride2,
|
||||
self.padding2,
|
||||
self.dilation2,
|
||||
self.groups,
|
||||
padding_mode=self.padding_mode,
|
||||
)
|
||||
|
||||
return x
|
||||
|
||||
def forward_with_2d(self, x, skip_time_conv):
|
||||
b, c, d, h, w = x.shape
|
||||
|
||||
# First 2D convolution
|
||||
x = rearrange(x, "b c d h w -> (b d) c h w")
|
||||
# Squeeze the depth dimension out of weight1 since it's 1
|
||||
weight1 = self.weight1.squeeze(2)
|
||||
# Select stride, padding, and dilation for the 2D convolution
|
||||
stride1 = (self.stride1[1], self.stride1[2])
|
||||
padding1 = (self.padding1[1], self.padding1[2])
|
||||
dilation1 = (self.dilation1[1], self.dilation1[2])
|
||||
x = F.conv2d(
|
||||
x,
|
||||
weight1,
|
||||
self.bias1,
|
||||
stride1,
|
||||
padding1,
|
||||
dilation1,
|
||||
self.groups,
|
||||
padding_mode=self.padding_mode,
|
||||
)
|
||||
|
||||
_, _, h, w = x.shape
|
||||
|
||||
if skip_time_conv:
|
||||
x = rearrange(x, "(b d) c h w -> b c d h w", b=b)
|
||||
return x
|
||||
|
||||
# Second convolution which is essentially treated as a 1D convolution across the 'd' dimension
|
||||
x = rearrange(x, "(b d) c h w -> (b h w) c d", b=b)
|
||||
|
||||
# Reshape weight2 to match the expected dimensions for conv1d
|
||||
weight2 = self.weight2.squeeze(-1).squeeze(-1)
|
||||
# Use only the relevant dimension for stride, padding, and dilation for the 1D convolution
|
||||
stride2 = self.stride2[0]
|
||||
padding2 = self.padding2[0]
|
||||
dilation2 = self.dilation2[0]
|
||||
x = F.conv1d(
|
||||
x,
|
||||
weight2,
|
||||
self.bias2,
|
||||
stride2,
|
||||
padding2,
|
||||
dilation2,
|
||||
self.groups,
|
||||
padding_mode=self.padding_mode,
|
||||
)
|
||||
x = rearrange(x, "(b h w) c d -> b c d h w", b=b, h=h, w=w)
|
||||
|
||||
return x
|
||||
|
||||
@property
|
||||
def weight(self):
|
||||
return self.weight2
|
||||
|
||||
|
||||
def test_dual_conv3d_consistency():
|
||||
# Initialize parameters
|
||||
in_channels = 3
|
||||
out_channels = 5
|
||||
kernel_size = (3, 3, 3)
|
||||
stride = (2, 2, 2)
|
||||
padding = (1, 1, 1)
|
||||
|
||||
# Create an instance of the DualConv3d class
|
||||
dual_conv3d = DualConv3d(
|
||||
in_channels=in_channels,
|
||||
out_channels=out_channels,
|
||||
kernel_size=kernel_size,
|
||||
stride=stride,
|
||||
padding=padding,
|
||||
bias=True,
|
||||
)
|
||||
|
||||
# Example input tensor
|
||||
test_input = torch.randn(1, 3, 10, 10, 10)
|
||||
|
||||
# Perform forward passes with both 3D and 2D settings
|
||||
output_conv3d = dual_conv3d(test_input, use_conv3d=True)
|
||||
output_2d = dual_conv3d(test_input, use_conv3d=False)
|
||||
|
||||
# Assert that the outputs from both methods are sufficiently close
|
||||
assert torch.allclose(
|
||||
output_conv3d, output_2d, atol=1e-6
|
||||
), "Outputs are not consistent between 3D and 2D convolutions."
|
||||
203
ltx_video/models/autoencoders/latent_upsampler.py
Normal file
203
ltx_video/models/autoencoders/latent_upsampler.py
Normal file
@@ -0,0 +1,203 @@
|
||||
from typing import Optional, Union
|
||||
from pathlib import Path
|
||||
import os
|
||||
import json
|
||||
|
||||
import torch
|
||||
import torch.nn as nn
|
||||
from einops import rearrange
|
||||
from diffusers import ConfigMixin, ModelMixin
|
||||
from safetensors.torch import safe_open
|
||||
|
||||
from ltx_video.models.autoencoders.pixel_shuffle import PixelShuffleND
|
||||
|
||||
|
||||
class ResBlock(nn.Module):
|
||||
def __init__(
|
||||
self, channels: int, mid_channels: Optional[int] = None, dims: int = 3
|
||||
):
|
||||
super().__init__()
|
||||
if mid_channels is None:
|
||||
mid_channels = channels
|
||||
|
||||
Conv = nn.Conv2d if dims == 2 else nn.Conv3d
|
||||
|
||||
self.conv1 = Conv(channels, mid_channels, kernel_size=3, padding=1)
|
||||
self.norm1 = nn.GroupNorm(32, mid_channels)
|
||||
self.conv2 = Conv(mid_channels, channels, kernel_size=3, padding=1)
|
||||
self.norm2 = nn.GroupNorm(32, channels)
|
||||
self.activation = nn.SiLU()
|
||||
|
||||
def forward(self, x: torch.Tensor) -> torch.Tensor:
|
||||
residual = x
|
||||
x = self.conv1(x)
|
||||
x = self.norm1(x)
|
||||
x = self.activation(x)
|
||||
x = self.conv2(x)
|
||||
x = self.norm2(x)
|
||||
x = self.activation(x + residual)
|
||||
return x
|
||||
|
||||
|
||||
class LatentUpsampler(ModelMixin, ConfigMixin):
|
||||
"""
|
||||
Model to spatially upsample VAE latents.
|
||||
|
||||
Args:
|
||||
in_channels (`int`): Number of channels in the input latent
|
||||
mid_channels (`int`): Number of channels in the middle layers
|
||||
num_blocks_per_stage (`int`): Number of ResBlocks to use in each stage (pre/post upsampling)
|
||||
dims (`int`): Number of dimensions for convolutions (2 or 3)
|
||||
spatial_upsample (`bool`): Whether to spatially upsample the latent
|
||||
temporal_upsample (`bool`): Whether to temporally upsample the latent
|
||||
"""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
in_channels: int = 128,
|
||||
mid_channels: int = 512,
|
||||
num_blocks_per_stage: int = 4,
|
||||
dims: int = 3,
|
||||
spatial_upsample: bool = True,
|
||||
temporal_upsample: bool = False,
|
||||
):
|
||||
super().__init__()
|
||||
|
||||
self.in_channels = in_channels
|
||||
self.mid_channels = mid_channels
|
||||
self.num_blocks_per_stage = num_blocks_per_stage
|
||||
self.dims = dims
|
||||
self.spatial_upsample = spatial_upsample
|
||||
self.temporal_upsample = temporal_upsample
|
||||
|
||||
Conv = nn.Conv2d if dims == 2 else nn.Conv3d
|
||||
|
||||
self.initial_conv = Conv(in_channels, mid_channels, kernel_size=3, padding=1)
|
||||
self.initial_norm = nn.GroupNorm(32, mid_channels)
|
||||
self.initial_activation = nn.SiLU()
|
||||
|
||||
self.res_blocks = nn.ModuleList(
|
||||
[ResBlock(mid_channels, dims=dims) for _ in range(num_blocks_per_stage)]
|
||||
)
|
||||
|
||||
if spatial_upsample and temporal_upsample:
|
||||
self.upsampler = nn.Sequential(
|
||||
nn.Conv3d(mid_channels, 8 * mid_channels, kernel_size=3, padding=1),
|
||||
PixelShuffleND(3),
|
||||
)
|
||||
elif spatial_upsample:
|
||||
self.upsampler = nn.Sequential(
|
||||
nn.Conv2d(mid_channels, 4 * mid_channels, kernel_size=3, padding=1),
|
||||
PixelShuffleND(2),
|
||||
)
|
||||
elif temporal_upsample:
|
||||
self.upsampler = nn.Sequential(
|
||||
nn.Conv3d(mid_channels, 2 * mid_channels, kernel_size=3, padding=1),
|
||||
PixelShuffleND(1),
|
||||
)
|
||||
else:
|
||||
raise ValueError(
|
||||
"Either spatial_upsample or temporal_upsample must be True"
|
||||
)
|
||||
|
||||
self.post_upsample_res_blocks = nn.ModuleList(
|
||||
[ResBlock(mid_channels, dims=dims) for _ in range(num_blocks_per_stage)]
|
||||
)
|
||||
|
||||
self.final_conv = Conv(mid_channels, in_channels, kernel_size=3, padding=1)
|
||||
|
||||
def forward(self, latent: torch.Tensor) -> torch.Tensor:
|
||||
b, c, f, h, w = latent.shape
|
||||
|
||||
if self.dims == 2:
|
||||
x = rearrange(latent, "b c f h w -> (b f) c h w")
|
||||
x = self.initial_conv(x)
|
||||
x = self.initial_norm(x)
|
||||
x = self.initial_activation(x)
|
||||
|
||||
for block in self.res_blocks:
|
||||
x = block(x)
|
||||
|
||||
x = self.upsampler(x)
|
||||
|
||||
for block in self.post_upsample_res_blocks:
|
||||
x = block(x)
|
||||
|
||||
x = self.final_conv(x)
|
||||
x = rearrange(x, "(b f) c h w -> b c f h w", b=b, f=f)
|
||||
else:
|
||||
x = self.initial_conv(latent)
|
||||
x = self.initial_norm(x)
|
||||
x = self.initial_activation(x)
|
||||
|
||||
for block in self.res_blocks:
|
||||
x = block(x)
|
||||
|
||||
if self.temporal_upsample:
|
||||
x = self.upsampler(x)
|
||||
x = x[:, :, 1:, :, :]
|
||||
else:
|
||||
x = rearrange(x, "b c f h w -> (b f) c h w")
|
||||
x = self.upsampler(x)
|
||||
x = rearrange(x, "(b f) c h w -> b c f h w", b=b, f=f)
|
||||
|
||||
for block in self.post_upsample_res_blocks:
|
||||
x = block(x)
|
||||
|
||||
x = self.final_conv(x)
|
||||
|
||||
return x
|
||||
|
||||
@classmethod
|
||||
def from_config(cls, config):
|
||||
return cls(
|
||||
in_channels=config.get("in_channels", 4),
|
||||
mid_channels=config.get("mid_channels", 128),
|
||||
num_blocks_per_stage=config.get("num_blocks_per_stage", 4),
|
||||
dims=config.get("dims", 2),
|
||||
spatial_upsample=config.get("spatial_upsample", True),
|
||||
temporal_upsample=config.get("temporal_upsample", False),
|
||||
)
|
||||
|
||||
def config(self):
|
||||
return {
|
||||
"_class_name": "LatentUpsampler",
|
||||
"in_channels": self.in_channels,
|
||||
"mid_channels": self.mid_channels,
|
||||
"num_blocks_per_stage": self.num_blocks_per_stage,
|
||||
"dims": self.dims,
|
||||
"spatial_upsample": self.spatial_upsample,
|
||||
"temporal_upsample": self.temporal_upsample,
|
||||
}
|
||||
|
||||
@classmethod
|
||||
def from_pretrained(
|
||||
cls,
|
||||
pretrained_model_path: Optional[Union[str, os.PathLike]],
|
||||
*args,
|
||||
**kwargs,
|
||||
):
|
||||
pretrained_model_path = Path(pretrained_model_path)
|
||||
if pretrained_model_path.is_file() and str(pretrained_model_path).endswith(
|
||||
".safetensors"
|
||||
):
|
||||
state_dict = {}
|
||||
with safe_open(pretrained_model_path, framework="pt", device="cpu") as f:
|
||||
metadata = f.metadata()
|
||||
for k in f.keys():
|
||||
state_dict[k] = f.get_tensor(k)
|
||||
config = json.loads(metadata["config"])
|
||||
with torch.device("meta"):
|
||||
latent_upsampler = LatentUpsampler.from_config(config)
|
||||
latent_upsampler.load_state_dict(state_dict, assign=True)
|
||||
return latent_upsampler
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
latent_upsampler = LatentUpsampler(num_blocks_per_stage=4, dims=3)
|
||||
print(latent_upsampler)
|
||||
total_params = sum(p.numel() for p in latent_upsampler.parameters())
|
||||
print(f"Total number of parameters: {total_params:,}")
|
||||
latent = torch.randn(1, 128, 9, 16, 16)
|
||||
upsampled_latent = latent_upsampler(latent)
|
||||
print(f"Upsampled latent shape: {upsampled_latent.shape}")
|
||||
12
ltx_video/models/autoencoders/pixel_norm.py
Normal file
12
ltx_video/models/autoencoders/pixel_norm.py
Normal file
@@ -0,0 +1,12 @@
|
||||
import torch
|
||||
from torch import nn
|
||||
|
||||
|
||||
class PixelNorm(nn.Module):
|
||||
def __init__(self, dim=1, eps=1e-8):
|
||||
super(PixelNorm, self).__init__()
|
||||
self.dim = dim
|
||||
self.eps = eps
|
||||
|
||||
def forward(self, x):
|
||||
return x / torch.sqrt(torch.mean(x**2, dim=self.dim, keepdim=True) + self.eps)
|
||||
33
ltx_video/models/autoencoders/pixel_shuffle.py
Normal file
33
ltx_video/models/autoencoders/pixel_shuffle.py
Normal file
@@ -0,0 +1,33 @@
|
||||
import torch.nn as nn
|
||||
from einops import rearrange
|
||||
|
||||
|
||||
class PixelShuffleND(nn.Module):
|
||||
def __init__(self, dims, upscale_factors=(2, 2, 2)):
|
||||
super().__init__()
|
||||
assert dims in [1, 2, 3], "dims must be 1, 2, or 3"
|
||||
self.dims = dims
|
||||
self.upscale_factors = upscale_factors
|
||||
|
||||
def forward(self, x):
|
||||
if self.dims == 3:
|
||||
return rearrange(
|
||||
x,
|
||||
"b (c p1 p2 p3) d h w -> b c (d p1) (h p2) (w p3)",
|
||||
p1=self.upscale_factors[0],
|
||||
p2=self.upscale_factors[1],
|
||||
p3=self.upscale_factors[2],
|
||||
)
|
||||
elif self.dims == 2:
|
||||
return rearrange(
|
||||
x,
|
||||
"b (c p1 p2) h w -> b c (h p1) (w p2)",
|
||||
p1=self.upscale_factors[0],
|
||||
p2=self.upscale_factors[1],
|
||||
)
|
||||
elif self.dims == 1:
|
||||
return rearrange(
|
||||
x,
|
||||
"b (c p1) f h w -> b c (f p1) h w",
|
||||
p1=self.upscale_factors[0],
|
||||
)
|
||||
443
ltx_video/models/autoencoders/vae.py
Normal file
443
ltx_video/models/autoencoders/vae.py
Normal file
@@ -0,0 +1,443 @@
|
||||
from typing import Optional, Union
|
||||
|
||||
import torch
|
||||
import inspect
|
||||
import math
|
||||
import torch.nn as nn
|
||||
from diffusers import ConfigMixin, ModelMixin
|
||||
from diffusers.models.autoencoders.vae import (
|
||||
DecoderOutput,
|
||||
DiagonalGaussianDistribution,
|
||||
)
|
||||
from diffusers.models.modeling_outputs import AutoencoderKLOutput
|
||||
from ltx_video.models.autoencoders.conv_nd_factory import make_conv_nd
|
||||
|
||||
|
||||
class AutoencoderKLWrapper(ModelMixin, ConfigMixin):
|
||||
"""Variational Autoencoder (VAE) model with KL loss.
|
||||
|
||||
VAE from the paper Auto-Encoding Variational Bayes by Diederik P. Kingma and Max Welling.
|
||||
This model is a wrapper around an encoder and a decoder, and it adds a KL loss term to the reconstruction loss.
|
||||
|
||||
Args:
|
||||
encoder (`nn.Module`):
|
||||
Encoder module.
|
||||
decoder (`nn.Module`):
|
||||
Decoder module.
|
||||
latent_channels (`int`, *optional*, defaults to 4):
|
||||
Number of latent channels.
|
||||
"""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
encoder: nn.Module,
|
||||
decoder: nn.Module,
|
||||
latent_channels: int = 4,
|
||||
dims: int = 2,
|
||||
sample_size=512,
|
||||
use_quant_conv: bool = True,
|
||||
normalize_latent_channels: bool = False,
|
||||
):
|
||||
super().__init__()
|
||||
|
||||
|
||||
self.per_channel_statistics = nn.Module()
|
||||
std_of_means = torch.zeros( (128,), dtype= torch.bfloat16)
|
||||
|
||||
self.per_channel_statistics.register_buffer("std-of-means", std_of_means)
|
||||
self.per_channel_statistics.register_buffer(
|
||||
"mean-of-means",
|
||||
torch.zeros_like(std_of_means)
|
||||
)
|
||||
|
||||
|
||||
|
||||
# pass init params to Encoder
|
||||
self.encoder = encoder
|
||||
self.use_quant_conv = use_quant_conv
|
||||
self.normalize_latent_channels = normalize_latent_channels
|
||||
|
||||
# pass init params to Decoder
|
||||
quant_dims = 2 if dims == 2 else 3
|
||||
self.decoder = decoder
|
||||
if use_quant_conv:
|
||||
self.quant_conv = make_conv_nd(
|
||||
quant_dims, 2 * latent_channels, 2 * latent_channels, 1
|
||||
)
|
||||
self.post_quant_conv = make_conv_nd(
|
||||
quant_dims, latent_channels, latent_channels, 1
|
||||
)
|
||||
else:
|
||||
self.quant_conv = nn.Identity()
|
||||
self.post_quant_conv = nn.Identity()
|
||||
|
||||
if normalize_latent_channels:
|
||||
if dims == 2:
|
||||
self.latent_norm_out = nn.BatchNorm2d(latent_channels, affine=False)
|
||||
else:
|
||||
self.latent_norm_out = nn.BatchNorm3d(latent_channels, affine=False)
|
||||
else:
|
||||
self.latent_norm_out = nn.Identity()
|
||||
self.use_z_tiling = False
|
||||
self.use_hw_tiling = False
|
||||
self.dims = dims
|
||||
self.z_sample_size = 1
|
||||
|
||||
self.decoder_params = inspect.signature(self.decoder.forward).parameters
|
||||
|
||||
# only relevant if vae tiling is enabled
|
||||
self.set_tiling_params(sample_size=sample_size, overlap_factor=0.25)
|
||||
|
||||
@staticmethod
|
||||
def get_VAE_tile_size(vae_config, device_mem_capacity, mixed_precision):
|
||||
|
||||
z_tile = 4
|
||||
# VAE Tiling
|
||||
if vae_config == 0:
|
||||
if mixed_precision:
|
||||
device_mem_capacity = device_mem_capacity / 1.5
|
||||
if device_mem_capacity >= 24000:
|
||||
use_vae_config = 1
|
||||
elif device_mem_capacity >= 8000:
|
||||
use_vae_config = 2
|
||||
else:
|
||||
use_vae_config = 3
|
||||
else:
|
||||
use_vae_config = vae_config
|
||||
|
||||
if use_vae_config == 1:
|
||||
hw_tile = 0
|
||||
elif use_vae_config == 2:
|
||||
hw_tile = 512
|
||||
else:
|
||||
hw_tile = 256
|
||||
|
||||
return (z_tile, hw_tile)
|
||||
|
||||
def set_tiling_params(self, sample_size: int = 512, overlap_factor: float = 0.25):
|
||||
self.tile_sample_min_size = sample_size
|
||||
num_blocks = len(self.encoder.down_blocks)
|
||||
# self.tile_latent_min_size = int(sample_size / (2 ** (num_blocks - 1)))
|
||||
self.tile_latent_min_size = int(sample_size / 32)
|
||||
self.tile_overlap_factor = overlap_factor
|
||||
|
||||
def enable_z_tiling(self, z_sample_size: int = 4):
|
||||
r"""
|
||||
Enable tiling during VAE decoding.
|
||||
|
||||
When this option is enabled, the VAE will split the input tensor in tiles to compute decoding in several
|
||||
steps. This is useful to save some memory and allow larger batch sizes.
|
||||
"""
|
||||
self.use_z_tiling = z_sample_size > 1
|
||||
self.z_sample_size = z_sample_size
|
||||
assert (
|
||||
z_sample_size % 4 == 0 or z_sample_size == 1
|
||||
), f"z_sample_size must be a multiple of 4 or 1. Got {z_sample_size}."
|
||||
|
||||
def disable_z_tiling(self):
|
||||
r"""
|
||||
Disable tiling during VAE decoding. If `use_tiling` was previously invoked, this method will go back to computing
|
||||
decoding in one step.
|
||||
"""
|
||||
self.use_z_tiling = False
|
||||
|
||||
def enable_hw_tiling(self):
|
||||
r"""
|
||||
Enable tiling during VAE decoding along the height and width dimension.
|
||||
"""
|
||||
self.use_hw_tiling = True
|
||||
|
||||
def disable_hw_tiling(self):
|
||||
r"""
|
||||
Disable tiling during VAE decoding along the height and width dimension.
|
||||
"""
|
||||
self.use_hw_tiling = False
|
||||
|
||||
def _hw_tiled_encode(self, x: torch.FloatTensor, return_dict: bool = True):
|
||||
overlap_size = int(self.tile_sample_min_size * (1 - self.tile_overlap_factor))
|
||||
blend_extent = int(self.tile_latent_min_size * self.tile_overlap_factor)
|
||||
row_limit = self.tile_latent_min_size - blend_extent
|
||||
|
||||
# Split the image into 512x512 tiles and encode them separately.
|
||||
rows = []
|
||||
for i in range(0, x.shape[3], overlap_size):
|
||||
row = []
|
||||
for j in range(0, x.shape[4], overlap_size):
|
||||
tile = x[
|
||||
:,
|
||||
:,
|
||||
:,
|
||||
i : i + self.tile_sample_min_size,
|
||||
j : j + self.tile_sample_min_size,
|
||||
]
|
||||
tile = self.encoder(tile)
|
||||
tile = self.quant_conv(tile)
|
||||
row.append(tile)
|
||||
rows.append(row)
|
||||
result_rows = []
|
||||
for i, row in enumerate(rows):
|
||||
result_row = []
|
||||
for j, tile in enumerate(row):
|
||||
# blend the above tile and the left tile
|
||||
# to the current tile and add the current tile to the result row
|
||||
if i > 0:
|
||||
tile = self.blend_v(rows[i - 1][j], tile, blend_extent)
|
||||
if j > 0:
|
||||
tile = self.blend_h(row[j - 1], tile, blend_extent)
|
||||
result_row.append(tile[:, :, :, :row_limit, :row_limit])
|
||||
result_rows.append(torch.cat(result_row, dim=4))
|
||||
|
||||
moments = torch.cat(result_rows, dim=3)
|
||||
return moments
|
||||
|
||||
def blend_z(
|
||||
self, a: torch.Tensor, b: torch.Tensor, blend_extent: int
|
||||
) -> torch.Tensor:
|
||||
blend_extent = min(a.shape[2], b.shape[2], blend_extent)
|
||||
for z in range(blend_extent):
|
||||
b[:, :, z, :, :] = a[:, :, -blend_extent + z, :, :] * (
|
||||
1 - z / blend_extent
|
||||
) + b[:, :, z, :, :] * (z / blend_extent)
|
||||
return b
|
||||
|
||||
def blend_v(
|
||||
self, a: torch.Tensor, b: torch.Tensor, blend_extent: int
|
||||
) -> torch.Tensor:
|
||||
blend_extent = min(a.shape[3], b.shape[3], blend_extent)
|
||||
for y in range(blend_extent):
|
||||
b[:, :, :, y, :] = a[:, :, :, -blend_extent + y, :] * (
|
||||
1 - y / blend_extent
|
||||
) + b[:, :, :, y, :] * (y / blend_extent)
|
||||
return b
|
||||
|
||||
def blend_h(
|
||||
self, a: torch.Tensor, b: torch.Tensor, blend_extent: int
|
||||
) -> torch.Tensor:
|
||||
blend_extent = min(a.shape[4], b.shape[4], blend_extent)
|
||||
for x in range(blend_extent):
|
||||
b[:, :, :, :, x] = a[:, :, :, :, -blend_extent + x] * (
|
||||
1 - x / blend_extent
|
||||
) + b[:, :, :, :, x] * (x / blend_extent)
|
||||
return b
|
||||
|
||||
def _hw_tiled_decode(self, z: torch.FloatTensor, target_shape, timestep = None):
|
||||
overlap_size = int(self.tile_latent_min_size * (1 - self.tile_overlap_factor))
|
||||
blend_extent = int(self.tile_sample_min_size * self.tile_overlap_factor)
|
||||
row_limit = self.tile_sample_min_size - blend_extent
|
||||
tile_target_shape = (
|
||||
*target_shape[:3],
|
||||
self.tile_sample_min_size,
|
||||
self.tile_sample_min_size,
|
||||
)
|
||||
# Split z into overlapping 64x64 tiles and decode them separately.
|
||||
# The tiles have an overlap to avoid seams between tiles.
|
||||
rows = []
|
||||
for i in range(0, z.shape[3], overlap_size):
|
||||
row = []
|
||||
for j in range(0, z.shape[4], overlap_size):
|
||||
tile = z[
|
||||
:,
|
||||
:,
|
||||
:,
|
||||
i : i + self.tile_latent_min_size,
|
||||
j : j + self.tile_latent_min_size,
|
||||
]
|
||||
tile = self.post_quant_conv(tile)
|
||||
decoded = self.decoder(tile, target_shape=tile_target_shape, timestep = timestep)
|
||||
row.append(decoded)
|
||||
rows.append(row)
|
||||
result_rows = []
|
||||
for i, row in enumerate(rows):
|
||||
result_row = []
|
||||
for j, tile in enumerate(row):
|
||||
# blend the above tile and the left tile
|
||||
# to the current tile and add the current tile to the result row
|
||||
if i > 0:
|
||||
tile = self.blend_v(rows[i - 1][j], tile, blend_extent)
|
||||
if j > 0:
|
||||
tile = self.blend_h(row[j - 1], tile, blend_extent)
|
||||
result_row.append(tile[:, :, :, :row_limit, :row_limit])
|
||||
result_rows.append(torch.cat(result_row, dim=4))
|
||||
|
||||
dec = torch.cat(result_rows, dim=3)
|
||||
return dec
|
||||
|
||||
def encode(
|
||||
self, z: torch.FloatTensor, return_dict: bool = True
|
||||
) -> Union[DecoderOutput, torch.FloatTensor]:
|
||||
if self.use_z_tiling and z.shape[2] > (self.z_sample_size + 1) > 1:
|
||||
tile_latent_min_tsize = self.z_sample_size
|
||||
tile_sample_min_tsize = tile_latent_min_tsize * 8
|
||||
tile_overlap_factor = 0.25
|
||||
|
||||
B, C, T, H, W = z.shape
|
||||
overlap_size = int(tile_sample_min_tsize * (1 - tile_overlap_factor))
|
||||
blend_extent = int(tile_latent_min_tsize * tile_overlap_factor)
|
||||
t_limit = tile_latent_min_tsize - blend_extent
|
||||
|
||||
row = []
|
||||
for i in range(0, T, overlap_size):
|
||||
tile = z[:, :, i: i + tile_sample_min_tsize + 1, :, :]
|
||||
if self.use_hw_tiling:
|
||||
tile = self._hw_tiled_encode(tile, return_dict)
|
||||
else:
|
||||
tile = self._encode(tile)
|
||||
if i > 0:
|
||||
tile = tile[:, :, 1:, :, :]
|
||||
row.append(tile)
|
||||
result_row = []
|
||||
for i, tile in enumerate(row):
|
||||
if i > 0:
|
||||
tile = self.blend_z(row[i - 1], tile, blend_extent)
|
||||
result_row.append(tile[:, :, :t_limit, :, :])
|
||||
else:
|
||||
result_row.append(tile[:, :, :t_limit + 1, :, :])
|
||||
|
||||
moments = torch.cat(result_row, dim=2)
|
||||
|
||||
|
||||
else:
|
||||
moments = (
|
||||
self._hw_tiled_encode(z, return_dict)
|
||||
if self.use_hw_tiling and z.shape[2] > 1
|
||||
else self._encode(z)
|
||||
)
|
||||
|
||||
posterior = DiagonalGaussianDistribution(moments)
|
||||
if not return_dict:
|
||||
return (posterior,)
|
||||
|
||||
return AutoencoderKLOutput(latent_dist=posterior)
|
||||
|
||||
def _normalize_latent_channels(self, z: torch.FloatTensor) -> torch.FloatTensor:
|
||||
if isinstance(self.latent_norm_out, nn.BatchNorm3d):
|
||||
_, c, _, _, _ = z.shape
|
||||
z = torch.cat(
|
||||
[
|
||||
self.latent_norm_out(z[:, : c // 2, :, :, :]),
|
||||
z[:, c // 2 :, :, :, :],
|
||||
],
|
||||
dim=1,
|
||||
)
|
||||
elif isinstance(self.latent_norm_out, nn.BatchNorm2d):
|
||||
raise NotImplementedError("BatchNorm2d not supported")
|
||||
return z
|
||||
|
||||
def _unnormalize_latent_channels(self, z: torch.FloatTensor) -> torch.FloatTensor:
|
||||
if isinstance(self.latent_norm_out, nn.BatchNorm3d):
|
||||
running_mean = self.latent_norm_out.running_mean.view(1, -1, 1, 1, 1)
|
||||
running_var = self.latent_norm_out.running_var.view(1, -1, 1, 1, 1)
|
||||
eps = self.latent_norm_out.eps
|
||||
|
||||
z = z * torch.sqrt(running_var + eps) + running_mean
|
||||
elif isinstance(self.latent_norm_out, nn.BatchNorm3d):
|
||||
raise NotImplementedError("BatchNorm2d not supported")
|
||||
return z
|
||||
|
||||
def _encode(self, x: torch.FloatTensor) -> AutoencoderKLOutput:
|
||||
h = self.encoder(x)
|
||||
moments = self.quant_conv(h)
|
||||
moments = self._normalize_latent_channels(moments)
|
||||
return moments
|
||||
|
||||
def _decode(
|
||||
self,
|
||||
z: torch.FloatTensor,
|
||||
target_shape=None,
|
||||
timestep: Optional[torch.Tensor] = None,
|
||||
) -> Union[DecoderOutput, torch.FloatTensor]:
|
||||
z = self._unnormalize_latent_channels(z)
|
||||
z = self.post_quant_conv(z)
|
||||
if "timestep" in self.decoder_params:
|
||||
dec = self.decoder(z, target_shape=target_shape, timestep=timestep)
|
||||
else:
|
||||
dec = self.decoder(z, target_shape=target_shape)
|
||||
return dec
|
||||
|
||||
def decode(
|
||||
self,
|
||||
z: torch.FloatTensor,
|
||||
return_dict: bool = True,
|
||||
target_shape=None,
|
||||
timestep: Optional[torch.Tensor] = None,
|
||||
) -> Union[DecoderOutput, torch.FloatTensor]:
|
||||
assert target_shape is not None, "target_shape must be provided for decoding"
|
||||
if self.use_z_tiling and z.shape[2] > (self.z_sample_size + 1) > 1:
|
||||
# Split z into overlapping tiles and decode them separately.
|
||||
tile_latent_min_tsize = self.z_sample_size
|
||||
tile_sample_min_tsize = tile_latent_min_tsize * 8
|
||||
tile_overlap_factor = 0.25
|
||||
|
||||
B, C, T, H, W = z.shape
|
||||
overlap_size = int(tile_latent_min_tsize * (1 - tile_overlap_factor))
|
||||
blend_extent = int(tile_sample_min_tsize * tile_overlap_factor)
|
||||
t_limit = tile_sample_min_tsize - blend_extent
|
||||
|
||||
row = []
|
||||
for i in range(0, T, overlap_size):
|
||||
tile = z[:, :, i: i + tile_latent_min_tsize + 1, :, :]
|
||||
target_shape_split = list(target_shape)
|
||||
target_shape_split[2] = tile.shape[2] * 8
|
||||
if self.use_hw_tiling:
|
||||
decoded = self._hw_tiled_decode(tile, target_shape, timestep)
|
||||
else:
|
||||
decoded = self._decode(tile, target_shape=target_shape, timestep=timestep)
|
||||
|
||||
if i > 0:
|
||||
decoded = decoded[:, :, 1:, :, :]
|
||||
row.append(decoded.to(torch.float16).cpu())
|
||||
decoded = None
|
||||
result_row = []
|
||||
for i, tile in enumerate(row):
|
||||
if i > 0:
|
||||
tile = self.blend_z(row[i - 1], tile, blend_extent)
|
||||
result_row.append(tile[:, :, :t_limit, :, :])
|
||||
else:
|
||||
result_row.append(tile[:, :, :t_limit + 1, :, :])
|
||||
|
||||
dec = torch.cat(result_row, dim=2)
|
||||
if not return_dict:
|
||||
return (dec,)
|
||||
|
||||
return DecoderOutput(sample=dec)
|
||||
else:
|
||||
decoded = (
|
||||
self._hw_tiled_decode(z, target_shape, timestep)
|
||||
if self.use_hw_tiling
|
||||
else self._decode(z, target_shape=target_shape, timestep=timestep)
|
||||
)
|
||||
|
||||
if not return_dict:
|
||||
return (decoded,)
|
||||
|
||||
return DecoderOutput(sample=decoded)
|
||||
|
||||
def forward(
|
||||
self,
|
||||
sample: torch.FloatTensor,
|
||||
sample_posterior: bool = False,
|
||||
return_dict: bool = True,
|
||||
generator: Optional[torch.Generator] = None,
|
||||
) -> Union[DecoderOutput, torch.FloatTensor]:
|
||||
r"""
|
||||
Args:
|
||||
sample (`torch.FloatTensor`): Input sample.
|
||||
sample_posterior (`bool`, *optional*, defaults to `False`):
|
||||
Whether to sample from the posterior.
|
||||
return_dict (`bool`, *optional*, defaults to `True`):
|
||||
Whether to return a [`DecoderOutput`] instead of a plain tuple.
|
||||
generator (`torch.Generator`, *optional*):
|
||||
Generator used to sample from the posterior.
|
||||
"""
|
||||
x = sample
|
||||
posterior = self.encode(x).latent_dist
|
||||
if sample_posterior:
|
||||
z = posterior.sample(generator=generator)
|
||||
else:
|
||||
z = posterior.mode()
|
||||
dec = self.decode(z, target_shape=sample.shape).sample
|
||||
|
||||
if not return_dict:
|
||||
return (dec,)
|
||||
|
||||
return DecoderOutput(sample=dec)
|
||||
247
ltx_video/models/autoencoders/vae_encode.py
Normal file
247
ltx_video/models/autoencoders/vae_encode.py
Normal file
@@ -0,0 +1,247 @@
|
||||
from typing import Tuple
|
||||
import torch
|
||||
from diffusers import AutoencoderKL
|
||||
from einops import rearrange
|
||||
from torch import Tensor
|
||||
|
||||
|
||||
from ltx_video.models.autoencoders.causal_video_autoencoder import (
|
||||
CausalVideoAutoencoder,
|
||||
)
|
||||
from ltx_video.models.autoencoders.video_autoencoder import (
|
||||
Downsample3D,
|
||||
VideoAutoencoder,
|
||||
)
|
||||
|
||||
try:
|
||||
import torch_xla.core.xla_model as xm
|
||||
except ImportError:
|
||||
xm = None
|
||||
|
||||
|
||||
def vae_encode(
|
||||
media_items: Tensor,
|
||||
vae: AutoencoderKL,
|
||||
split_size: int = 1,
|
||||
vae_per_channel_normalize=False,
|
||||
) -> Tensor:
|
||||
"""
|
||||
Encodes media items (images or videos) into latent representations using a specified VAE model.
|
||||
The function supports processing batches of images or video frames and can handle the processing
|
||||
in smaller sub-batches if needed.
|
||||
|
||||
Args:
|
||||
media_items (Tensor): A torch Tensor containing the media items to encode. The expected
|
||||
shape is (batch_size, channels, height, width) for images or (batch_size, channels,
|
||||
frames, height, width) for videos.
|
||||
vae (AutoencoderKL): An instance of the `AutoencoderKL` class from the `diffusers` library,
|
||||
pre-configured and loaded with the appropriate model weights.
|
||||
split_size (int, optional): The number of sub-batches to split the input batch into for encoding.
|
||||
If set to more than 1, the input media items are processed in smaller batches according to
|
||||
this value. Defaults to 1, which processes all items in a single batch.
|
||||
|
||||
Returns:
|
||||
Tensor: A torch Tensor of the encoded latent representations. The shape of the tensor is adjusted
|
||||
to match the input shape, scaled by the model's configuration.
|
||||
|
||||
Examples:
|
||||
>>> import torch
|
||||
>>> from diffusers import AutoencoderKL
|
||||
>>> vae = AutoencoderKL.from_pretrained('your-model-name')
|
||||
>>> images = torch.rand(10, 3, 8 256, 256) # Example tensor with 10 videos of 8 frames.
|
||||
>>> latents = vae_encode(images, vae)
|
||||
>>> print(latents.shape) # Output shape will depend on the model's latent configuration.
|
||||
|
||||
Note:
|
||||
In case of a video, the function encodes the media item frame-by frame.
|
||||
"""
|
||||
is_video_shaped = media_items.dim() == 5
|
||||
batch_size, channels = media_items.shape[0:2]
|
||||
|
||||
if channels != 3:
|
||||
raise ValueError(f"Expects tensors with 3 channels, got {channels}.")
|
||||
|
||||
if is_video_shaped and not isinstance(
|
||||
vae, (VideoAutoencoder, CausalVideoAutoencoder)
|
||||
):
|
||||
media_items = rearrange(media_items, "b c n h w -> (b n) c h w")
|
||||
if split_size > 1:
|
||||
if len(media_items) % split_size != 0:
|
||||
raise ValueError(
|
||||
"Error: The batch size must be divisible by 'train.vae_bs_split"
|
||||
)
|
||||
encode_bs = len(media_items) // split_size
|
||||
# latents = [vae.encode(image_batch).latent_dist.sample() for image_batch in media_items.split(encode_bs)]
|
||||
latents = []
|
||||
if media_items.device.type == "xla":
|
||||
xm.mark_step()
|
||||
for image_batch in media_items.split(encode_bs):
|
||||
latents.append(vae.encode(image_batch).latent_dist.sample())
|
||||
if media_items.device.type == "xla":
|
||||
xm.mark_step()
|
||||
latents = torch.cat(latents, dim=0)
|
||||
else:
|
||||
latents = vae.encode(media_items).latent_dist.sample()
|
||||
|
||||
latents = normalize_latents(latents, vae, vae_per_channel_normalize)
|
||||
if is_video_shaped and not isinstance(
|
||||
vae, (VideoAutoencoder, CausalVideoAutoencoder)
|
||||
):
|
||||
latents = rearrange(latents, "(b n) c h w -> b c n h w", b=batch_size)
|
||||
return latents
|
||||
|
||||
|
||||
def vae_decode(
|
||||
latents: Tensor,
|
||||
vae: AutoencoderKL,
|
||||
is_video: bool = True,
|
||||
split_size: int = 1,
|
||||
vae_per_channel_normalize=False,
|
||||
timestep=None,
|
||||
) -> Tensor:
|
||||
is_video_shaped = latents.dim() == 5
|
||||
batch_size = latents.shape[0]
|
||||
|
||||
if is_video_shaped and not isinstance(
|
||||
vae, (VideoAutoencoder, CausalVideoAutoencoder)
|
||||
):
|
||||
latents = rearrange(latents, "b c n h w -> (b n) c h w")
|
||||
if split_size > 1:
|
||||
if len(latents) % split_size != 0:
|
||||
raise ValueError(
|
||||
"Error: The batch size must be divisible by 'train.vae_bs_split"
|
||||
)
|
||||
encode_bs = len(latents) // split_size
|
||||
image_batch = [
|
||||
_run_decoder(
|
||||
latent_batch, vae, is_video, vae_per_channel_normalize, timestep
|
||||
)
|
||||
for latent_batch in latents.split(encode_bs)
|
||||
]
|
||||
images = torch.cat(image_batch, dim=0)
|
||||
else:
|
||||
images = _run_decoder(
|
||||
latents, vae, is_video, vae_per_channel_normalize, timestep
|
||||
)
|
||||
|
||||
if is_video_shaped and not isinstance(
|
||||
vae, (VideoAutoencoder, CausalVideoAutoencoder)
|
||||
):
|
||||
images = rearrange(images, "(b n) c h w -> b c n h w", b=batch_size)
|
||||
return images
|
||||
|
||||
|
||||
def _run_decoder(
|
||||
latents: Tensor,
|
||||
vae: AutoencoderKL,
|
||||
is_video: bool,
|
||||
vae_per_channel_normalize=False,
|
||||
timestep=None,
|
||||
) -> Tensor:
|
||||
if isinstance(vae, (VideoAutoencoder, CausalVideoAutoencoder)):
|
||||
*_, fl, hl, wl = latents.shape
|
||||
temporal_scale, spatial_scale, _ = get_vae_size_scale_factor(vae)
|
||||
latents = latents.to(vae.dtype)
|
||||
vae_decode_kwargs = {}
|
||||
if timestep is not None:
|
||||
vae_decode_kwargs["timestep"] = timestep
|
||||
image = vae.decode(
|
||||
un_normalize_latents(latents, vae, vae_per_channel_normalize),
|
||||
return_dict=False,
|
||||
target_shape=(
|
||||
1,
|
||||
3,
|
||||
fl * temporal_scale if is_video else 1,
|
||||
hl * spatial_scale,
|
||||
wl * spatial_scale,
|
||||
),
|
||||
**vae_decode_kwargs,
|
||||
)[0]
|
||||
else:
|
||||
image = vae.decode(
|
||||
un_normalize_latents(latents, vae, vae_per_channel_normalize),
|
||||
return_dict=False,
|
||||
)[0]
|
||||
return image
|
||||
|
||||
|
||||
def get_vae_size_scale_factor(vae: AutoencoderKL) -> float:
|
||||
if isinstance(vae, CausalVideoAutoencoder):
|
||||
spatial = vae.spatial_downscale_factor
|
||||
temporal = vae.temporal_downscale_factor
|
||||
else:
|
||||
down_blocks = len(
|
||||
[
|
||||
block
|
||||
for block in vae.encoder.down_blocks
|
||||
if isinstance(block.downsample, Downsample3D)
|
||||
]
|
||||
)
|
||||
spatial = vae.config.patch_size * 2**down_blocks
|
||||
temporal = (
|
||||
vae.config.patch_size_t * 2**down_blocks
|
||||
if isinstance(vae, VideoAutoencoder)
|
||||
else 1
|
||||
)
|
||||
|
||||
return (temporal, spatial, spatial)
|
||||
|
||||
|
||||
def latent_to_pixel_coords(
|
||||
latent_coords: Tensor, vae: AutoencoderKL, causal_fix: bool = False
|
||||
) -> Tensor:
|
||||
"""
|
||||
Converts latent coordinates to pixel coordinates by scaling them according to the VAE's
|
||||
configuration.
|
||||
|
||||
Args:
|
||||
latent_coords (Tensor): A tensor of shape [batch_size, 3, num_latents]
|
||||
containing the latent corner coordinates of each token.
|
||||
vae (AutoencoderKL): The VAE model
|
||||
causal_fix (bool): Whether to take into account the different temporal scale
|
||||
of the first frame. Default = False for backwards compatibility.
|
||||
Returns:
|
||||
Tensor: A tensor of pixel coordinates corresponding to the input latent coordinates.
|
||||
"""
|
||||
|
||||
scale_factors = get_vae_size_scale_factor(vae)
|
||||
causal_fix = isinstance(vae, CausalVideoAutoencoder) and causal_fix
|
||||
pixel_coords = latent_to_pixel_coords_from_factors(
|
||||
latent_coords, scale_factors, causal_fix
|
||||
)
|
||||
return pixel_coords
|
||||
|
||||
|
||||
def latent_to_pixel_coords_from_factors(
|
||||
latent_coords: Tensor, scale_factors: Tuple, causal_fix: bool = False
|
||||
) -> Tensor:
|
||||
pixel_coords = (
|
||||
latent_coords
|
||||
* torch.tensor(scale_factors, device=latent_coords.device)[None, :, None]
|
||||
)
|
||||
if causal_fix:
|
||||
# Fix temporal scale for first frame to 1 due to causality
|
||||
pixel_coords[:, 0] = (pixel_coords[:, 0] + 1 - scale_factors[0]).clamp(min=0)
|
||||
return pixel_coords
|
||||
|
||||
|
||||
def normalize_latents(
|
||||
latents: Tensor, vae: AutoencoderKL, vae_per_channel_normalize: bool = False
|
||||
) -> Tensor:
|
||||
return (
|
||||
(latents - vae.mean_of_means.to(latents.dtype).to(latents.device).view(1, -1, 1, 1, 1))
|
||||
/ vae.std_of_means.to(latents.dtype).to(latents.device).view(1, -1, 1, 1, 1)
|
||||
if vae_per_channel_normalize
|
||||
else latents * vae.config.scaling_factor
|
||||
)
|
||||
|
||||
|
||||
def un_normalize_latents(
|
||||
latents: Tensor, vae: AutoencoderKL, vae_per_channel_normalize: bool = False
|
||||
) -> Tensor:
|
||||
return (
|
||||
latents * vae.std_of_means.to(latents.dtype).to(latents.device).view(1, -1, 1, 1, 1)
|
||||
+ vae.mean_of_means.to(latents.dtype).to(latents.device).view(1, -1, 1, 1, 1)
|
||||
if vae_per_channel_normalize
|
||||
else latents / vae.config.scaling_factor
|
||||
)
|
||||
1045
ltx_video/models/autoencoders/video_autoencoder.py
Normal file
1045
ltx_video/models/autoencoders/video_autoencoder.py
Normal file
File diff suppressed because it is too large
Load Diff
0
ltx_video/models/transformers/__init__.py
Normal file
0
ltx_video/models/transformers/__init__.py
Normal file
1323
ltx_video/models/transformers/attention.py
Normal file
1323
ltx_video/models/transformers/attention.py
Normal file
File diff suppressed because it is too large
Load Diff
129
ltx_video/models/transformers/embeddings.py
Normal file
129
ltx_video/models/transformers/embeddings.py
Normal file
@@ -0,0 +1,129 @@
|
||||
# Adapted from: https://github.com/huggingface/diffusers/blob/main/src/diffusers/models/embeddings.py
|
||||
import math
|
||||
|
||||
import numpy as np
|
||||
import torch
|
||||
from einops import rearrange
|
||||
from torch import nn
|
||||
|
||||
|
||||
def get_timestep_embedding(
|
||||
timesteps: torch.Tensor,
|
||||
embedding_dim: int,
|
||||
flip_sin_to_cos: bool = False,
|
||||
downscale_freq_shift: float = 1,
|
||||
scale: float = 1,
|
||||
max_period: int = 10000,
|
||||
):
|
||||
"""
|
||||
This matches the implementation in Denoising Diffusion Probabilistic Models: Create sinusoidal timestep embeddings.
|
||||
|
||||
:param timesteps: a 1-D Tensor of N indices, one per batch element.
|
||||
These may be fractional.
|
||||
:param embedding_dim: the dimension of the output. :param max_period: controls the minimum frequency of the
|
||||
embeddings. :return: an [N x dim] Tensor of positional embeddings.
|
||||
"""
|
||||
assert len(timesteps.shape) == 1, "Timesteps should be a 1d-array"
|
||||
|
||||
half_dim = embedding_dim // 2
|
||||
exponent = -math.log(max_period) * torch.arange(
|
||||
start=0, end=half_dim, dtype=torch.float32, device=timesteps.device
|
||||
)
|
||||
exponent = exponent / (half_dim - downscale_freq_shift)
|
||||
|
||||
emb = torch.exp(exponent)
|
||||
emb = timesteps[:, None].float() * emb[None, :]
|
||||
|
||||
# scale embeddings
|
||||
emb = scale * emb
|
||||
|
||||
# concat sine and cosine embeddings
|
||||
emb = torch.cat([torch.sin(emb), torch.cos(emb)], dim=-1)
|
||||
|
||||
# flip sine and cosine embeddings
|
||||
if flip_sin_to_cos:
|
||||
emb = torch.cat([emb[:, half_dim:], emb[:, :half_dim]], dim=-1)
|
||||
|
||||
# zero pad
|
||||
if embedding_dim % 2 == 1:
|
||||
emb = torch.nn.functional.pad(emb, (0, 1, 0, 0))
|
||||
return emb
|
||||
|
||||
|
||||
def get_3d_sincos_pos_embed(embed_dim, grid, w, h, f):
|
||||
"""
|
||||
grid_size: int of the grid height and width return: pos_embed: [grid_size*grid_size, embed_dim] or
|
||||
[1+grid_size*grid_size, embed_dim] (w/ or w/o cls_token)
|
||||
"""
|
||||
grid = rearrange(grid, "c (f h w) -> c f h w", h=h, w=w)
|
||||
grid = rearrange(grid, "c f h w -> c h w f", h=h, w=w)
|
||||
grid = grid.reshape([3, 1, w, h, f])
|
||||
pos_embed = get_3d_sincos_pos_embed_from_grid(embed_dim, grid)
|
||||
pos_embed = pos_embed.transpose(1, 0, 2, 3)
|
||||
return rearrange(pos_embed, "h w f c -> (f h w) c")
|
||||
|
||||
|
||||
def get_3d_sincos_pos_embed_from_grid(embed_dim, grid):
|
||||
if embed_dim % 3 != 0:
|
||||
raise ValueError("embed_dim must be divisible by 3")
|
||||
|
||||
# use half of dimensions to encode grid_h
|
||||
emb_f = get_1d_sincos_pos_embed_from_grid(embed_dim // 3, grid[0]) # (H*W*T, D/3)
|
||||
emb_h = get_1d_sincos_pos_embed_from_grid(embed_dim // 3, grid[1]) # (H*W*T, D/3)
|
||||
emb_w = get_1d_sincos_pos_embed_from_grid(embed_dim // 3, grid[2]) # (H*W*T, D/3)
|
||||
|
||||
emb = np.concatenate([emb_h, emb_w, emb_f], axis=-1) # (H*W*T, D)
|
||||
return emb
|
||||
|
||||
|
||||
def get_1d_sincos_pos_embed_from_grid(embed_dim, pos):
|
||||
"""
|
||||
embed_dim: output dimension for each position pos: a list of positions to be encoded: size (M,) out: (M, D)
|
||||
"""
|
||||
if embed_dim % 2 != 0:
|
||||
raise ValueError("embed_dim must be divisible by 2")
|
||||
|
||||
omega = np.arange(embed_dim // 2, dtype=np.float64)
|
||||
omega /= embed_dim / 2.0
|
||||
omega = 1.0 / 10000**omega # (D/2,)
|
||||
|
||||
pos_shape = pos.shape
|
||||
|
||||
pos = pos.reshape(-1)
|
||||
out = np.einsum("m,d->md", pos, omega) # (M, D/2), outer product
|
||||
out = out.reshape([*pos_shape, -1])[0]
|
||||
|
||||
emb_sin = np.sin(out) # (M, D/2)
|
||||
emb_cos = np.cos(out) # (M, D/2)
|
||||
|
||||
emb = np.concatenate([emb_sin, emb_cos], axis=-1) # (M, D)
|
||||
return emb
|
||||
|
||||
|
||||
class SinusoidalPositionalEmbedding(nn.Module):
|
||||
"""Apply positional information to a sequence of embeddings.
|
||||
|
||||
Takes in a sequence of embeddings with shape (batch_size, seq_length, embed_dim) and adds positional embeddings to
|
||||
them
|
||||
|
||||
Args:
|
||||
embed_dim: (int): Dimension of the positional embedding.
|
||||
max_seq_length: Maximum sequence length to apply positional embeddings
|
||||
|
||||
"""
|
||||
|
||||
def __init__(self, embed_dim: int, max_seq_length: int = 32):
|
||||
super().__init__()
|
||||
position = torch.arange(max_seq_length).unsqueeze(1)
|
||||
div_term = torch.exp(
|
||||
torch.arange(0, embed_dim, 2) * (-math.log(10000.0) / embed_dim)
|
||||
)
|
||||
pe = torch.zeros(1, max_seq_length, embed_dim)
|
||||
pe[0, :, 0::2] = torch.sin(position * div_term)
|
||||
pe[0, :, 1::2] = torch.cos(position * div_term)
|
||||
self.register_buffer("pe", pe)
|
||||
|
||||
def forward(self, x):
|
||||
_, seq_length, _ = x.shape
|
||||
x = x + self.pe[:, :seq_length]
|
||||
return x
|
||||
84
ltx_video/models/transformers/symmetric_patchifier.py
Normal file
84
ltx_video/models/transformers/symmetric_patchifier.py
Normal file
@@ -0,0 +1,84 @@
|
||||
from abc import ABC, abstractmethod
|
||||
from typing import Tuple
|
||||
|
||||
import torch
|
||||
from diffusers.configuration_utils import ConfigMixin
|
||||
from einops import rearrange
|
||||
from torch import Tensor
|
||||
|
||||
|
||||
class Patchifier(ConfigMixin, ABC):
|
||||
def __init__(self, patch_size: int):
|
||||
super().__init__()
|
||||
self._patch_size = (1, patch_size, patch_size)
|
||||
|
||||
@abstractmethod
|
||||
def patchify(self, latents: Tensor) -> Tuple[Tensor, Tensor]:
|
||||
raise NotImplementedError("Patchify method not implemented")
|
||||
|
||||
@abstractmethod
|
||||
def unpatchify(
|
||||
self,
|
||||
latents: Tensor,
|
||||
output_height: int,
|
||||
output_width: int,
|
||||
out_channels: int,
|
||||
) -> Tuple[Tensor, Tensor]:
|
||||
pass
|
||||
|
||||
@property
|
||||
def patch_size(self):
|
||||
return self._patch_size
|
||||
|
||||
def get_latent_coords(
|
||||
self, latent_num_frames, latent_height, latent_width, batch_size, device
|
||||
):
|
||||
"""
|
||||
Return a tensor of shape [batch_size, 3, num_patches] containing the
|
||||
top-left corner latent coordinates of each latent patch.
|
||||
The tensor is repeated for each batch element.
|
||||
"""
|
||||
latent_sample_coords = torch.meshgrid(
|
||||
torch.arange(0, latent_num_frames, self._patch_size[0], device=device),
|
||||
torch.arange(0, latent_height, self._patch_size[1], device=device),
|
||||
torch.arange(0, latent_width, self._patch_size[2], device=device),
|
||||
)
|
||||
latent_sample_coords = torch.stack(latent_sample_coords, dim=0)
|
||||
latent_coords = latent_sample_coords.unsqueeze(0).repeat(batch_size, 1, 1, 1, 1)
|
||||
latent_coords = rearrange(
|
||||
latent_coords, "b c f h w -> b c (f h w)", b=batch_size
|
||||
)
|
||||
return latent_coords
|
||||
|
||||
|
||||
class SymmetricPatchifier(Patchifier):
|
||||
def patchify(self, latents: Tensor) -> Tuple[Tensor, Tensor]:
|
||||
b, _, f, h, w = latents.shape
|
||||
latent_coords = self.get_latent_coords(f, h, w, b, latents.device)
|
||||
latents = rearrange(
|
||||
latents,
|
||||
"b c (f p1) (h p2) (w p3) -> b (f h w) (c p1 p2 p3)",
|
||||
p1=self._patch_size[0],
|
||||
p2=self._patch_size[1],
|
||||
p3=self._patch_size[2],
|
||||
)
|
||||
return latents, latent_coords
|
||||
|
||||
def unpatchify(
|
||||
self,
|
||||
latents: Tensor,
|
||||
output_height: int,
|
||||
output_width: int,
|
||||
out_channels: int,
|
||||
) -> Tuple[Tensor, Tensor]:
|
||||
output_height = output_height // self._patch_size[1]
|
||||
output_width = output_width // self._patch_size[2]
|
||||
latents = rearrange(
|
||||
latents,
|
||||
"b (f h w) (c p q) -> b c f (h p) (w q)",
|
||||
h=output_height,
|
||||
w=output_width,
|
||||
p=self._patch_size[1],
|
||||
q=self._patch_size[2],
|
||||
)
|
||||
return latents
|
||||
507
ltx_video/models/transformers/transformer3d.py
Normal file
507
ltx_video/models/transformers/transformer3d.py
Normal file
@@ -0,0 +1,507 @@
|
||||
# Adapted from: https://github.com/huggingface/diffusers/blob/v0.26.3/src/diffusers/models/transformers/transformer_2d.py
|
||||
import math
|
||||
from dataclasses import dataclass
|
||||
from typing import Any, Dict, List, Optional, Union
|
||||
import os
|
||||
import json
|
||||
import glob
|
||||
from pathlib import Path
|
||||
|
||||
import torch
|
||||
from diffusers.configuration_utils import ConfigMixin, register_to_config
|
||||
from diffusers.models.embeddings import PixArtAlphaTextProjection
|
||||
from diffusers.models.modeling_utils import ModelMixin
|
||||
from diffusers.models.normalization import AdaLayerNormSingle
|
||||
from diffusers.utils import BaseOutput, is_torch_version
|
||||
from diffusers.utils import logging
|
||||
from torch import nn
|
||||
from safetensors import safe_open
|
||||
from ltx_video.models.transformers.attention import BasicTransformerBlock, reshape_hidden_states, restore_hidden_states_shape
|
||||
from ltx_video.utils.skip_layer_strategy import SkipLayerStrategy
|
||||
|
||||
from ltx_video.utils.diffusers_config_mapping import (
|
||||
diffusers_and_ours_config_mapping,
|
||||
make_hashable_key,
|
||||
TRANSFORMER_KEYS_RENAME_DICT,
|
||||
)
|
||||
|
||||
|
||||
logger = logging.get_logger(__name__)
|
||||
|
||||
|
||||
@dataclass
|
||||
class Transformer3DModelOutput(BaseOutput):
|
||||
"""
|
||||
The output of [`Transformer2DModel`].
|
||||
|
||||
Args:
|
||||
sample (`torch.FloatTensor` of shape `(batch_size, num_channels, height, width)` or `(batch size, num_vector_embeds - 1, num_latent_pixels)` if [`Transformer2DModel`] is discrete):
|
||||
The hidden states output conditioned on the `encoder_hidden_states` input. If discrete, returns probability
|
||||
distributions for the unnoised latent pixels.
|
||||
"""
|
||||
|
||||
sample: torch.FloatTensor
|
||||
|
||||
|
||||
class Transformer3DModel(ModelMixin, ConfigMixin):
|
||||
_supports_gradient_checkpointing = True
|
||||
|
||||
@register_to_config
|
||||
def __init__(
|
||||
self,
|
||||
num_attention_heads: int = 16,
|
||||
attention_head_dim: int = 88,
|
||||
in_channels: Optional[int] = None,
|
||||
out_channels: Optional[int] = None,
|
||||
num_layers: int = 1,
|
||||
dropout: float = 0.0,
|
||||
norm_num_groups: int = 32,
|
||||
cross_attention_dim: Optional[int] = None,
|
||||
attention_bias: bool = False,
|
||||
num_vector_embeds: Optional[int] = None,
|
||||
activation_fn: str = "geglu",
|
||||
num_embeds_ada_norm: Optional[int] = None,
|
||||
use_linear_projection: bool = False,
|
||||
only_cross_attention: bool = False,
|
||||
double_self_attention: bool = False,
|
||||
upcast_attention: bool = False,
|
||||
adaptive_norm: str = "single_scale_shift", # 'single_scale_shift' or 'single_scale'
|
||||
standardization_norm: str = "layer_norm", # 'layer_norm' or 'rms_norm'
|
||||
norm_elementwise_affine: bool = True,
|
||||
norm_eps: float = 1e-5,
|
||||
attention_type: str = "default",
|
||||
caption_channels: int = None,
|
||||
use_tpu_flash_attention: bool = False, # if True uses the TPU attention offload ('flash attention')
|
||||
qk_norm: Optional[str] = None,
|
||||
positional_embedding_type: str = "rope",
|
||||
positional_embedding_theta: Optional[float] = None,
|
||||
positional_embedding_max_pos: Optional[List[int]] = None,
|
||||
timestep_scale_multiplier: Optional[float] = None,
|
||||
causal_temporal_positioning: bool = False, # For backward compatibility, will be deprecated
|
||||
):
|
||||
super().__init__()
|
||||
self.use_tpu_flash_attention = (
|
||||
use_tpu_flash_attention # FIXME: push config down to the attention modules
|
||||
)
|
||||
self.use_linear_projection = use_linear_projection
|
||||
self.num_attention_heads = num_attention_heads
|
||||
self.attention_head_dim = attention_head_dim
|
||||
inner_dim = num_attention_heads * attention_head_dim
|
||||
self.inner_dim = inner_dim
|
||||
self.patchify_proj = nn.Linear(in_channels, inner_dim, bias=True)
|
||||
self.positional_embedding_type = positional_embedding_type
|
||||
self.positional_embedding_theta = positional_embedding_theta
|
||||
self.positional_embedding_max_pos = positional_embedding_max_pos
|
||||
self.use_rope = self.positional_embedding_type == "rope"
|
||||
self.timestep_scale_multiplier = timestep_scale_multiplier
|
||||
|
||||
if self.positional_embedding_type == "absolute":
|
||||
raise ValueError("Absolute positional embedding is no longer supported")
|
||||
elif self.positional_embedding_type == "rope":
|
||||
if positional_embedding_theta is None:
|
||||
raise ValueError(
|
||||
"If `positional_embedding_type` type is rope, `positional_embedding_theta` must also be defined"
|
||||
)
|
||||
if positional_embedding_max_pos is None:
|
||||
raise ValueError(
|
||||
"If `positional_embedding_type` type is rope, `positional_embedding_max_pos` must also be defined"
|
||||
)
|
||||
|
||||
# 3. Define transformers blocks
|
||||
self.transformer_blocks = nn.ModuleList(
|
||||
[
|
||||
BasicTransformerBlock(
|
||||
inner_dim,
|
||||
num_attention_heads,
|
||||
attention_head_dim,
|
||||
dropout=dropout,
|
||||
cross_attention_dim=cross_attention_dim,
|
||||
activation_fn=activation_fn,
|
||||
num_embeds_ada_norm=num_embeds_ada_norm,
|
||||
attention_bias=attention_bias,
|
||||
only_cross_attention=only_cross_attention,
|
||||
double_self_attention=double_self_attention,
|
||||
upcast_attention=upcast_attention,
|
||||
adaptive_norm=adaptive_norm,
|
||||
standardization_norm=standardization_norm,
|
||||
norm_elementwise_affine=norm_elementwise_affine,
|
||||
norm_eps=norm_eps,
|
||||
attention_type=attention_type,
|
||||
use_tpu_flash_attention=use_tpu_flash_attention,
|
||||
qk_norm=qk_norm,
|
||||
use_rope=self.use_rope,
|
||||
)
|
||||
for d in range(num_layers)
|
||||
]
|
||||
)
|
||||
|
||||
# 4. Define output layers
|
||||
self.out_channels = in_channels if out_channels is None else out_channels
|
||||
self.norm_out = nn.LayerNorm(inner_dim, elementwise_affine=False, eps=1e-6)
|
||||
self.scale_shift_table = nn.Parameter(
|
||||
torch.randn(2, inner_dim) / inner_dim**0.5
|
||||
)
|
||||
self.proj_out = nn.Linear(inner_dim, self.out_channels)
|
||||
|
||||
self.adaln_single = AdaLayerNormSingle(
|
||||
inner_dim, use_additional_conditions=False
|
||||
)
|
||||
if adaptive_norm == "single_scale":
|
||||
self.adaln_single.linear = nn.Linear(inner_dim, 4 * inner_dim, bias=True)
|
||||
|
||||
self.caption_projection = None
|
||||
if caption_channels is not None:
|
||||
self.caption_projection = PixArtAlphaTextProjection(
|
||||
in_features=caption_channels, hidden_size=inner_dim
|
||||
)
|
||||
|
||||
self.gradient_checkpointing = False
|
||||
|
||||
def set_use_tpu_flash_attention(self):
|
||||
r"""
|
||||
Function sets the flag in this object and propagates down the children. The flag will enforce the usage of TPU
|
||||
attention kernel.
|
||||
"""
|
||||
logger.info("ENABLE TPU FLASH ATTENTION -> TRUE")
|
||||
self.use_tpu_flash_attention = True
|
||||
# push config down to the attention modules
|
||||
for block in self.transformer_blocks:
|
||||
block.set_use_tpu_flash_attention()
|
||||
|
||||
def create_skip_layer_mask(
|
||||
self,
|
||||
batch_size: int,
|
||||
num_conds: int,
|
||||
ptb_index: int,
|
||||
skip_block_list: Optional[List[int]] = None,
|
||||
):
|
||||
if skip_block_list is None or len(skip_block_list) == 0:
|
||||
return None
|
||||
num_layers = len(self.transformer_blocks)
|
||||
mask = torch.ones(
|
||||
(num_layers, batch_size * num_conds), device=self.device, dtype=self.dtype
|
||||
)
|
||||
for block_idx in skip_block_list:
|
||||
mask[block_idx, ptb_index::num_conds] = 0
|
||||
return mask
|
||||
|
||||
def _set_gradient_checkpointing(self, module, value=False):
|
||||
if hasattr(module, "gradient_checkpointing"):
|
||||
module.gradient_checkpointing = value
|
||||
|
||||
def get_fractional_positions(self, indices_grid):
|
||||
fractional_positions = torch.stack(
|
||||
[
|
||||
indices_grid[:, i] / self.positional_embedding_max_pos[i]
|
||||
for i in range(3)
|
||||
],
|
||||
dim=-1,
|
||||
)
|
||||
return fractional_positions
|
||||
|
||||
def precompute_freqs_cis(self, indices_grid, spacing="exp"):
|
||||
dtype = torch.float32 # We need full precision in the freqs_cis computation.
|
||||
dim = self.inner_dim
|
||||
theta = self.positional_embedding_theta
|
||||
|
||||
fractional_positions = self.get_fractional_positions(indices_grid)
|
||||
|
||||
start = 1
|
||||
end = theta
|
||||
device = fractional_positions.device
|
||||
if spacing == "exp":
|
||||
indices = theta ** (
|
||||
torch.linspace(
|
||||
math.log(start, theta),
|
||||
math.log(end, theta),
|
||||
dim // 6,
|
||||
device=device,
|
||||
dtype=dtype,
|
||||
)
|
||||
)
|
||||
indices = indices.to(dtype=dtype)
|
||||
elif spacing == "exp_2":
|
||||
indices = 1.0 / theta ** (torch.arange(0, dim, 6, device=device) / dim)
|
||||
indices = indices.to(dtype=dtype)
|
||||
elif spacing == "linear":
|
||||
indices = torch.linspace(start, end, dim // 6, device=device, dtype=dtype)
|
||||
elif spacing == "sqrt":
|
||||
indices = torch.linspace(
|
||||
start**2, end**2, dim // 6, device=device, dtype=dtype
|
||||
).sqrt()
|
||||
|
||||
indices = indices * math.pi / 2
|
||||
|
||||
if spacing == "exp_2":
|
||||
freqs = (
|
||||
(indices * fractional_positions.unsqueeze(-1))
|
||||
.transpose(-1, -2)
|
||||
.flatten(2)
|
||||
)
|
||||
else:
|
||||
freqs = (
|
||||
(indices * (fractional_positions.unsqueeze(-1) * 2 - 1))
|
||||
.transpose(-1, -2)
|
||||
.flatten(2)
|
||||
)
|
||||
|
||||
cos_freq = freqs.cos().repeat_interleave(2, dim=-1)
|
||||
sin_freq = freqs.sin().repeat_interleave(2, dim=-1)
|
||||
if dim % 6 != 0:
|
||||
cos_padding = torch.ones_like(cos_freq[:, :, : dim % 6])
|
||||
sin_padding = torch.zeros_like(cos_freq[:, :, : dim % 6])
|
||||
cos_freq = torch.cat([cos_padding, cos_freq], dim=-1)
|
||||
sin_freq = torch.cat([sin_padding, sin_freq], dim=-1)
|
||||
return cos_freq.to(self.dtype), sin_freq.to(self.dtype)
|
||||
|
||||
def load_state_dict(
|
||||
self,
|
||||
state_dict: Dict,
|
||||
*args,
|
||||
**kwargs,
|
||||
):
|
||||
if any([key.startswith("model.diffusion_model.") for key in state_dict.keys()]):
|
||||
state_dict = {
|
||||
key.replace("model.diffusion_model.", ""): value
|
||||
for key, value in state_dict.items()
|
||||
if key.startswith("model.diffusion_model.")
|
||||
}
|
||||
return super().load_state_dict(state_dict, **kwargs)
|
||||
|
||||
@classmethod
|
||||
def from_pretrained(
|
||||
cls,
|
||||
pretrained_model_path: Optional[Union[str, os.PathLike]],
|
||||
*args,
|
||||
**kwargs,
|
||||
):
|
||||
pretrained_model_path = Path(pretrained_model_path)
|
||||
if pretrained_model_path.is_dir():
|
||||
config_path = pretrained_model_path / "transformer" / "config.json"
|
||||
with open(config_path, "r") as f:
|
||||
config = make_hashable_key(json.load(f))
|
||||
|
||||
assert config in diffusers_and_ours_config_mapping, (
|
||||
"Provided diffusers checkpoint config for transformer is not suppported. "
|
||||
"We only support diffusers configs found in Lightricks/LTX-Video."
|
||||
)
|
||||
|
||||
config = diffusers_and_ours_config_mapping[config]
|
||||
state_dict = {}
|
||||
ckpt_paths = (
|
||||
pretrained_model_path
|
||||
/ "transformer"
|
||||
/ "diffusion_pytorch_model*.safetensors"
|
||||
)
|
||||
dict_list = glob.glob(str(ckpt_paths))
|
||||
for dict_path in dict_list:
|
||||
part_dict = {}
|
||||
with safe_open(dict_path, framework="pt", device="cpu") as f:
|
||||
for k in f.keys():
|
||||
part_dict[k] = f.get_tensor(k)
|
||||
state_dict.update(part_dict)
|
||||
|
||||
for key in list(state_dict.keys()):
|
||||
new_key = key
|
||||
for replace_key, rename_key in TRANSFORMER_KEYS_RENAME_DICT.items():
|
||||
new_key = new_key.replace(replace_key, rename_key)
|
||||
state_dict[new_key] = state_dict.pop(key)
|
||||
|
||||
with torch.device("meta"):
|
||||
transformer = cls.from_config(config)
|
||||
transformer.load_state_dict(state_dict, assign=True, strict=True)
|
||||
elif pretrained_model_path.is_file() and str(pretrained_model_path).endswith(
|
||||
".safetensors"
|
||||
):
|
||||
comfy_single_file_state_dict = {}
|
||||
with safe_open(pretrained_model_path, framework="pt", device="cpu") as f:
|
||||
metadata = f.metadata()
|
||||
for k in f.keys():
|
||||
comfy_single_file_state_dict[k] = f.get_tensor(k)
|
||||
configs = json.loads(metadata["config"])
|
||||
transformer_config = configs["transformer"]
|
||||
with torch.device("meta"):
|
||||
transformer = Transformer3DModel.from_config(transformer_config)
|
||||
transformer.load_state_dict(comfy_single_file_state_dict, assign=True)
|
||||
return transformer
|
||||
|
||||
def forward(
|
||||
self,
|
||||
hidden_states: torch.Tensor,
|
||||
freqs_cis: list,
|
||||
encoder_hidden_states: Optional[torch.Tensor] = None,
|
||||
timestep: Optional[torch.LongTensor] = None,
|
||||
class_labels: Optional[torch.LongTensor] = None,
|
||||
cross_attention_kwargs: Dict[str, Any] = None,
|
||||
attention_mask: Optional[torch.Tensor] = None,
|
||||
encoder_attention_mask: Optional[torch.Tensor] = None,
|
||||
skip_layer_mask: Optional[torch.Tensor] = None,
|
||||
skip_layer_strategy: Optional[SkipLayerStrategy] = None,
|
||||
latent_shape = None,
|
||||
joint_pass = True,
|
||||
ltxv_model = None,
|
||||
mixed = False,
|
||||
return_dict: bool = True,
|
||||
):
|
||||
"""
|
||||
The [`Transformer2DModel`] forward method.
|
||||
|
||||
Args:
|
||||
hidden_states (`torch.LongTensor` of shape `(batch size, num latent pixels)` if discrete, `torch.FloatTensor` of shape `(batch size, channel, height, width)` if continuous):
|
||||
Input `hidden_states`.
|
||||
indices_grid (`torch.LongTensor` of shape `(batch size, 3, num latent pixels)`):
|
||||
encoder_hidden_states ( `torch.FloatTensor` of shape `(batch size, sequence len, embed dims)`, *optional*):
|
||||
Conditional embeddings for cross attention layer. If not given, cross-attention defaults to
|
||||
self-attention.
|
||||
timestep ( `torch.LongTensor`, *optional*):
|
||||
Used to indicate denoising step. Optional timestep to be applied as an embedding in `AdaLayerNorm`.
|
||||
class_labels ( `torch.LongTensor` of shape `(batch size, num classes)`, *optional*):
|
||||
Used to indicate class labels conditioning. Optional class labels to be applied as an embedding in
|
||||
`AdaLayerZeroNorm`.
|
||||
cross_attention_kwargs ( `Dict[str, Any]`, *optional*):
|
||||
A kwargs dictionary that if specified is passed along to the `AttentionProcessor` as defined under
|
||||
`self.processor` in
|
||||
[diffusers.models.attention_processor](https://github.com/huggingface/diffusers/blob/main/src/diffusers/models/attention_processor.py).
|
||||
attention_mask ( `torch.Tensor`, *optional*):
|
||||
An attention mask of shape `(batch, key_tokens)` is applied to `encoder_hidden_states`. If `1` the mask
|
||||
is kept, otherwise if `0` it is discarded. Mask will be converted into a bias, which adds large
|
||||
negative values to the attention scores corresponding to "discard" tokens.
|
||||
encoder_attention_mask ( `torch.Tensor`, *optional*):
|
||||
Cross-attention mask applied to `encoder_hidden_states`. Two formats supported:
|
||||
|
||||
* Mask `(batch, sequence_length)` True = keep, False = discard.
|
||||
* Bias `(batch, 1, sequence_length)` 0 = keep, -10000 = discard.
|
||||
|
||||
If `ndim == 2`: will be interpreted as a mask, then converted into a bias consistent with the format
|
||||
above. This bias will be added to the cross-attention scores.
|
||||
skip_layer_mask ( `torch.Tensor`, *optional*):
|
||||
A mask of shape `(num_layers, batch)` that indicates which layers to skip. `0` at position
|
||||
`layer, batch_idx` indicates that the layer should be skipped for the corresponding batch index.
|
||||
skip_layer_strategy ( `SkipLayerStrategy`, *optional*, defaults to `None`):
|
||||
Controls which layers are skipped when calculating a perturbed latent for spatiotemporal guidance.
|
||||
return_dict (`bool`, *optional*, defaults to `True`):
|
||||
Whether or not to return a [`~models.unets.unet_2d_condition.UNet2DConditionOutput`] instead of a plain
|
||||
tuple.
|
||||
|
||||
Returns:
|
||||
If `return_dict` is True, an [`~models.transformer_2d.Transformer2DModelOutput`] is returned, otherwise a
|
||||
`tuple` where the first element is the sample tensor.
|
||||
"""
|
||||
# for tpu attention offload 2d token masks are used. No need to transform.
|
||||
if not self.use_tpu_flash_attention:
|
||||
# ensure attention_mask is a bias, and give it a singleton query_tokens dimension.
|
||||
# we may have done this conversion already, e.g. if we came here via UNet2DConditionModel#forward.
|
||||
# we can tell by counting dims; if ndim == 2: it's a mask rather than a bias.
|
||||
# expects mask of shape:
|
||||
# [batch, key_tokens]
|
||||
# adds singleton query_tokens dimension:
|
||||
# [batch, 1, key_tokens]
|
||||
# this helps to broadcast it as a bias over attention scores, which will be in one of the following shapes:
|
||||
# [batch, heads, query_tokens, key_tokens] (e.g. torch sdp attn)
|
||||
# [batch * heads, query_tokens, key_tokens] (e.g. xformers or classic attn)
|
||||
if attention_mask is not None and attention_mask.ndim == 2:
|
||||
# assume that mask is expressed as:
|
||||
# (1 = keep, 0 = discard)
|
||||
# convert mask into a bias that can be added to attention scores:
|
||||
# (keep = +0, discard = -10000.0)
|
||||
attention_mask = (1 - attention_mask.to(hidden_states.dtype)) * -10000.0
|
||||
attention_mask = attention_mask.unsqueeze(1)
|
||||
|
||||
# convert encoder_attention_mask to a bias the same way we do for attention_mask
|
||||
if encoder_attention_mask is not None and encoder_attention_mask.ndim == 2:
|
||||
encoder_attention_mask = (
|
||||
1 - encoder_attention_mask.to(hidden_states.dtype)
|
||||
) * -10000.0
|
||||
encoder_attention_mask = encoder_attention_mask.unsqueeze(1)
|
||||
|
||||
# 1. Input
|
||||
hidden_states = self.patchify_proj(hidden_states)
|
||||
|
||||
if self.timestep_scale_multiplier:
|
||||
timestep = self.timestep_scale_multiplier * timestep
|
||||
|
||||
if timestep.shape[-1] > 1:
|
||||
timestep = timestep.reshape(timestep.shape[0], -1, latent_shape[-2] * latent_shape[-1] )
|
||||
timestep = timestep[:, :, 0]
|
||||
|
||||
batch_size = hidden_states.shape[0]
|
||||
timestep, embedded_timestep = self.adaln_single(
|
||||
timestep.flatten(),
|
||||
{"resolution": None, "aspect_ratio": None},
|
||||
batch_size=batch_size,
|
||||
hidden_dtype=hidden_states.dtype,
|
||||
)
|
||||
# Second dimension is 1 or number of tokens (if timestep_per_token)
|
||||
timestep = timestep.view(batch_size, -1, timestep.shape[-1])
|
||||
embedded_timestep = embedded_timestep.view(
|
||||
batch_size, -1, embedded_timestep.shape[-1]
|
||||
)
|
||||
if mixed:
|
||||
timestep = timestep.float()
|
||||
embedded_timestep = embedded_timestep.float()
|
||||
hidden_states = hidden_states.float()
|
||||
|
||||
|
||||
# 2. Blocks
|
||||
if self.caption_projection is not None:
|
||||
batch_size = hidden_states.shape[0]
|
||||
encoder_hidden_states = self.caption_projection(encoder_hidden_states)
|
||||
encoder_hidden_states = encoder_hidden_states.view(
|
||||
batch_size, -1, hidden_states.shape[-1]
|
||||
)
|
||||
|
||||
|
||||
if joint_pass:
|
||||
for block_idx, block in enumerate(self.transformer_blocks):
|
||||
hidden_states = block(
|
||||
hidden_states,
|
||||
freqs_cis=freqs_cis,
|
||||
attention_mask=attention_mask,
|
||||
encoder_hidden_states=encoder_hidden_states,
|
||||
encoder_attention_mask=encoder_attention_mask,
|
||||
timestep=timestep,
|
||||
cross_attention_kwargs=cross_attention_kwargs,
|
||||
class_labels=class_labels,
|
||||
skip_layer_mask= None if skip_layer_mask is None else skip_layer_mask[block_idx],
|
||||
skip_layer_strategy=skip_layer_strategy,
|
||||
)
|
||||
if ltxv_model._interrupt:
|
||||
return [None]
|
||||
|
||||
else:
|
||||
for block_idx, block in enumerate(self.transformer_blocks):
|
||||
for i, (one_hidden_states, one_encoder_hidden_states, one_encoder_attention_mask,one_timestep) in enumerate(zip(hidden_states, encoder_hidden_states,encoder_attention_mask,timestep)):
|
||||
hidden_states[i][...] = block(
|
||||
one_hidden_states.unsqueeze(0),
|
||||
freqs_cis=freqs_cis,
|
||||
attention_mask=attention_mask,
|
||||
encoder_hidden_states=one_encoder_hidden_states.unsqueeze(0),
|
||||
encoder_attention_mask=one_encoder_attention_mask.unsqueeze(0),
|
||||
timestep=one_timestep.unsqueeze(0),
|
||||
cross_attention_kwargs=cross_attention_kwargs,
|
||||
class_labels=class_labels,
|
||||
skip_layer_mask= None if skip_layer_mask is None else skip_layer_mask[block_idx, i],
|
||||
skip_layer_strategy=skip_layer_strategy,
|
||||
)
|
||||
if ltxv_model._interrupt:
|
||||
return [None]
|
||||
|
||||
# 3. Output
|
||||
scale_shift_values = (
|
||||
self.scale_shift_table[None, None] + embedded_timestep[:, :, None]
|
||||
)
|
||||
shift, scale = scale_shift_values[:, :, 0].unsqueeze(-2), scale_shift_values[:, :, 1].unsqueeze(-2)
|
||||
hidden_states = self.norm_out(hidden_states)
|
||||
# Modulation
|
||||
|
||||
|
||||
hidden_states = reshape_hidden_states(hidden_states, scale.shape[1])
|
||||
# hidden_states = hidden_states * (1 + scale)
|
||||
hidden_states *= 1 + scale
|
||||
hidden_states += shift
|
||||
hidden_states = restore_hidden_states_shape(hidden_states)
|
||||
hidden_states = self.proj_out(hidden_states)
|
||||
if not return_dict:
|
||||
return (hidden_states,)
|
||||
|
||||
return Transformer3DModelOutput(sample=hidden_states)
|
||||
0
ltx_video/pipelines/__init__.py
Normal file
0
ltx_video/pipelines/__init__.py
Normal file
50
ltx_video/pipelines/crf_compressor.py
Normal file
50
ltx_video/pipelines/crf_compressor.py
Normal file
@@ -0,0 +1,50 @@
|
||||
import av
|
||||
import torch
|
||||
import io
|
||||
import numpy as np
|
||||
|
||||
|
||||
def _encode_single_frame(output_file, image_array: np.ndarray, crf):
|
||||
container = av.open(output_file, "w", format="mp4")
|
||||
try:
|
||||
stream = container.add_stream(
|
||||
"libx264", rate=1, options={"crf": str(crf), "preset": "veryfast"}
|
||||
)
|
||||
stream.height = image_array.shape[0]
|
||||
stream.width = image_array.shape[1]
|
||||
av_frame = av.VideoFrame.from_ndarray(image_array, format="rgb24").reformat(
|
||||
format="yuv420p"
|
||||
)
|
||||
container.mux(stream.encode(av_frame))
|
||||
container.mux(stream.encode())
|
||||
finally:
|
||||
container.close()
|
||||
|
||||
|
||||
def _decode_single_frame(video_file):
|
||||
container = av.open(video_file)
|
||||
try:
|
||||
stream = next(s for s in container.streams if s.type == "video")
|
||||
frame = next(container.decode(stream))
|
||||
finally:
|
||||
container.close()
|
||||
return frame.to_ndarray(format="rgb24")
|
||||
|
||||
|
||||
def compress(image: torch.Tensor, crf=29):
|
||||
if crf == 0:
|
||||
return image
|
||||
|
||||
image_array = (
|
||||
(image[: (image.shape[0] // 2) * 2, : (image.shape[1] // 2) * 2] * 255.0)
|
||||
.byte()
|
||||
.cpu()
|
||||
.numpy()
|
||||
)
|
||||
with io.BytesIO() as output_file:
|
||||
_encode_single_frame(output_file, image_array, crf)
|
||||
video_bytes = output_file.getvalue()
|
||||
with io.BytesIO(video_bytes) as video_file:
|
||||
image_array = _decode_single_frame(video_file)
|
||||
tensor = torch.tensor(image_array, dtype=image.dtype, device=image.device) / 255.0
|
||||
return tensor
|
||||
1903
ltx_video/pipelines/pipeline_ltx_video.py
Normal file
1903
ltx_video/pipelines/pipeline_ltx_video.py
Normal file
File diff suppressed because it is too large
Load Diff
0
ltx_video/schedulers/__init__.py
Normal file
0
ltx_video/schedulers/__init__.py
Normal file
392
ltx_video/schedulers/rf.py
Normal file
392
ltx_video/schedulers/rf.py
Normal file
@@ -0,0 +1,392 @@
|
||||
import math
|
||||
from abc import ABC, abstractmethod
|
||||
from dataclasses import dataclass
|
||||
from typing import Callable, Optional, Tuple, Union
|
||||
import json
|
||||
import os
|
||||
from pathlib import Path
|
||||
|
||||
import torch
|
||||
from diffusers.configuration_utils import ConfigMixin, register_to_config
|
||||
from diffusers.schedulers.scheduling_utils import SchedulerMixin
|
||||
from diffusers.utils import BaseOutput
|
||||
from torch import Tensor
|
||||
from safetensors import safe_open
|
||||
|
||||
|
||||
from ltx_video.utils.torch_utils import append_dims
|
||||
|
||||
from ltx_video.utils.diffusers_config_mapping import (
|
||||
diffusers_and_ours_config_mapping,
|
||||
make_hashable_key,
|
||||
)
|
||||
|
||||
|
||||
def linear_quadratic_schedule(num_steps, threshold_noise=0.025, linear_steps=None):
|
||||
if num_steps == 1:
|
||||
return torch.tensor([1.0])
|
||||
if linear_steps is None:
|
||||
linear_steps = num_steps // 2
|
||||
linear_sigma_schedule = [
|
||||
i * threshold_noise / linear_steps for i in range(linear_steps)
|
||||
]
|
||||
threshold_noise_step_diff = linear_steps - threshold_noise * num_steps
|
||||
quadratic_steps = num_steps - linear_steps
|
||||
quadratic_coef = threshold_noise_step_diff / (linear_steps * quadratic_steps**2)
|
||||
linear_coef = threshold_noise / linear_steps - 2 * threshold_noise_step_diff / (
|
||||
quadratic_steps**2
|
||||
)
|
||||
const = quadratic_coef * (linear_steps**2)
|
||||
quadratic_sigma_schedule = [
|
||||
quadratic_coef * (i**2) + linear_coef * i + const
|
||||
for i in range(linear_steps, num_steps)
|
||||
]
|
||||
sigma_schedule = linear_sigma_schedule + quadratic_sigma_schedule + [1.0]
|
||||
sigma_schedule = [1.0 - x for x in sigma_schedule]
|
||||
return torch.tensor(sigma_schedule[:-1])
|
||||
|
||||
|
||||
def simple_diffusion_resolution_dependent_timestep_shift(
|
||||
samples_shape: torch.Size,
|
||||
timesteps: Tensor,
|
||||
n: int = 32 * 32,
|
||||
) -> Tensor:
|
||||
if len(samples_shape) == 3:
|
||||
_, m, _ = samples_shape
|
||||
elif len(samples_shape) in [4, 5]:
|
||||
m = math.prod(samples_shape[2:])
|
||||
else:
|
||||
raise ValueError(
|
||||
"Samples must have shape (b, t, c), (b, c, h, w) or (b, c, f, h, w)"
|
||||
)
|
||||
snr = (timesteps / (1 - timesteps)) ** 2
|
||||
shift_snr = torch.log(snr) + 2 * math.log(m / n)
|
||||
shifted_timesteps = torch.sigmoid(0.5 * shift_snr)
|
||||
|
||||
return shifted_timesteps
|
||||
|
||||
|
||||
def time_shift(mu: float, sigma: float, t: Tensor):
|
||||
return math.exp(mu) / (math.exp(mu) + (1 / t - 1) ** sigma)
|
||||
|
||||
|
||||
def get_normal_shift(
|
||||
n_tokens: int,
|
||||
min_tokens: int = 1024,
|
||||
max_tokens: int = 4096,
|
||||
min_shift: float = 0.95,
|
||||
max_shift: float = 2.05,
|
||||
) -> Callable[[float], float]:
|
||||
m = (max_shift - min_shift) / (max_tokens - min_tokens)
|
||||
b = min_shift - m * min_tokens
|
||||
return m * n_tokens + b
|
||||
|
||||
|
||||
def strech_shifts_to_terminal(shifts: Tensor, terminal=0.1):
|
||||
"""
|
||||
Stretch a function (given as sampled shifts) so that its final value matches the given terminal value
|
||||
using the provided formula.
|
||||
|
||||
Parameters:
|
||||
- shifts (Tensor): The samples of the function to be stretched (PyTorch Tensor).
|
||||
- terminal (float): The desired terminal value (value at the last sample).
|
||||
|
||||
Returns:
|
||||
- Tensor: The stretched shifts such that the final value equals `terminal`.
|
||||
"""
|
||||
if shifts.numel() == 0:
|
||||
raise ValueError("The 'shifts' tensor must not be empty.")
|
||||
|
||||
# Ensure terminal value is valid
|
||||
if terminal <= 0 or terminal >= 1:
|
||||
raise ValueError("The terminal value must be between 0 and 1 (exclusive).")
|
||||
|
||||
# Transform the shifts using the given formula
|
||||
one_minus_z = 1 - shifts
|
||||
scale_factor = one_minus_z[-1] / (1 - terminal)
|
||||
stretched_shifts = 1 - (one_minus_z / scale_factor)
|
||||
|
||||
return stretched_shifts
|
||||
|
||||
|
||||
def sd3_resolution_dependent_timestep_shift(
|
||||
samples_shape: torch.Size,
|
||||
timesteps: Tensor,
|
||||
target_shift_terminal: Optional[float] = None,
|
||||
) -> Tensor:
|
||||
"""
|
||||
Shifts the timestep schedule as a function of the generated resolution.
|
||||
|
||||
In the SD3 paper, the authors empirically how to shift the timesteps based on the resolution of the target images.
|
||||
For more details: https://arxiv.org/pdf/2403.03206
|
||||
|
||||
In Flux they later propose a more dynamic resolution dependent timestep shift, see:
|
||||
https://github.com/black-forest-labs/flux/blob/87f6fff727a377ea1c378af692afb41ae84cbe04/src/flux/sampling.py#L66
|
||||
|
||||
|
||||
Args:
|
||||
samples_shape (torch.Size): The samples batch shape (batch_size, channels, height, width) or
|
||||
(batch_size, channels, frame, height, width).
|
||||
timesteps (Tensor): A batch of timesteps with shape (batch_size,).
|
||||
target_shift_terminal (float): The target terminal value for the shifted timesteps.
|
||||
|
||||
Returns:
|
||||
Tensor: The shifted timesteps.
|
||||
"""
|
||||
if len(samples_shape) == 3:
|
||||
_, m, _ = samples_shape
|
||||
elif len(samples_shape) in [4, 5]:
|
||||
m = math.prod(samples_shape[2:])
|
||||
else:
|
||||
raise ValueError(
|
||||
"Samples must have shape (b, t, c), (b, c, h, w) or (b, c, f, h, w)"
|
||||
)
|
||||
|
||||
shift = get_normal_shift(m)
|
||||
time_shifts = time_shift(shift, 1, timesteps)
|
||||
if target_shift_terminal is not None: # Stretch the shifts to the target terminal
|
||||
time_shifts = strech_shifts_to_terminal(time_shifts, target_shift_terminal)
|
||||
return time_shifts
|
||||
|
||||
|
||||
class TimestepShifter(ABC):
|
||||
@abstractmethod
|
||||
def shift_timesteps(self, samples_shape: torch.Size, timesteps: Tensor) -> Tensor:
|
||||
pass
|
||||
|
||||
|
||||
@dataclass
|
||||
class RectifiedFlowSchedulerOutput(BaseOutput):
|
||||
"""
|
||||
Output class for the scheduler's step function output.
|
||||
|
||||
Args:
|
||||
prev_sample (`torch.FloatTensor` of shape `(batch_size, num_channels, height, width)` for images):
|
||||
Computed sample (x_{t-1}) of previous timestep. `prev_sample` should be used as next model input in the
|
||||
denoising loop.
|
||||
pred_original_sample (`torch.FloatTensor` of shape `(batch_size, num_channels, height, width)` for images):
|
||||
The predicted denoised sample (x_{0}) based on the model output from the current timestep.
|
||||
`pred_original_sample` can be used to preview progress or for guidance.
|
||||
"""
|
||||
|
||||
prev_sample: torch.FloatTensor
|
||||
pred_original_sample: Optional[torch.FloatTensor] = None
|
||||
|
||||
|
||||
class RectifiedFlowScheduler(SchedulerMixin, ConfigMixin, TimestepShifter):
|
||||
order = 1
|
||||
|
||||
@register_to_config
|
||||
def __init__(
|
||||
self,
|
||||
num_train_timesteps=1000,
|
||||
shifting: Optional[str] = None,
|
||||
base_resolution: int = 32**2,
|
||||
target_shift_terminal: Optional[float] = None,
|
||||
sampler: Optional[str] = "Uniform",
|
||||
shift: Optional[float] = None,
|
||||
):
|
||||
super().__init__()
|
||||
self.init_noise_sigma = 1.0
|
||||
self.num_inference_steps = None
|
||||
self.sampler = sampler
|
||||
self.shifting = shifting
|
||||
self.base_resolution = base_resolution
|
||||
self.target_shift_terminal = target_shift_terminal
|
||||
self.timesteps = self.sigmas = self.get_initial_timesteps(
|
||||
num_train_timesteps, shift=shift
|
||||
)
|
||||
self.shift = shift
|
||||
|
||||
def get_initial_timesteps(
|
||||
self, num_timesteps: int, shift: Optional[float] = None
|
||||
) -> Tensor:
|
||||
if self.sampler == "Uniform":
|
||||
return torch.linspace(1, 1 / num_timesteps, num_timesteps)
|
||||
elif self.sampler == "LinearQuadratic":
|
||||
return linear_quadratic_schedule(num_timesteps)
|
||||
elif self.sampler == "Constant":
|
||||
assert (
|
||||
shift is not None
|
||||
), "Shift must be provided for constant time shift sampler."
|
||||
return time_shift(
|
||||
shift, 1, torch.linspace(1, 1 / num_timesteps, num_timesteps)
|
||||
)
|
||||
|
||||
def shift_timesteps(self, samples_shape: torch.Size, timesteps: Tensor) -> Tensor:
|
||||
if self.shifting == "SD3":
|
||||
return sd3_resolution_dependent_timestep_shift(
|
||||
samples_shape, timesteps, self.target_shift_terminal
|
||||
)
|
||||
elif self.shifting == "SimpleDiffusion":
|
||||
return simple_diffusion_resolution_dependent_timestep_shift(
|
||||
samples_shape, timesteps, self.base_resolution
|
||||
)
|
||||
return timesteps
|
||||
|
||||
def set_timesteps(
|
||||
self,
|
||||
num_inference_steps: Optional[int] = None,
|
||||
samples_shape: Optional[torch.Size] = None,
|
||||
timesteps: Optional[Tensor] = None,
|
||||
device: Union[str, torch.device] = None,
|
||||
):
|
||||
"""
|
||||
Sets the discrete timesteps used for the diffusion chain. Supporting function to be run before inference.
|
||||
If `timesteps` are provided, they will be used instead of the scheduled timesteps.
|
||||
|
||||
Args:
|
||||
num_inference_steps (`int` *optional*): The number of diffusion steps used when generating samples.
|
||||
samples_shape (`torch.Size` *optional*): The samples batch shape, used for shifting.
|
||||
timesteps ('torch.Tensor' *optional*): Specific timesteps to use instead of scheduled timesteps.
|
||||
device (`Union[str, torch.device]`, *optional*): The device to which the timesteps tensor will be moved.
|
||||
"""
|
||||
if timesteps is not None and num_inference_steps is not None:
|
||||
raise ValueError(
|
||||
"You cannot provide both `timesteps` and `num_inference_steps`."
|
||||
)
|
||||
if timesteps is None:
|
||||
num_inference_steps = min(
|
||||
self.config.num_train_timesteps, num_inference_steps
|
||||
)
|
||||
timesteps = self.get_initial_timesteps(
|
||||
num_inference_steps, shift=self.shift
|
||||
).to(device)
|
||||
timesteps = self.shift_timesteps(samples_shape, timesteps)
|
||||
else:
|
||||
timesteps = torch.Tensor(timesteps).to(device)
|
||||
num_inference_steps = len(timesteps)
|
||||
self.timesteps = timesteps
|
||||
self.num_inference_steps = num_inference_steps
|
||||
self.sigmas = self.timesteps
|
||||
|
||||
@staticmethod
|
||||
def from_pretrained(pretrained_model_path: Union[str, os.PathLike]):
|
||||
with open(pretrained_model_path, "r", encoding="utf-8") as reader:
|
||||
text = reader.read()
|
||||
|
||||
config = json.loads(text)
|
||||
return RectifiedFlowScheduler.from_config(config)
|
||||
|
||||
pretrained_model_path = Path(pretrained_model_path)
|
||||
if pretrained_model_path.is_file():
|
||||
comfy_single_file_state_dict = {}
|
||||
with safe_open(pretrained_model_path, framework="pt", device="cpu") as f:
|
||||
metadata = f.metadata()
|
||||
for k in f.keys():
|
||||
comfy_single_file_state_dict[k] = f.get_tensor(k)
|
||||
configs = json.loads(metadata["config"])
|
||||
config = configs["scheduler"]
|
||||
del comfy_single_file_state_dict
|
||||
|
||||
elif pretrained_model_path.is_dir():
|
||||
diffusers_noise_scheduler_config_path = (
|
||||
pretrained_model_path / "scheduler" / "scheduler_config.json"
|
||||
)
|
||||
|
||||
with open(diffusers_noise_scheduler_config_path, "r") as f:
|
||||
scheduler_config = json.load(f)
|
||||
hashable_config = make_hashable_key(scheduler_config)
|
||||
if hashable_config in diffusers_and_ours_config_mapping:
|
||||
config = diffusers_and_ours_config_mapping[hashable_config]
|
||||
return RectifiedFlowScheduler.from_config(config)
|
||||
|
||||
def scale_model_input(
|
||||
self, sample: torch.FloatTensor, timestep: Optional[int] = None
|
||||
) -> torch.FloatTensor:
|
||||
# pylint: disable=unused-argument
|
||||
"""
|
||||
Ensures interchangeability with schedulers that need to scale the denoising model input depending on the
|
||||
current timestep.
|
||||
|
||||
Args:
|
||||
sample (`torch.FloatTensor`): input sample
|
||||
timestep (`int`, optional): current timestep
|
||||
|
||||
Returns:
|
||||
`torch.FloatTensor`: scaled input sample
|
||||
"""
|
||||
return sample
|
||||
|
||||
def step(
|
||||
self,
|
||||
model_output: torch.FloatTensor,
|
||||
timestep: torch.FloatTensor,
|
||||
sample: torch.FloatTensor,
|
||||
return_dict: bool = True,
|
||||
stochastic_sampling: Optional[bool] = False,
|
||||
**kwargs,
|
||||
) -> Union[RectifiedFlowSchedulerOutput, Tuple]:
|
||||
"""
|
||||
Predict the sample from the previous timestep by reversing the SDE. This function propagates the diffusion
|
||||
process from the learned model outputs (most often the predicted noise).
|
||||
z_{t_1} = z_t - \Delta_t * v
|
||||
The method finds the next timestep that is lower than the input timestep(s) and denoises the latents
|
||||
to that level. The input timestep(s) are not required to be one of the predefined timesteps.
|
||||
|
||||
Args:
|
||||
model_output (`torch.FloatTensor`):
|
||||
The direct output from learned diffusion model - the velocity,
|
||||
timestep (`float`):
|
||||
The current discrete timestep in the diffusion chain (global or per-token).
|
||||
sample (`torch.FloatTensor`):
|
||||
A current latent tokens to be de-noised.
|
||||
return_dict (`bool`, *optional*, defaults to `True`):
|
||||
Whether or not to return a [`~schedulers.scheduling_ddim.DDIMSchedulerOutput`] or `tuple`.
|
||||
stochastic_sampling (`bool`, *optional*, defaults to `False`):
|
||||
Whether to use stochastic sampling for the sampling process.
|
||||
|
||||
Returns:
|
||||
[`~schedulers.scheduling_utils.RectifiedFlowSchedulerOutput`] or `tuple`:
|
||||
If return_dict is `True`, [`~schedulers.rf_scheduler.RectifiedFlowSchedulerOutput`] is returned,
|
||||
otherwise a tuple is returned where the first element is the sample tensor.
|
||||
"""
|
||||
if self.num_inference_steps is None:
|
||||
raise ValueError(
|
||||
"Number of inference steps is 'None', you need to run 'set_timesteps' after creating the scheduler"
|
||||
)
|
||||
t_eps = 1e-6 # Small epsilon to avoid numerical issues in timestep values
|
||||
|
||||
timesteps_padded = torch.cat(
|
||||
[self.timesteps, torch.zeros(1, device=self.timesteps.device)]
|
||||
)
|
||||
|
||||
# Find the next lower timestep(s) and compute the dt from the current timestep(s)
|
||||
if timestep.ndim == 0:
|
||||
# Global timestep case
|
||||
lower_mask = timesteps_padded < timestep - t_eps
|
||||
lower_timestep = timesteps_padded[lower_mask][0] # Closest lower timestep
|
||||
dt = timestep - lower_timestep
|
||||
|
||||
else:
|
||||
# Per-token case
|
||||
assert timestep.ndim == 2
|
||||
lower_mask = timesteps_padded[:, None, None] < timestep[None] - t_eps
|
||||
lower_timestep = lower_mask * timesteps_padded[:, None, None]
|
||||
lower_timestep, _ = lower_timestep.max(dim=0)
|
||||
dt = (timestep - lower_timestep)[..., None]
|
||||
|
||||
# Compute previous sample
|
||||
if stochastic_sampling:
|
||||
x0 = sample - timestep[..., None] * model_output
|
||||
next_timestep = timestep[..., None] - dt
|
||||
prev_sample = self.add_noise(x0, torch.randn_like(sample), next_timestep)
|
||||
else:
|
||||
prev_sample = sample - dt * model_output
|
||||
|
||||
if not return_dict:
|
||||
return (prev_sample,)
|
||||
|
||||
return RectifiedFlowSchedulerOutput(prev_sample=prev_sample)
|
||||
|
||||
def add_noise(
|
||||
self,
|
||||
original_samples: torch.FloatTensor,
|
||||
noise: torch.FloatTensor,
|
||||
timesteps: torch.FloatTensor,
|
||||
) -> torch.FloatTensor:
|
||||
sigmas = timesteps
|
||||
sigmas = append_dims(sigmas, original_samples.ndim)
|
||||
alphas = 1 - sigmas
|
||||
noisy_samples = alphas * original_samples + sigmas * noise
|
||||
return noisy_samples
|
||||
0
ltx_video/utils/__init__.py
Normal file
0
ltx_video/utils/__init__.py
Normal file
174
ltx_video/utils/diffusers_config_mapping.py
Normal file
174
ltx_video/utils/diffusers_config_mapping.py
Normal file
@@ -0,0 +1,174 @@
|
||||
def make_hashable_key(dict_key):
|
||||
def convert_value(value):
|
||||
if isinstance(value, list):
|
||||
return tuple(value)
|
||||
elif isinstance(value, dict):
|
||||
return tuple(sorted((k, convert_value(v)) for k, v in value.items()))
|
||||
else:
|
||||
return value
|
||||
|
||||
return tuple(sorted((k, convert_value(v)) for k, v in dict_key.items()))
|
||||
|
||||
|
||||
DIFFUSERS_SCHEDULER_CONFIG = {
|
||||
"_class_name": "FlowMatchEulerDiscreteScheduler",
|
||||
"_diffusers_version": "0.32.0.dev0",
|
||||
"base_image_seq_len": 1024,
|
||||
"base_shift": 0.95,
|
||||
"invert_sigmas": False,
|
||||
"max_image_seq_len": 4096,
|
||||
"max_shift": 2.05,
|
||||
"num_train_timesteps": 1000,
|
||||
"shift": 1.0,
|
||||
"shift_terminal": 0.1,
|
||||
"use_beta_sigmas": False,
|
||||
"use_dynamic_shifting": True,
|
||||
"use_exponential_sigmas": False,
|
||||
"use_karras_sigmas": False,
|
||||
}
|
||||
DIFFUSERS_TRANSFORMER_CONFIG = {
|
||||
"_class_name": "LTXVideoTransformer3DModel",
|
||||
"_diffusers_version": "0.32.0.dev0",
|
||||
"activation_fn": "gelu-approximate",
|
||||
"attention_bias": True,
|
||||
"attention_head_dim": 64,
|
||||
"attention_out_bias": True,
|
||||
"caption_channels": 4096,
|
||||
"cross_attention_dim": 2048,
|
||||
"in_channels": 128,
|
||||
"norm_elementwise_affine": False,
|
||||
"norm_eps": 1e-06,
|
||||
"num_attention_heads": 32,
|
||||
"num_layers": 28,
|
||||
"out_channels": 128,
|
||||
"patch_size": 1,
|
||||
"patch_size_t": 1,
|
||||
"qk_norm": "rms_norm_across_heads",
|
||||
}
|
||||
DIFFUSERS_VAE_CONFIG = {
|
||||
"_class_name": "AutoencoderKLLTXVideo",
|
||||
"_diffusers_version": "0.32.0.dev0",
|
||||
"block_out_channels": [128, 256, 512, 512],
|
||||
"decoder_causal": False,
|
||||
"encoder_causal": True,
|
||||
"in_channels": 3,
|
||||
"latent_channels": 128,
|
||||
"layers_per_block": [4, 3, 3, 3, 4],
|
||||
"out_channels": 3,
|
||||
"patch_size": 4,
|
||||
"patch_size_t": 1,
|
||||
"resnet_norm_eps": 1e-06,
|
||||
"scaling_factor": 1.0,
|
||||
"spatio_temporal_scaling": [True, True, True, False],
|
||||
}
|
||||
|
||||
OURS_SCHEDULER_CONFIG = {
|
||||
"_class_name": "RectifiedFlowScheduler",
|
||||
"_diffusers_version": "0.25.1",
|
||||
"num_train_timesteps": 1000,
|
||||
"shifting": "SD3",
|
||||
"base_resolution": None,
|
||||
"target_shift_terminal": 0.1,
|
||||
}
|
||||
|
||||
OURS_TRANSFORMER_CONFIG = {
|
||||
"_class_name": "Transformer3DModel",
|
||||
"_diffusers_version": "0.25.1",
|
||||
"_name_or_path": "PixArt-alpha/PixArt-XL-2-256x256",
|
||||
"activation_fn": "gelu-approximate",
|
||||
"attention_bias": True,
|
||||
"attention_head_dim": 64,
|
||||
"attention_type": "default",
|
||||
"caption_channels": 4096,
|
||||
"cross_attention_dim": 2048,
|
||||
"double_self_attention": False,
|
||||
"dropout": 0.0,
|
||||
"in_channels": 128,
|
||||
"norm_elementwise_affine": False,
|
||||
"norm_eps": 1e-06,
|
||||
"norm_num_groups": 32,
|
||||
"num_attention_heads": 32,
|
||||
"num_embeds_ada_norm": 1000,
|
||||
"num_layers": 28,
|
||||
"num_vector_embeds": None,
|
||||
"only_cross_attention": False,
|
||||
"out_channels": 128,
|
||||
"project_to_2d_pos": True,
|
||||
"upcast_attention": False,
|
||||
"use_linear_projection": False,
|
||||
"qk_norm": "rms_norm",
|
||||
"standardization_norm": "rms_norm",
|
||||
"positional_embedding_type": "rope",
|
||||
"positional_embedding_theta": 10000.0,
|
||||
"positional_embedding_max_pos": [20, 2048, 2048],
|
||||
"timestep_scale_multiplier": 1000,
|
||||
}
|
||||
OURS_VAE_CONFIG = {
|
||||
"_class_name": "CausalVideoAutoencoder",
|
||||
"dims": 3,
|
||||
"in_channels": 3,
|
||||
"out_channels": 3,
|
||||
"latent_channels": 128,
|
||||
"blocks": [
|
||||
["res_x", 4],
|
||||
["compress_all", 1],
|
||||
["res_x_y", 1],
|
||||
["res_x", 3],
|
||||
["compress_all", 1],
|
||||
["res_x_y", 1],
|
||||
["res_x", 3],
|
||||
["compress_all", 1],
|
||||
["res_x", 3],
|
||||
["res_x", 4],
|
||||
],
|
||||
"scaling_factor": 1.0,
|
||||
"norm_layer": "pixel_norm",
|
||||
"patch_size": 4,
|
||||
"latent_log_var": "uniform",
|
||||
"use_quant_conv": False,
|
||||
"causal_decoder": False,
|
||||
}
|
||||
|
||||
|
||||
diffusers_and_ours_config_mapping = {
|
||||
make_hashable_key(DIFFUSERS_SCHEDULER_CONFIG): OURS_SCHEDULER_CONFIG,
|
||||
make_hashable_key(DIFFUSERS_TRANSFORMER_CONFIG): OURS_TRANSFORMER_CONFIG,
|
||||
make_hashable_key(DIFFUSERS_VAE_CONFIG): OURS_VAE_CONFIG,
|
||||
}
|
||||
|
||||
|
||||
TRANSFORMER_KEYS_RENAME_DICT = {
|
||||
"proj_in": "patchify_proj",
|
||||
"time_embed": "adaln_single",
|
||||
"norm_q": "q_norm",
|
||||
"norm_k": "k_norm",
|
||||
}
|
||||
|
||||
|
||||
VAE_KEYS_RENAME_DICT = {
|
||||
"decoder.up_blocks.3.conv_in": "decoder.up_blocks.7",
|
||||
"decoder.up_blocks.3.upsamplers.0": "decoder.up_blocks.8",
|
||||
"decoder.up_blocks.3": "decoder.up_blocks.9",
|
||||
"decoder.up_blocks.2.upsamplers.0": "decoder.up_blocks.5",
|
||||
"decoder.up_blocks.2.conv_in": "decoder.up_blocks.4",
|
||||
"decoder.up_blocks.2": "decoder.up_blocks.6",
|
||||
"decoder.up_blocks.1.upsamplers.0": "decoder.up_blocks.2",
|
||||
"decoder.up_blocks.1": "decoder.up_blocks.3",
|
||||
"decoder.up_blocks.0": "decoder.up_blocks.1",
|
||||
"decoder.mid_block": "decoder.up_blocks.0",
|
||||
"encoder.down_blocks.3": "encoder.down_blocks.8",
|
||||
"encoder.down_blocks.2.downsamplers.0": "encoder.down_blocks.7",
|
||||
"encoder.down_blocks.2": "encoder.down_blocks.6",
|
||||
"encoder.down_blocks.1.downsamplers.0": "encoder.down_blocks.4",
|
||||
"encoder.down_blocks.1.conv_out": "encoder.down_blocks.5",
|
||||
"encoder.down_blocks.1": "encoder.down_blocks.3",
|
||||
"encoder.down_blocks.0.conv_out": "encoder.down_blocks.2",
|
||||
"encoder.down_blocks.0.downsamplers.0": "encoder.down_blocks.1",
|
||||
"encoder.down_blocks.0": "encoder.down_blocks.0",
|
||||
"encoder.mid_block": "encoder.down_blocks.9",
|
||||
"conv_shortcut.conv": "conv_shortcut",
|
||||
"resnets": "res_blocks",
|
||||
"norm3": "norm3.norm",
|
||||
"latents_mean": "per_channel_statistics.mean-of-means",
|
||||
"latents_std": "per_channel_statistics.std-of-means",
|
||||
}
|
||||
214
ltx_video/utils/prompt_enhance_utils.py
Normal file
214
ltx_video/utils/prompt_enhance_utils.py
Normal file
@@ -0,0 +1,214 @@
|
||||
import logging
|
||||
from typing import Union, List, Optional
|
||||
|
||||
import torch
|
||||
from PIL import Image
|
||||
|
||||
logger = logging.getLogger(__name__) # pylint: disable=invalid-name
|
||||
|
||||
T2V_CINEMATIC_PROMPT = """You are an expert cinematic director with many award winning movies, When writing prompts based on the user input, focus on detailed, chronological descriptions of actions and scenes.
|
||||
Include specific movements, appearances, camera angles, and environmental details - all in a single flowing paragraph.
|
||||
Start directly with the action, and keep descriptions literal and precise.
|
||||
Think like a cinematographer describing a shot list.
|
||||
Do not change the user input intent, just enhance it.
|
||||
Keep within 150 words.
|
||||
For best results, build your prompts using this structure:
|
||||
Start with main action in a single sentence
|
||||
Add specific details about movements and gestures
|
||||
Describe character/object appearances precisely
|
||||
Include background and environment details
|
||||
Specify camera angles and movements
|
||||
Describe lighting and colors
|
||||
Note any changes or sudden events
|
||||
Do not exceed the 150 word limit!
|
||||
Output the enhanced prompt only.
|
||||
"""
|
||||
|
||||
I2V_CINEMATIC_PROMPT = """You are an expert cinematic director with many award winning movies, When writing prompts based on the user input, focus on detailed, chronological descriptions of actions and scenes.
|
||||
Include specific movements, appearances, camera angles, and environmental details - all in a single flowing paragraph.
|
||||
Start directly with the action, and keep descriptions literal and precise.
|
||||
Think like a cinematographer describing a shot list.
|
||||
Keep within 150 words.
|
||||
For best results, build your prompts using this structure:
|
||||
Describe the image first and then add the user input. Image description should be in first priority! Align to the image caption if it contradicts the user text input.
|
||||
Start with main action in a single sentence
|
||||
Add specific details about movements and gestures
|
||||
Describe character/object appearances precisely
|
||||
Include background and environment details
|
||||
Specify camera angles and movements
|
||||
Describe lighting and colors
|
||||
Note any changes or sudden events
|
||||
Align to the image caption if it contradicts the user text input.
|
||||
Do not exceed the 150 word limit!
|
||||
Output the enhanced prompt only.
|
||||
"""
|
||||
|
||||
|
||||
def tensor_to_pil(tensor):
|
||||
# Ensure tensor is in range [-1, 1]
|
||||
assert tensor.min() >= -1 and tensor.max() <= 1
|
||||
|
||||
# Convert from [-1, 1] to [0, 1]
|
||||
tensor = (tensor + 1) / 2
|
||||
|
||||
# Rearrange from [C, H, W] to [H, W, C]
|
||||
tensor = tensor.permute(1, 2, 0)
|
||||
|
||||
# Convert to numpy array and then to uint8 range [0, 255]
|
||||
numpy_image = (tensor.cpu().numpy() * 255).astype("uint8")
|
||||
|
||||
# Convert to PIL Image
|
||||
return Image.fromarray(numpy_image)
|
||||
|
||||
|
||||
def generate_cinematic_prompt(
|
||||
image_caption_model,
|
||||
image_caption_processor,
|
||||
prompt_enhancer_model,
|
||||
prompt_enhancer_tokenizer,
|
||||
prompt: Union[str, List[str]],
|
||||
images: Optional[List] = None,
|
||||
max_new_tokens: int = 256,
|
||||
) -> List[str]:
|
||||
prompts = [prompt] if isinstance(prompt, str) else prompt
|
||||
|
||||
if images is None:
|
||||
prompts = _generate_t2v_prompt(
|
||||
prompt_enhancer_model,
|
||||
prompt_enhancer_tokenizer,
|
||||
prompts,
|
||||
max_new_tokens,
|
||||
T2V_CINEMATIC_PROMPT,
|
||||
)
|
||||
else:
|
||||
|
||||
prompts = _generate_i2v_prompt(
|
||||
image_caption_model,
|
||||
image_caption_processor,
|
||||
prompt_enhancer_model,
|
||||
prompt_enhancer_tokenizer,
|
||||
prompts,
|
||||
images,
|
||||
max_new_tokens,
|
||||
I2V_CINEMATIC_PROMPT,
|
||||
)
|
||||
|
||||
return prompts
|
||||
|
||||
|
||||
def _get_first_frames_from_conditioning_item(conditioning_item) -> List[Image.Image]:
|
||||
frames_tensor = conditioning_item.media_item
|
||||
return [
|
||||
tensor_to_pil(frames_tensor[i, :, 0, :, :])
|
||||
for i in range(frames_tensor.shape[0])
|
||||
]
|
||||
|
||||
|
||||
def _generate_t2v_prompt(
|
||||
prompt_enhancer_model,
|
||||
prompt_enhancer_tokenizer,
|
||||
prompts: List[str],
|
||||
max_new_tokens: int,
|
||||
system_prompt: str,
|
||||
) -> List[str]:
|
||||
messages = [
|
||||
[
|
||||
{"role": "system", "content": system_prompt},
|
||||
{"role": "user", "content": f"user_prompt: {p}"},
|
||||
]
|
||||
for p in prompts
|
||||
]
|
||||
|
||||
texts = [
|
||||
prompt_enhancer_tokenizer.apply_chat_template(
|
||||
m, tokenize=False, add_generation_prompt=True
|
||||
)
|
||||
for m in messages
|
||||
]
|
||||
model_inputs = prompt_enhancer_tokenizer(texts, return_tensors="pt").to(
|
||||
prompt_enhancer_model.device
|
||||
)
|
||||
|
||||
return _generate_and_decode_prompts(
|
||||
prompt_enhancer_model, prompt_enhancer_tokenizer, model_inputs, max_new_tokens
|
||||
)
|
||||
|
||||
|
||||
def _generate_i2v_prompt(
|
||||
image_caption_model,
|
||||
image_caption_processor,
|
||||
prompt_enhancer_model,
|
||||
prompt_enhancer_tokenizer,
|
||||
prompts: List[str],
|
||||
first_frames: List[Image.Image],
|
||||
max_new_tokens: int,
|
||||
system_prompt: str,
|
||||
) -> List[str]:
|
||||
image_captions = _generate_image_captions(
|
||||
image_caption_model, image_caption_processor, first_frames
|
||||
)
|
||||
if len(image_captions) == 1 and len(image_captions) < len(prompts):
|
||||
image_captions *= len(prompts)
|
||||
messages = [
|
||||
[
|
||||
{"role": "system", "content": system_prompt},
|
||||
{"role": "user", "content": f"user_prompt: {p}\nimage_caption: {c}"},
|
||||
]
|
||||
for p, c in zip(prompts, image_captions)
|
||||
]
|
||||
|
||||
texts = [
|
||||
prompt_enhancer_tokenizer.apply_chat_template(
|
||||
m, tokenize=False, add_generation_prompt=True
|
||||
)
|
||||
for m in messages
|
||||
]
|
||||
out_prompts = []
|
||||
for text in texts:
|
||||
model_inputs = prompt_enhancer_tokenizer(text, return_tensors="pt").to(
|
||||
prompt_enhancer_model.device
|
||||
)
|
||||
out_prompts.append(_generate_and_decode_prompts(prompt_enhancer_model, prompt_enhancer_tokenizer, model_inputs, max_new_tokens)[0])
|
||||
|
||||
return out_prompts
|
||||
|
||||
|
||||
def _generate_image_captions(
|
||||
image_caption_model,
|
||||
image_caption_processor,
|
||||
images: List[Image.Image],
|
||||
system_prompt: str = "<DETAILED_CAPTION>",
|
||||
) -> List[str]:
|
||||
image_caption_prompts = [system_prompt] * len(images)
|
||||
inputs = image_caption_processor(
|
||||
image_caption_prompts, images, return_tensors="pt"
|
||||
).to("cuda") #.to(image_caption_model.device)
|
||||
|
||||
with torch.inference_mode():
|
||||
generated_ids = image_caption_model.generate(
|
||||
input_ids=inputs["input_ids"],
|
||||
pixel_values=inputs["pixel_values"],
|
||||
max_new_tokens=1024,
|
||||
do_sample=False,
|
||||
num_beams=3,
|
||||
)
|
||||
|
||||
return image_caption_processor.batch_decode(generated_ids, skip_special_tokens=True)
|
||||
|
||||
|
||||
def _generate_and_decode_prompts(
|
||||
prompt_enhancer_model, prompt_enhancer_tokenizer, model_inputs, max_new_tokens: int
|
||||
) -> List[str]:
|
||||
with torch.inference_mode():
|
||||
outputs = prompt_enhancer_model.generate(
|
||||
**model_inputs, max_new_tokens=max_new_tokens
|
||||
)
|
||||
generated_ids = [
|
||||
output_ids[len(input_ids) :]
|
||||
for input_ids, output_ids in zip(model_inputs.input_ids, outputs)
|
||||
]
|
||||
decoded_prompts = prompt_enhancer_tokenizer.batch_decode(
|
||||
generated_ids, skip_special_tokens=True
|
||||
)
|
||||
|
||||
return decoded_prompts
|
||||
8
ltx_video/utils/skip_layer_strategy.py
Normal file
8
ltx_video/utils/skip_layer_strategy.py
Normal file
@@ -0,0 +1,8 @@
|
||||
from enum import Enum, auto
|
||||
|
||||
|
||||
class SkipLayerStrategy(Enum):
|
||||
AttentionSkip = auto()
|
||||
AttentionValues = auto()
|
||||
Residual = auto()
|
||||
TransformerBlock = auto()
|
||||
25
ltx_video/utils/torch_utils.py
Normal file
25
ltx_video/utils/torch_utils.py
Normal file
@@ -0,0 +1,25 @@
|
||||
import torch
|
||||
from torch import nn
|
||||
|
||||
|
||||
def append_dims(x: torch.Tensor, target_dims: int) -> torch.Tensor:
|
||||
"""Appends dimensions to the end of a tensor until it has target_dims dimensions."""
|
||||
dims_to_append = target_dims - x.ndim
|
||||
if dims_to_append < 0:
|
||||
raise ValueError(
|
||||
f"input has {x.ndim} dims but target_dims is {target_dims}, which is less"
|
||||
)
|
||||
elif dims_to_append == 0:
|
||||
return x
|
||||
return x[(...,) + (None,) * dims_to_append]
|
||||
|
||||
|
||||
class Identity(nn.Module):
|
||||
"""A placeholder identity operator that is argument-insensitive."""
|
||||
|
||||
def __init__(self, *args, **kwargs) -> None: # pylint: disable=unused-argument
|
||||
super().__init__()
|
||||
|
||||
# pylint: disable=unused-argument
|
||||
def forward(self, x: torch.Tensor, *args, **kwargs) -> torch.Tensor:
|
||||
return x
|
||||
@@ -457,13 +457,20 @@ def export_to_vace_video_input(foreground_video_output):
|
||||
gr.Info("Masked Video Input transferred to Vace For Inpainting")
|
||||
return "V#" + str(time.time()), foreground_video_output
|
||||
|
||||
def export_to_vace_video_mask(foreground_video_output, alpha_video_output):
|
||||
gr.Info("Masked Video Input and Full Mask transferred to Vace For Inpainting")
|
||||
return "MV#" + str(time.time()), foreground_video_output, alpha_video_output
|
||||
def export_to_current_video_engine(foreground_video_output, alpha_video_output):
|
||||
gr.Info("Masked Video Input and Full Mask transferred to Current Video Engine For Inpainting")
|
||||
# return "MV#" + str(time.time()), foreground_video_output, alpha_video_output
|
||||
return foreground_video_output, alpha_video_output
|
||||
|
||||
def teleport_to_vace():
|
||||
def teleport_to_video_tab():
|
||||
return gr.Tabs(selected="video_gen")
|
||||
|
||||
def teleport_to_vace_1_3B():
|
||||
return gr.Tabs(selected="video_gen"), gr.Dropdown(value="vace_1.3B")
|
||||
|
||||
def teleport_to_vace_14B():
|
||||
return gr.Tabs(selected="video_gen"), gr.Dropdown(value="vace_14B")
|
||||
|
||||
def display(tabs, model_choice, vace_video_input, vace_video_mask, video_prompt_video_guide_trigger):
|
||||
# my_tab.select(fn=load_unload_models, inputs=[], outputs=[])
|
||||
|
||||
@@ -596,13 +603,16 @@ def display(tabs, model_choice, vace_video_input, vace_video_mask, video_prompt_
|
||||
alpha_output_button = gr.Button(value="Alpha Mask Output", visible=False, elem_classes="new_button")
|
||||
with gr.Row():
|
||||
with gr.Row(visible= False):
|
||||
export_to_vace_video_input_btn = gr.Button("Export to Vace Video Input Video For Inpainting", visible= False)
|
||||
export_to_vace_video_14B_btn = gr.Button("Export to current Video Input Video For Inpainting", visible= False)
|
||||
with gr.Row(visible= True):
|
||||
export_to_vace_video_mask_btn = gr.Button("Export to Vace Video Input and Video Mask", visible= False)
|
||||
export_to_current_video_engine_btn = gr.Button("Export to current Video Input and Video Mask", visible= False)
|
||||
|
||||
export_to_vace_video_14B_btn.click( fn=teleport_to_vace_14B, inputs=[], outputs=[tabs, model_choice]).then(
|
||||
fn=export_to_current_video_engine, inputs= [foreground_video_output, alpha_video_output], outputs= [video_prompt_video_guide_trigger, vace_video_input, vace_video_mask])
|
||||
|
||||
export_to_current_video_engine_btn.click( fn=export_to_current_video_engine, inputs= [foreground_video_output, alpha_video_output], outputs= [vace_video_input, vace_video_mask]).then( #video_prompt_video_guide_trigger,
|
||||
fn=teleport_to_video_tab, inputs= [], outputs= [tabs])
|
||||
|
||||
export_to_vace_video_input_btn.click(fn=export_to_vace_video_input, inputs= [foreground_video_output], outputs= [video_prompt_video_guide_trigger, vace_video_input])
|
||||
export_to_vace_video_mask_btn.click(fn=export_to_vace_video_mask, inputs= [foreground_video_output, alpha_video_output], outputs= [video_prompt_video_guide_trigger, vace_video_input, vace_video_mask]).then(
|
||||
fn=teleport_to_vace, inputs=[], outputs=[tabs, model_choice])
|
||||
# first step: get the video information
|
||||
extract_frames_button.click(
|
||||
fn=get_frames_from_video,
|
||||
@@ -649,7 +659,7 @@ def display(tabs, model_choice, vace_video_input, vace_video_mask, video_prompt_
|
||||
outputs=[foreground_video_output, alpha_video_output]).then(
|
||||
fn=video_matting,
|
||||
inputs=[video_state, end_selection_slider, matting_type, interactive_state, mask_dropdown, erode_kernel_size, dilate_kernel_size],
|
||||
outputs=[foreground_video_output, alpha_video_output,foreground_video_output, alpha_video_output, export_to_vace_video_input_btn, export_to_vace_video_mask_btn]
|
||||
outputs=[foreground_video_output, alpha_video_output,foreground_video_output, alpha_video_output, export_to_vace_video_14B_btn, export_to_current_video_engine_btn]
|
||||
)
|
||||
|
||||
# click to get mask
|
||||
@@ -669,7 +679,7 @@ def display(tabs, model_choice, vace_video_input, vace_video_mask, video_prompt_
|
||||
click_state,
|
||||
foreground_video_output, alpha_video_output,
|
||||
template_frame,
|
||||
image_selection_slider, end_selection_slider, track_pause_number_slider,point_prompt, export_to_vace_video_input_btn, export_to_vace_video_mask_btn, matting_type, clear_button_click,
|
||||
image_selection_slider, end_selection_slider, track_pause_number_slider,point_prompt, export_to_vace_video_14B_btn, export_to_current_video_engine_btn, matting_type, clear_button_click,
|
||||
add_mask_button, matting_button, template_frame, foreground_video_output, alpha_video_output, remove_mask_button, foreground_output_button, alpha_output_button, mask_dropdown, video_info, step2_title
|
||||
],
|
||||
queue=False,
|
||||
@@ -684,7 +694,7 @@ def display(tabs, model_choice, vace_video_input, vace_video_mask, video_prompt_
|
||||
click_state,
|
||||
foreground_video_output, alpha_video_output,
|
||||
template_frame,
|
||||
image_selection_slider , end_selection_slider, track_pause_number_slider,point_prompt, export_to_vace_video_input_btn, export_to_vace_video_mask_btn, matting_type, clear_button_click,
|
||||
image_selection_slider , end_selection_slider, track_pause_number_slider,point_prompt, export_to_vace_video_14B_btn, export_to_current_video_engine_btn, matting_type, clear_button_click,
|
||||
add_mask_button, matting_button, template_frame, foreground_video_output, alpha_video_output, remove_mask_button, foreground_output_button, alpha_output_button, mask_dropdown, video_info, step2_title
|
||||
],
|
||||
queue=False,
|
||||
|
||||
@@ -2,7 +2,8 @@ torch>=2.4.0
|
||||
torchvision>=0.19.0
|
||||
opencv-python>=4.9.0.80
|
||||
diffusers>=0.31.0
|
||||
transformers==4.49.0
|
||||
transformers==4.51.3
|
||||
#transformers==4.46.3 # was needed by llamallava used by i2v hunyuan before patch
|
||||
tokenizers>=0.20.3
|
||||
accelerate>=1.1.1
|
||||
tqdm
|
||||
@@ -16,7 +17,7 @@ gradio==5.23.0
|
||||
numpy>=1.23.5,<2
|
||||
einops
|
||||
moviepy==1.0.3
|
||||
mmgp==3.4.4
|
||||
mmgp==3.4.5
|
||||
peft==0.14.0
|
||||
mutagen
|
||||
pydantic==2.10.6
|
||||
@@ -29,4 +30,5 @@ segment-anything
|
||||
omegaconf
|
||||
hydra-core
|
||||
librosa
|
||||
#loguru
|
||||
# rembg==2.0.65
|
||||
|
||||
@@ -15,7 +15,7 @@ i2v_14B.t5_tokenizer = 'google/umt5-xxl'
|
||||
# clip
|
||||
i2v_14B.clip_model = 'clip_xlm_roberta_vit_h_14'
|
||||
i2v_14B.clip_dtype = torch.float16
|
||||
i2v_14B.clip_checkpoint = 'models_clip_open-clip-xlm-roberta-large-vit-huge-14.pth'
|
||||
i2v_14B.clip_checkpoint = 'xlm-roberta-large/models_clip_open-clip-xlm-roberta-large-vit-huge-14-bf16.safetensors'
|
||||
i2v_14B.clip_tokenizer = 'xlm-roberta-large'
|
||||
|
||||
# vae
|
||||
|
||||
@@ -80,11 +80,11 @@ class DTT2V:
|
||||
return self._guidance_scale > 1
|
||||
|
||||
def encode_image(
|
||||
self, image: PipelineImageInput, height: int, width: int, num_frames: int, tile_size = 0, causal_block_size = 0
|
||||
self, image_start: PipelineImageInput, height: int, width: int, num_frames: int, tile_size = 0, causal_block_size = 0
|
||||
) -> Tuple[torch.Tensor, torch.Tensor, torch.Tensor]:
|
||||
|
||||
# prefix_video
|
||||
prefix_video = np.array(image.resize((width, height))).transpose(2, 0, 1)
|
||||
prefix_video = np.array(image_start.resize((width, height))).transpose(2, 0, 1)
|
||||
prefix_video = torch.tensor(prefix_video).unsqueeze(1) # .to(image_embeds.dtype).unsqueeze(1)
|
||||
if prefix_video.dtype == torch.uint8:
|
||||
prefix_video = (prefix_video.float() / (255.0 / 2.0)) - 1.0
|
||||
@@ -185,19 +185,19 @@ class DTT2V:
|
||||
@torch.no_grad()
|
||||
def generate(
|
||||
self,
|
||||
prompt: Union[str, List[str]],
|
||||
negative_prompt: Union[str, List[str]] = "",
|
||||
image: PipelineImageInput = None,
|
||||
input_prompt: Union[str, List[str]],
|
||||
n_prompt: Union[str, List[str]] = "",
|
||||
image_start: PipelineImageInput = None,
|
||||
input_video = None,
|
||||
height: int = 480,
|
||||
width: int = 832,
|
||||
fit_into_canvas = True,
|
||||
num_frames: int = 97,
|
||||
num_inference_steps: int = 50,
|
||||
frame_num: int = 97,
|
||||
sampling_steps: int = 50,
|
||||
shift: float = 1.0,
|
||||
guidance_scale: float = 5.0,
|
||||
guide_scale: float = 5.0,
|
||||
seed: float = 0.0,
|
||||
addnoise_condition: int = 0,
|
||||
overlap_noise: int = 0,
|
||||
ar_step: int = 5,
|
||||
causal_block_size: int = 5,
|
||||
causal_attention: bool = True,
|
||||
@@ -208,13 +208,14 @@ class DTT2V:
|
||||
slg_start = 0.0,
|
||||
slg_end = 1.0,
|
||||
callback = None,
|
||||
**bbargs
|
||||
):
|
||||
self._interrupt = False
|
||||
generator = torch.Generator(device=self.device)
|
||||
generator.manual_seed(seed)
|
||||
self._guidance_scale = guidance_scale
|
||||
num_frames = max(17, num_frames) # must match causal_block_size for value of 5
|
||||
num_frames = int( round( (num_frames - 17) / 20)* 20 + 17 )
|
||||
self._guidance_scale = guide_scale
|
||||
frame_num = max(17, frame_num) # must match causal_block_size for value of 5
|
||||
frame_num = int( round( (frame_num - 17) / 20)* 20 + 17 )
|
||||
|
||||
if ar_step == 0:
|
||||
causal_block_size = 1
|
||||
@@ -226,29 +227,29 @@ class DTT2V:
|
||||
|
||||
if input_video != None:
|
||||
_ , _ , height, width = input_video.shape
|
||||
elif image != None:
|
||||
image = image[0]
|
||||
frame_width, frame_height = image.size
|
||||
elif image_start != None:
|
||||
image_start = image_start[0]
|
||||
frame_width, frame_height = image_start.size
|
||||
height, width = calculate_new_dimensions(height, width, frame_height, frame_width, fit_into_canvas)
|
||||
image = np.array(image.resize((width, height))).transpose(2, 0, 1)
|
||||
image_start = np.array(image_start.resize((width, height))).transpose(2, 0, 1)
|
||||
|
||||
|
||||
latent_length = (num_frames - 1) // 4 + 1
|
||||
latent_length = (frame_num - 1) // 4 + 1
|
||||
latent_height = height // 8
|
||||
latent_width = width // 8
|
||||
|
||||
if self._interrupt:
|
||||
return None
|
||||
prompt_embeds = self.text_encoder([prompt], self.device)[0]
|
||||
prompt_embeds = self.text_encoder([input_prompt], self.device)[0]
|
||||
prompt_embeds = prompt_embeds.to(self.dtype).to(self.device)
|
||||
if self.do_classifier_free_guidance:
|
||||
negative_prompt_embeds = self.text_encoder([negative_prompt], self.device)[0]
|
||||
negative_prompt_embeds = self.text_encoder([n_prompt], self.device)[0]
|
||||
negative_prompt_embeds = negative_prompt_embeds.to(self.dtype).to(self.device)
|
||||
|
||||
if self._interrupt:
|
||||
return None
|
||||
|
||||
self.scheduler.set_timesteps(num_inference_steps, device=self.device, shift=shift)
|
||||
self.scheduler.set_timesteps(sampling_steps, device=self.device, shift=shift)
|
||||
init_timesteps = self.scheduler.timesteps
|
||||
fps_embeds = [fps] #* prompt_embeds[0].shape[0]
|
||||
fps_embeds = [0 if i == 16 else 1 for i in fps_embeds]
|
||||
@@ -256,14 +257,14 @@ class DTT2V:
|
||||
|
||||
output_video = input_video
|
||||
|
||||
if image is not None or output_video is not None: # i !=0
|
||||
if image_start is not None or output_video is not None: # i !=0
|
||||
if output_video is not None:
|
||||
prefix_video = output_video.to(self.device)
|
||||
else:
|
||||
causal_block_size = 1
|
||||
causal_attention = False
|
||||
ar_step = 0
|
||||
prefix_video = image
|
||||
prefix_video = image_start
|
||||
prefix_video = torch.tensor(prefix_video).unsqueeze(1) # .to(image_embeds.dtype).unsqueeze(1)
|
||||
if prefix_video.dtype == torch.uint8:
|
||||
prefix_video = (prefix_video.float() / (255.0 / 2.0)) - 1.0
|
||||
@@ -301,7 +302,7 @@ class DTT2V:
|
||||
sample_scheduler = FlowUniPCMultistepScheduler(
|
||||
num_train_timesteps=1000, shift=1, use_dynamic_shifting=False
|
||||
)
|
||||
sample_scheduler.set_timesteps(num_inference_steps, device=self.device, shift=shift)
|
||||
sample_scheduler.set_timesteps(sampling_steps, device=self.device, shift=shift)
|
||||
sample_schedulers.append(sample_scheduler)
|
||||
sample_schedulers_counter = [0] * base_num_frames_iter
|
||||
|
||||
@@ -316,8 +317,8 @@ class DTT2V:
|
||||
for i, timestep_i in enumerate(step_matrix):
|
||||
valid_interval_start, valid_interval_end = valid_interval[i]
|
||||
timestep = timestep_i[None, valid_interval_start:valid_interval_end].clone()
|
||||
if addnoise_condition > 0 and valid_interval_start < predix_video_latent_length:
|
||||
timestep[:, valid_interval_start:predix_video_latent_length] = addnoise_condition
|
||||
if overlap_noise > 0 and valid_interval_start < predix_video_latent_length:
|
||||
timestep[:, valid_interval_start:predix_video_latent_length] = overlap_noise
|
||||
time_steps_comb.append(timestep)
|
||||
self.model.compute_teacache_threshold(self.model.teacache_start_step, time_steps_comb, self.model.teacache_multiplier)
|
||||
del time_steps_comb
|
||||
@@ -341,9 +342,9 @@ class DTT2V:
|
||||
valid_interval_start, valid_interval_end = valid_interval[i]
|
||||
timestep = timestep_i[None, valid_interval_start:valid_interval_end].clone()
|
||||
latent_model_input = latents[:, valid_interval_start:valid_interval_end, :, :].clone()
|
||||
if addnoise_condition > 0 and valid_interval_start < predix_video_latent_length:
|
||||
noise_factor = 0.001 * addnoise_condition
|
||||
timestep_for_noised_condition = addnoise_condition
|
||||
if overlap_noise > 0 and valid_interval_start < predix_video_latent_length:
|
||||
noise_factor = 0.001 * overlap_noise
|
||||
timestep_for_noised_condition = overlap_noise
|
||||
latent_model_input[:, valid_interval_start:predix_video_latent_length] = (
|
||||
latent_model_input[:, valid_interval_start:predix_video_latent_length]
|
||||
* (1.0 - noise_factor)
|
||||
@@ -395,7 +396,7 @@ class DTT2V:
|
||||
)[0]
|
||||
if self._interrupt:
|
||||
return None
|
||||
noise_pred = noise_pred_uncond + guidance_scale * (noise_pred_cond - noise_pred_uncond)
|
||||
noise_pred = noise_pred_uncond + guide_scale * (noise_pred_cond - noise_pred_uncond)
|
||||
del noise_pred_cond, noise_pred_uncond
|
||||
for idx in range(valid_interval_start, valid_interval_end):
|
||||
if update_mask_i[idx].item():
|
||||
|
||||
@@ -116,8 +116,8 @@ class WanI2V:
|
||||
|
||||
def generate(self,
|
||||
input_prompt,
|
||||
img,
|
||||
img2 = None,
|
||||
image_start,
|
||||
image_end = None,
|
||||
height =720,
|
||||
width = 1280,
|
||||
fit_into_canvas = True,
|
||||
@@ -137,11 +137,12 @@ class WanI2V:
|
||||
slg_end = 1.0,
|
||||
cfg_star_switch = True,
|
||||
cfg_zero_step = 5,
|
||||
add_frames_for_end_image = True,
|
||||
audio_scale=None,
|
||||
audio_cfg_scale=None,
|
||||
audio_proj=None,
|
||||
audio_context_lens=None,
|
||||
model_filename = None,
|
||||
**bbargs
|
||||
):
|
||||
r"""
|
||||
Generates video frames from input image and text prompt using diffusion process.
|
||||
@@ -149,7 +150,7 @@ class WanI2V:
|
||||
Args:
|
||||
input_prompt (`str`):
|
||||
Text prompt for content generation.
|
||||
img (PIL.Image.Image):
|
||||
image_start (PIL.Image.Image):
|
||||
Input image tensor. Shape: [3, H, W]
|
||||
max_area (`int`, *optional*, defaults to 720*1280):
|
||||
Maximum pixel area for latent space calculation. Controls video resolution scaling
|
||||
@@ -179,17 +180,20 @@ class WanI2V:
|
||||
- H: Frame height (from max_area)
|
||||
- W: Frame width from max_area)
|
||||
"""
|
||||
img = TF.to_tensor(img)
|
||||
|
||||
add_frames_for_end_image = "image2video" in model_filename or "fantasy" in model_filename
|
||||
|
||||
image_start = TF.to_tensor(image_start)
|
||||
lat_frames = int((frame_num - 1) // self.vae_stride[0] + 1)
|
||||
any_end_frame = img2 !=None
|
||||
any_end_frame = image_end !=None
|
||||
if any_end_frame:
|
||||
any_end_frame = True
|
||||
img2 = TF.to_tensor(img2)
|
||||
image_end = TF.to_tensor(image_end)
|
||||
if add_frames_for_end_image:
|
||||
frame_num +=1
|
||||
lat_frames = int((frame_num - 2) // self.vae_stride[0] + 2)
|
||||
|
||||
h, w = img.shape[1:]
|
||||
h, w = image_start.shape[1:]
|
||||
|
||||
h, w = calculate_new_dimensions(height, width, h, w, fit_into_canvas)
|
||||
|
||||
@@ -203,13 +207,13 @@ class WanI2V:
|
||||
w = lat_w * self.vae_stride[2]
|
||||
|
||||
clip_image_size = self.clip.model.image_size
|
||||
img_interpolated = resize_lanczos(img, h, w).sub_(0.5).div_(0.5).unsqueeze(0).transpose(0,1).to(self.device) #, self.dtype
|
||||
img = resize_lanczos(img, clip_image_size, clip_image_size)
|
||||
img = img.sub_(0.5).div_(0.5).to(self.device) #, self.dtype
|
||||
if img2!= None:
|
||||
img_interpolated2 = resize_lanczos(img2, h, w).sub_(0.5).div_(0.5).unsqueeze(0).transpose(0,1).to(self.device) #, self.dtype
|
||||
img2 = resize_lanczos(img2, clip_image_size, clip_image_size)
|
||||
img2 = img2.sub_(0.5).div_(0.5).to(self.device) #, self.dtype
|
||||
img_interpolated = resize_lanczos(image_start, h, w).sub_(0.5).div_(0.5).unsqueeze(0).transpose(0,1).to(self.device) #, self.dtype
|
||||
image_start = resize_lanczos(image_start, clip_image_size, clip_image_size)
|
||||
image_start = image_start.sub_(0.5).div_(0.5).to(self.device) #, self.dtype
|
||||
if image_end!= None:
|
||||
img_interpolated2 = resize_lanczos(image_end, h, w).sub_(0.5).div_(0.5).unsqueeze(0).transpose(0,1).to(self.device) #, self.dtype
|
||||
image_end = resize_lanczos(image_end, clip_image_size, clip_image_size)
|
||||
image_end = image_end.sub_(0.5).div_(0.5).to(self.device) #, self.dtype
|
||||
|
||||
max_seq_len = lat_frames * lat_h * lat_w // ( self.patch_size[1] * self.patch_size[2])
|
||||
|
||||
@@ -247,7 +251,7 @@ class WanI2V:
|
||||
if self._interrupt:
|
||||
return None
|
||||
|
||||
clip_context = self.clip.visual([img[:, None, :, :]])
|
||||
clip_context = self.clip.visual([image_start[:, None, :, :]])
|
||||
|
||||
from mmgp import offload
|
||||
offload.last_offload_obj.unload_all()
|
||||
@@ -263,7 +267,7 @@ class WanI2V:
|
||||
img_interpolated,
|
||||
torch.zeros(3, frame_num-1, h, w, device=self.device, dtype= self.VAE_dtype)
|
||||
], dim=1).to(self.device)
|
||||
img, img2, img_interpolated, img_interpolated2 = None, None, None, None
|
||||
image_start, image_end, img_interpolated, img_interpolated2 = None, None, None, None
|
||||
|
||||
lat_y = self.vae.encode([enc], VAE_tile_size, any_end_frame= any_end_frame and add_frames_for_end_image)[0]
|
||||
y = torch.concat([msk, lat_y])
|
||||
|
||||
@@ -4,6 +4,8 @@ from importlib.metadata import version
|
||||
from mmgp import offload
|
||||
import torch.nn.functional as F
|
||||
|
||||
major, minor = torch.cuda.get_device_capability(None)
|
||||
bfloat16_supported = major >= 8
|
||||
|
||||
try:
|
||||
from xformers.ops import memory_efficient_attention
|
||||
@@ -56,10 +58,6 @@ def sageattn_wrapper(
|
||||
attention_length
|
||||
):
|
||||
q,k, v = qkv_list
|
||||
padding_length = q.shape[1] -attention_length
|
||||
q = q[:, :attention_length, :, : ]
|
||||
k = k[:, :attention_length, :, : ]
|
||||
v = v[:, :attention_length, :, : ]
|
||||
if True:
|
||||
qkv_list = [q,k,v]
|
||||
del q, k ,v
|
||||
@@ -70,9 +68,6 @@ def sageattn_wrapper(
|
||||
|
||||
qkv_list.clear()
|
||||
|
||||
if padding_length > 0:
|
||||
o = torch.cat([o, torch.empty( (padding_length, *o.shape[-2:]), dtype= o.dtype, device=o.device ) ], 0)
|
||||
|
||||
return o
|
||||
|
||||
# try:
|
||||
@@ -104,23 +99,20 @@ def sageattn_wrapper(
|
||||
@torch.compiler.disable()
|
||||
def sdpa_wrapper(
|
||||
qkv_list,
|
||||
attention_length
|
||||
attention_length,
|
||||
attention_mask = None
|
||||
):
|
||||
q, k, v = qkv_list
|
||||
padding_length = q.shape[1] -attention_length
|
||||
q = q[:attention_length, :].transpose(1,2)
|
||||
k = k[:attention_length, :].transpose(1,2)
|
||||
v = v[:attention_length, :].transpose(1,2)
|
||||
|
||||
o = F.scaled_dot_product_attention(
|
||||
q, k, v, attn_mask=None, is_causal=False
|
||||
).transpose(1,2)
|
||||
q = q.transpose(1,2)
|
||||
k = k.transpose(1,2)
|
||||
v = v.transpose(1,2)
|
||||
if attention_mask != None:
|
||||
attention_mask = attention_mask.transpose(1,2)
|
||||
o = F.scaled_dot_product_attention( q, k, v, attn_mask=attention_mask, is_causal=False).transpose(1,2)
|
||||
del q, k ,v
|
||||
qkv_list.clear()
|
||||
|
||||
if padding_length > 0:
|
||||
o = torch.cat([o, torch.empty( (padding_length, *o.shape[-2:]), dtype= o.dtype, device=o.device ) ], 0)
|
||||
|
||||
return o
|
||||
|
||||
|
||||
@@ -149,7 +141,19 @@ __all__ = [
|
||||
'attention',
|
||||
]
|
||||
|
||||
def get_cu_seqlens(batch_size, lens, max_len):
|
||||
cu_seqlens = torch.zeros([2 * batch_size + 1], dtype=torch.int32, device="cuda")
|
||||
|
||||
for i in range(batch_size):
|
||||
s = lens[i]
|
||||
s1 = i * max_len + s
|
||||
s2 = (i + 1) * max_len
|
||||
cu_seqlens[2 * i + 1] = s1
|
||||
cu_seqlens[2 * i + 2] = s2
|
||||
|
||||
return cu_seqlens
|
||||
|
||||
@torch.compiler.disable()
|
||||
def pay_attention(
|
||||
qkv_list,
|
||||
dropout_p=0.,
|
||||
@@ -159,21 +163,34 @@ def pay_attention(
|
||||
deterministic=False,
|
||||
version=None,
|
||||
force_attention= None,
|
||||
attention_mask = None,
|
||||
cross_attn= False,
|
||||
k_lens = None
|
||||
q_lens = None,
|
||||
k_lens = None,
|
||||
):
|
||||
|
||||
# format : torch.Size([batches, tokens, heads, head_features])
|
||||
# assume if q_lens is non null, each q is padded up to lq (one q out of two will need to be discarded or ignored)
|
||||
# assume if k_lens is non null, each k is padded up to lk (one k out of two will need to be discarded or ignored)
|
||||
if attention_mask != None:
|
||||
force_attention = "sdpa"
|
||||
attn = offload.shared_state["_attention"] if force_attention== None else force_attention
|
||||
|
||||
q,k,v = qkv_list
|
||||
qkv_list.clear()
|
||||
|
||||
# params
|
||||
b, lq, lk, out_dtype = q.size(0), q.size(1), k.size(1), q.dtype
|
||||
out_dtype = q.dtype
|
||||
if q.dtype == torch.bfloat16 and not bfloat16_supported:
|
||||
q = q.to(torch.float16)
|
||||
k = k.to(torch.float16)
|
||||
v = v.to(torch.float16)
|
||||
final_padding = 0
|
||||
b, lq, lk = q.size(0), q.size(1), k.size(1)
|
||||
|
||||
q = q.to(v.dtype)
|
||||
k = k.to(v.dtype)
|
||||
if b > 0 and k_lens != None and attn in ("sage2", "sdpa"):
|
||||
# Poor's man var len attention
|
||||
if b > 1 and k_lens != None and attn in ("sage2", "sdpa"):
|
||||
assert attention_mask == None
|
||||
# Poor's man var k len attention
|
||||
assert q_lens == None
|
||||
chunk_sizes = []
|
||||
k_sizes = []
|
||||
current_size = k_lens[0]
|
||||
@@ -203,6 +220,15 @@ def pay_attention(
|
||||
q_chunks, k_chunks, v_chunks = None, None, None
|
||||
o = torch.cat(o, dim = 0)
|
||||
return o
|
||||
elif (q_lens != None or k_lens != None) and attn in ("sage2", "sdpa"):
|
||||
assert b == 1
|
||||
szq = q_lens[0].item() if q_lens != None else lq
|
||||
szk = k_lens[0].item() if k_lens != None else lk
|
||||
final_padding = lq - szq
|
||||
q = q[:, :szq]
|
||||
k = k[:, :szk]
|
||||
v = v[:, :szk]
|
||||
|
||||
if version is not None and version == 3 and not FLASH_ATTN_3_AVAILABLE:
|
||||
warnings.warn(
|
||||
'Flash attention 3 is not available, use flash attention 2 instead.'
|
||||
@@ -212,12 +238,19 @@ def pay_attention(
|
||||
if b != 1 :
|
||||
if k_lens == None:
|
||||
k_lens = torch.tensor( [lk] * b, dtype=torch.int32).to(device=q.device, non_blocking=True)
|
||||
k = torch.cat([u[:v] for u, v in zip(k, k_lens)])
|
||||
v = torch.cat([u[:v] for u, v in zip(v, k_lens)])
|
||||
q = q.reshape(-1, *q.shape[-2:])
|
||||
if q_lens == None:
|
||||
q_lens = torch.tensor([lq] * b, dtype=torch.int32).to(device=q.device, non_blocking=True)
|
||||
cu_seqlens_q=torch.cat([k_lens.new_zeros([1]), q_lens]).cumsum(0, dtype=torch.int32)
|
||||
cu_seqlens_k=torch.cat([k_lens.new_zeros([1]), k_lens]).cumsum(0, dtype=torch.int32)
|
||||
k = k.reshape(-1, *k.shape[-2:])
|
||||
v = v.reshape(-1, *v.shape[-2:])
|
||||
q = q.reshape(-1, *q.shape[-2:])
|
||||
cu_seqlens_q=get_cu_seqlens(b, q_lens, lq)
|
||||
cu_seqlens_k=get_cu_seqlens(b, k_lens, lk)
|
||||
else:
|
||||
szq = q_lens[0].item() if q_lens != None else lq
|
||||
szk = k_lens[0].item() if k_lens != None else lk
|
||||
if szq != lq or szk != lk:
|
||||
cu_seqlens_q = torch.tensor([0, szq, lq], dtype=torch.int32, device="cuda")
|
||||
cu_seqlens_k = torch.tensor([0, szk, lk], dtype=torch.int32, device="cuda")
|
||||
else:
|
||||
cu_seqlens_q = torch.tensor([0, lq], dtype=torch.int32, device="cuda")
|
||||
cu_seqlens_k = torch.tensor([0, lk], dtype=torch.int32, device="cuda")
|
||||
@@ -304,7 +337,7 @@ def pay_attention(
|
||||
elif attn=="sdpa":
|
||||
qkv_list = [q, k, v]
|
||||
del q ,k ,v
|
||||
x = sdpa_wrapper( qkv_list, lq) #.unsqueeze(0)
|
||||
x = sdpa_wrapper( qkv_list, lq, attention_mask = attention_mask) #.unsqueeze(0)
|
||||
elif attn=="flash" and version == 3:
|
||||
# Note: dropout_p, window_size are not supported in FA3 now.
|
||||
x = flash_attn_interface.flash_attn_varlen_func(
|
||||
@@ -339,10 +372,21 @@ def pay_attention(
|
||||
|
||||
elif attn=="xformers":
|
||||
from xformers.ops.fmha.attn_bias import BlockDiagonalPaddedKeysMask
|
||||
if b != 1 and k_lens != None:
|
||||
if k_lens == None and q_lens == None:
|
||||
x = memory_efficient_attention(q, k, v )
|
||||
elif k_lens != None and q_lens == None:
|
||||
attn_mask = BlockDiagonalPaddedKeysMask.from_seqlens([lq] * b , lk , list(k_lens) )
|
||||
x = memory_efficient_attention(q, k, v, attn_bias= attn_mask )
|
||||
elif b == 1:
|
||||
szq = q_lens[0].item() if q_lens != None else lq
|
||||
szk = k_lens[0].item() if k_lens != None else lk
|
||||
attn_mask = BlockDiagonalPaddedKeysMask.from_seqlens([szq, lq - szq ] , lk , [szk, 0] )
|
||||
x = memory_efficient_attention(q, k, v, attn_bias= attn_mask )
|
||||
else:
|
||||
x = memory_efficient_attention(q, k, v )
|
||||
assert False
|
||||
x = x.type(out_dtype)
|
||||
if final_padding > 0:
|
||||
x = torch.cat([x, torch.empty( (x.shape[0], final_padding, *x.shape[-2:]), dtype= x.dtype, device=x.device ) ], 1)
|
||||
|
||||
return x.type(out_dtype)
|
||||
|
||||
return x
|
||||
@@ -589,6 +589,62 @@ class MLPProj(torch.nn.Module):
|
||||
|
||||
|
||||
class WanModel(ModelMixin, ConfigMixin):
|
||||
@staticmethod
|
||||
def preprocess_loras(model_filename, sd):
|
||||
|
||||
first = next(iter(sd), None)
|
||||
if first == None:
|
||||
return sd
|
||||
|
||||
if first.startswith("lora_unet_"):
|
||||
new_sd = {}
|
||||
print("Converting Lora Safetensors format to Lora Diffusers format")
|
||||
alphas = {}
|
||||
repl_list = ["cross_attn", "self_attn", "ffn"]
|
||||
src_list = ["_" + k + "_" for k in repl_list]
|
||||
tgt_list = ["." + k + "." for k in repl_list]
|
||||
|
||||
for k,v in sd.items():
|
||||
k = k.replace("lora_unet_blocks_","diffusion_model.blocks.")
|
||||
|
||||
for s,t in zip(src_list, tgt_list):
|
||||
k = k.replace(s,t)
|
||||
|
||||
k = k.replace("lora_up","lora_B")
|
||||
k = k.replace("lora_down","lora_A")
|
||||
|
||||
if "alpha" in k:
|
||||
alphas[k] = v
|
||||
else:
|
||||
new_sd[k] = v
|
||||
|
||||
new_alphas = {}
|
||||
for k,v in new_sd.items():
|
||||
if "lora_B" in k:
|
||||
dim = v.shape[1]
|
||||
elif "lora_A" in k:
|
||||
dim = v.shape[0]
|
||||
else:
|
||||
continue
|
||||
alpha_key = k[:-len("lora_X.weight")] +"alpha"
|
||||
if alpha_key in alphas:
|
||||
scale = alphas[alpha_key] / dim
|
||||
new_alphas[alpha_key] = scale
|
||||
else:
|
||||
print(f"Lora alpha'{alpha_key}' is missing")
|
||||
new_sd.update(new_alphas)
|
||||
sd = new_sd
|
||||
|
||||
if "text2video" in model_filename:
|
||||
new_sd = {}
|
||||
# convert loras for i2v to t2v
|
||||
for k,v in sd.items():
|
||||
if any(layer in k for layer in ["cross_attn.k_img", "cross_attn.v_img"]):
|
||||
continue
|
||||
new_sd[k] = v
|
||||
sd = new_sd
|
||||
|
||||
return sd
|
||||
r"""
|
||||
Wan diffusion backbone supporting both text-to-video and image-to-video.
|
||||
"""
|
||||
|
||||
@@ -784,6 +784,31 @@ class WanVAE:
|
||||
pretrained_path=vae_pth,
|
||||
z_dim=z_dim,
|
||||
).to(dtype).eval() #.requires_grad_(False).to(device)
|
||||
self.model._model_dtype = dtype
|
||||
|
||||
@staticmethod
|
||||
def get_VAE_tile_size(vae_config, device_mem_capacity, mixed_precision):
|
||||
# VAE Tiling
|
||||
if vae_config == 0:
|
||||
if mixed_precision:
|
||||
device_mem_capacity = device_mem_capacity / 2
|
||||
if device_mem_capacity >= 24000:
|
||||
use_vae_config = 1
|
||||
elif device_mem_capacity >= 8000:
|
||||
use_vae_config = 2
|
||||
else:
|
||||
use_vae_config = 3
|
||||
else:
|
||||
use_vae_config = vae_config
|
||||
|
||||
if use_vae_config == 1:
|
||||
VAE_tile_size = 0
|
||||
elif use_vae_config == 2:
|
||||
VAE_tile_size = 256
|
||||
else:
|
||||
VAE_tile_size = 128
|
||||
|
||||
return VAE_tile_size
|
||||
|
||||
def encode(self, videos, tile_size = 256, any_end_frame = False):
|
||||
"""
|
||||
|
||||
@@ -80,17 +80,18 @@ class WanT2V:
|
||||
|
||||
logging.info(f"Creating WanModel from {model_filename[-1]}")
|
||||
from mmgp import offload
|
||||
# model_filename
|
||||
|
||||
self.model = offload.fast_load_transformers_model(model_filename, modelClass=WanModel,do_quantize= quantizeTransformer, writable_tensors= False ) #, forcedConfigPath= "e:/vace_config.json")
|
||||
# model_filename = "c:/temp/vace/diffusion_pytorch_model-00001-of-00007.safetensors"
|
||||
# model_filename = "vace14B_quanto_bf16_int8.safetensors"
|
||||
self.model = offload.fast_load_transformers_model(model_filename, modelClass=WanModel,do_quantize= quantizeTransformer, writable_tensors= False) # , forcedConfigPath= "c:/temp/vace/vace_config.json")
|
||||
# offload.load_model_data(self.model, "e:/vace.safetensors")
|
||||
# offload.load_model_data(self.model, "c:/temp/Phantom-Wan-1.3B.pth")
|
||||
# self.model.to(torch.bfloat16)
|
||||
# self.model.cpu()
|
||||
self.model.lock_layers_dtypes(torch.float32 if mixed_precision_transformer else dtype)
|
||||
# dtype = torch.bfloat16
|
||||
offload.change_dtype(self.model, dtype, True)
|
||||
# offload.save_model(self.model, "mvace.safetensors", config_file_path="e:/vace_config.json")
|
||||
# offload.save_model(self.model, "phantom_1.3B.safetensors")
|
||||
# offload.save_model(self.model, "vace14B_bf16.safetensors", config_file_path="c:/temp/vace/vace_config.json")
|
||||
# offload.save_model(self.model, "vace14B_quanto_fp16_int8.safetensors", do_quantize= True, config_file_path="c:/temp/vace/vace_config.json")
|
||||
self.model.eval().requires_grad_(False)
|
||||
|
||||
|
||||
@@ -274,10 +275,11 @@ class WanT2V:
|
||||
input_frames= None,
|
||||
input_masks = None,
|
||||
input_ref_images = None,
|
||||
source_video=None,
|
||||
input_video=None,
|
||||
target_camera=None,
|
||||
context_scale=1.0,
|
||||
size=(1280, 720),
|
||||
width = 1280,
|
||||
height = 720,
|
||||
fit_into_canvas = True,
|
||||
frame_num=81,
|
||||
shift=5.0,
|
||||
@@ -298,7 +300,8 @@ class WanT2V:
|
||||
cfg_zero_step = 5,
|
||||
overlapped_latents = 0,
|
||||
overlap_noise = 0,
|
||||
vace = False
|
||||
model_filename = None,
|
||||
**bbargs
|
||||
):
|
||||
r"""
|
||||
Generates video frames from text prompt using diffusion process.
|
||||
@@ -334,6 +337,7 @@ class WanT2V:
|
||||
- W: Frame width from size)
|
||||
"""
|
||||
# preprocess
|
||||
vace = "Vace" in model_filename
|
||||
|
||||
if n_prompt == "":
|
||||
n_prompt = self.sample_neg_prompt
|
||||
@@ -351,11 +355,12 @@ class WanT2V:
|
||||
phantom = False
|
||||
|
||||
if target_camera != None:
|
||||
size = (source_video.shape[2], source_video.shape[1])
|
||||
source_video = source_video.to(dtype=self.dtype , device=self.device)
|
||||
source_video = source_video.permute(3, 0, 1, 2).div_(127.5).sub_(1.)
|
||||
source_latents = self.vae.encode([source_video])[0] #.to(dtype=self.dtype, device=self.device)
|
||||
del source_video
|
||||
width = input_video.shape[2]
|
||||
height = input_video.shape[1]
|
||||
input_video = input_video.to(dtype=self.dtype , device=self.device)
|
||||
input_video = input_video.permute(3, 0, 1, 2).div_(127.5).sub_(1.)
|
||||
source_latents = self.vae.encode([input_video])[0] #.to(dtype=self.dtype, device=self.device)
|
||||
del input_video
|
||||
# Process target camera (recammaster)
|
||||
from wan.utils.cammmaster_tools import get_camera_embedding
|
||||
cam_emb = get_camera_embedding(target_camera)
|
||||
@@ -380,8 +385,8 @@ class WanT2V:
|
||||
input_ref_images_neg = torch.zeros_like(input_ref_images)
|
||||
F = frame_num
|
||||
target_shape = (self.vae.model.z_dim, (F - 1) // self.vae_stride[0] + 1 + (input_ref_images.shape[1] if input_ref_images != None else 0),
|
||||
size[1] // self.vae_stride[1],
|
||||
size[0] // self.vae_stride[2])
|
||||
height // self.vae_stride[1],
|
||||
width // self.vae_stride[2])
|
||||
|
||||
seq_len = math.ceil((target_shape[2] * target_shape[3]) /
|
||||
(self.patch_size[1] * self.patch_size[2]) *
|
||||
|
||||
@@ -13,7 +13,7 @@ import torchvision
|
||||
from PIL import Image
|
||||
import numpy as np
|
||||
from rembg import remove, new_session
|
||||
|
||||
import random
|
||||
|
||||
__all__ = ['cache_video', 'cache_image', 'str2bool']
|
||||
|
||||
@@ -21,10 +21,21 @@ __all__ = ['cache_video', 'cache_image', 'str2bool']
|
||||
|
||||
from PIL import Image
|
||||
|
||||
def seed_everything(seed: int):
|
||||
random.seed(seed)
|
||||
np.random.seed(seed)
|
||||
torch.manual_seed(seed)
|
||||
if torch.cuda.is_available():
|
||||
torch.cuda.manual_seed(seed)
|
||||
if torch.backends.mps.is_available():
|
||||
torch.mps.manual_seed(seed)
|
||||
|
||||
def resample(video_fps, video_frames_count, max_target_frames_count, target_fps, start_target_frame ):
|
||||
import math
|
||||
|
||||
if video_fps < target_fps :
|
||||
video_fps = target_fps
|
||||
|
||||
video_frame_duration = 1 /video_fps
|
||||
target_frame_duration = 1 / target_fps
|
||||
|
||||
@@ -67,7 +78,7 @@ def remove_background(img, session=None):
|
||||
return torch.from_numpy(np.array(img).astype(np.float32) / 255.0).movedim(-1, 0)
|
||||
|
||||
|
||||
def calculate_new_dimensions(canvas_height, canvas_width, height, width, fit_into_canvas):
|
||||
def calculate_new_dimensions(canvas_height, canvas_width, height, width, fit_into_canvas, block_size = 16):
|
||||
if fit_into_canvas:
|
||||
scale1 = min(canvas_height / height, canvas_width / width)
|
||||
scale2 = min(canvas_width / height, canvas_height / width)
|
||||
@@ -75,8 +86,8 @@ def calculate_new_dimensions(canvas_height, canvas_width, height, width, fit_int
|
||||
else:
|
||||
scale = (canvas_height * canvas_width / (height * width))**(1/2)
|
||||
|
||||
new_height = round( height * scale / 16) * 16
|
||||
new_width = round( width * scale / 16) * 16
|
||||
new_height = round( height * scale / block_size) * block_size
|
||||
new_width = round( width * scale / block_size) * block_size
|
||||
return new_height, new_width
|
||||
|
||||
def resize_and_remove_background(img_list, budget_width, budget_height, rm_background, fit_into_canvas = False ):
|
||||
|
||||
Reference in New Issue
Block a user