huggingface · Mar 24, 2025
diff --git a/‎docs/source/en/api/pipelines/hunyuan_video.md
+2-1 b/‎docs/source/en/api/pipelines/hunyuan_video.md
+2-1
diff --git a/‎scripts/convert_hunyuan_video_to_diffusers.py
+22-1 b/‎scripts/convert_hunyuan_video_to_diffusers.py
+22-1
@@ -50,7 +50,8 @@ The following models are available for the image-to-video pipeline:
 | Model name | Description |
 |:---|:---|
 | [`Skywork/SkyReels-V1-Hunyuan-I2V`](https://huggingface.co/Skywork/SkyReels-V1-Hunyuan-I2V) | Skywork's custom finetune of HunyuanVideo (de-distilled). Performs best with `97x544x960` resolution. Performs best at `97x544x960` resolution, `guidance_scale=1.0`, `true_cfg_scale=6.0` and a negative prompt. |
-| [`hunyuanvideo-community/HunyuanVideo-I2V`](https://huggingface.co/hunyuanvideo-community/HunyuanVideo-I2V) | Tecent's official HunyuanVideo I2V model. Performs best at resolutions of 480, 720, 960, 1280. A higher `shift` value when initializing the scheduler is recommended (good values are between 7 and 20) |
+| [`hunyuanvideo-community/HunyuanVideo-I2V-33ch`](https://huggingface.co/hunyuanvideo-community/HunyuanVideo-I2V) | Tecent's official HunyuanVideo 33-channel I2V model. Performs best at resolutions of 480, 720, 960, 1280. A higher `shift` value when initializing the scheduler is recommended (good values are between 7 and 20). |
+| [`hunyuanvideo-community/HunyuanVideo-I2V`](https://huggingface.co/hunyuanvideo-community/HunyuanVideo-I2V) | Tecent's official HunyuanVideo 16-channel I2V model. Performs best at resolutions of 480, 720, 960, 1280. A higher `shift` value when initializing the scheduler is recommended (good values are between 7 and 20) |
 
 ## Quantization
 
 
@@ -160,8 +160,9 @@ def remap_single_transformer_blocks_(key, state_dict):
         "pooled_projection_dim": 768,
         "rope_theta": 256.0,
         "rope_axes_dim": (16, 56, 56),
+        "image_condition_type": None,
     },
-    "HYVideo-T/2-I2V": {
+    "HYVideo-T/2-I2V-33ch": {
         "in_channels": 16 * 2 + 1,
         "out_channels": 16,
         "num_attention_heads": 24,
@@ -178,6 +179,26 @@ def remap_single_transformer_blocks_(key, state_dict):
         "pooled_projection_dim": 768,
         "rope_theta": 256.0,
         "rope_axes_dim": (16, 56, 56),
+        "image_condition_type": "latent_concat",
+    },
+    "HYVideo-T/2-I2V-16ch": {
+        "in_channels": 16,
+        "out_channels": 16,
+        "num_attention_heads": 24,
+        "attention_head_dim": 128,
+        "num_layers": 20,
+        "num_single_layers": 40,
+        "num_refiner_layers": 2,
+        "mlp_ratio": 4.0,
+        "patch_size": 2,
+        "patch_size_t": 1,
+        "qk_norm": "rms_norm",
+        "guidance_embeds": True,
+        "text_embed_dim": 4096,
+        "pooled_projection_dim": 768,
+        "rope_theta": 256.0,
+        "rope_axes_dim": (16, 56, 56),
+        "image_condition_type": "token_replace",
     },
 }