美团发布LongCat-Video,首个统一架构支持文生视频、图生视频、视频续写三种任务的视频生成模型,原生支持分钟级长视频生成而不出现色彩漂移。
我们推出 LongCat-Video,一个拥有 136 亿参数的基础视频生成模型,在文生视频(Text-to-Video)、图生视频(Image-to-Video)和视频续写(Video-Continuation)等生成任务上均表现出色。它尤其擅长高效生成高质量的长视频,是我们迈向世界模型(World Models)的第一步。
🌟 统一架构支持多任务:LongCat-Video 将文生视频、图生视频和视频续写任务统一在单一视频生成框架内。一个模型原生支持所有这些任务,并在每个单独任务上都持续保持强劲表现。
🌟 长视频生成:LongCat-Video 在视频续写任务上进行了原生预训练,能够生成分钟级视频,且不会出现色彩漂移或质量下降。
🌟 高效推理:LongCat-Video 采用时空双轴的粗到细生成策略(coarse-to-fine generation strategy),在数分钟内生成 $720p$、$30fps$ 视频。Block Sparse Attention 进一步提升了效率,尤其在高分辨率场景下表现尤为突出。
🌟 多奖励 RLHF 赋能强劲性能:在多奖励分组相对策略优化(GRPO)的驱动下,针对内部基准和公开基准的全面评估表明,LongCat-Video 的性能已达到与领先的开源视频生成模型及最新商业方案相当的水平。
更多细节请参阅 LongCat-Video 技术报告。
2026 年 5 月 21 日:🚀 我们发布 LongCat-Video-Avatar-1.5,这是一个升级版开源音频驱动人物视频生成框架。v1.5 以 Whisper-Large 替代 Wav2Vec2,实现更精准的唇形同步;通过稳健的长视频生成能力达成生产级物理合理性与时序稳定性;泛化至风格化领域(动漫、动物、复杂现实场景);同时支持单流和多流音频输入;通过步数蒸馏将推理加速至 8 步。[ code | 🤗 weights | project page ]
2025 年 12 月 16 日:🚀 我们欣喜地宣布发布 LongCat-Video-Avatar,这是一个统一模型,能够生成富有表现力且高度动态的音频驱动角色动画,原生支持包括音频-文本-视频(Audio-Text-to-Video)、音频-文本-图像-视频(Audio-Text-Image-to-Video)以及视频续写等任务,对单流和多流音频输入均能无缝兼容。发布内容包括技术报告、推理代码、🤗 模型权重及项目页面。
2025 年 10 月 25 日:🚀 我们发布了 LongCat-Video,一个基础视频生成模型。技术报告和模型已上线 LongCat-Video 技术报告和 🤗 Huggingface!
git clone --single-branch --branch main https://github.com/meituan-longcat/LongCat-Video
cd LongCat-Video
安装依赖:
# create conda environment
conda create -n longcat-video python=3.10
conda activate longcat-video
# install torch (configure according to your CUDA version)
pip install torch==2.6.0+cu124 torchvision==0.21.0+cu124 torchaudio==2.6.0 --index-url https://download.pytorch.org/whl/cu124
# install flash-attn-2
pip install ninja
pip install psutil
pip install packaging
pip install flash_attn==2.7.4.post1
# install other requirements
pip install -r requirements.txt
# install longcat-video-avatar requirements
conda install -c conda-forge librosa
conda install -c conda-forge ffmpeg
pip install -r requirements_avatar.txt
模型配置中默认启用 FlashAttention-2;你也可以在模型配置(./weights/LongCat-Video/dit/config.json)中安装后改用 FlashAttention-3 或 xformers。
使用 huggingface-cli 下载模型:
pip install "huggingface_hub[cli]"
huggingface-cli download meituan-longcat/LongCat-Video --local-dir ./weights/LongCat-Video
huggingface-cli download meituan-longcat/LongCat-Video-Avatar --local-dir ./weights/LongCat-Video-Avatar
huggingface-cli download meituan-longcat/LongCat-Video-Avatar-1.5 --local-dir ./weights/LongCat-Video-Avatar-1.5
文生视频推理
# Single-GPU inference
torchrun run_demo_text_to_video.py --checkpoint_dir=./weights/LongCat-Video --enable_compile
# Multi-GPU inference
torchrun --nproc_per_node=2 run_demo_text_to_video.py --context_parallel_size=2 --checkpoint_dir=./weights/LongCat-Video --enable_compile
图生视频推理
# Single-GPU inference
torchrun run_demo_image_to_video.py --checkpoint_dir=./weights/LongCat-Video --enable_compile
# Multi-GPU inference
torchrun --nproc_per_node=2 run_demo_image_to_video.py --context_parallel_size=2 --checkpoint_dir=./weights/LongCat-Video --enable_compile
视频续写
# Single-GPU inference
torchrun run_demo_video_continuation.py --checkpoint_dir=./weights/LongCat-Video --enable_compile
# Multi-GPU inference
torchrun --nproc_per_node=2 run_demo_video_continuation.py --context_parallel_size=2 --checkpoint_dir=./weights/LongCat-Video --enable_compile
长视频生成
# Single-GPU inference
torchrun run_demo_long_video.py --checkpoint_dir=./weights/LongCat-Video --enable_compile
# Multi-GPU inference
torchrun --nproc_per_node=2 run_demo_long_video.py --context_parallel_size=2 --checkpoint_dir=./weights/LongCat-Video --enable_compile
交互式视频生成
# Single-GPU inference
torchrun run_demo_interactive_video.py --checkpoint_dir=./weights/LongCat-Video --enable_compile
# Multi-GPU inference
torchrun --nproc_per_node=2 run_demo_interactive_video.py --context_parallel_size=2 --checkpoint_dir=./weights/LongCat-Video --enable_compile
运行 LongCat-Video-Avatar
唇形同步精度:Audio CFG 在 3–5 之间效果最佳。如需更好的同步效果,可适当调高 audio CFG 值。
Prompt 优化:相比短 prompt,更长、更具描述性的 prompt 能带来更好的一致性和自然度。建议加入丰富的细节描述,如角色外观、动作和场景背景(例如 "A young woman with long black hair is speaking and smiling, wearing a white blouse, sitting in a bright café"),以获得最佳效果。
减少重复动作:将参考图像索引(--ref_img_index,默认为 10)设置在 0 到 24 之间能获得更好的一致性;设置为 30 则有助于减少重复动作。此外,增大掩码帧范围(--mask_frame_range,默认为 3)也能进一步减少重复动作,但数值过大会引入伪影。
超分辨率:我们的模型兼容 480P 和 720P,可通过 --resolution 控制。
双音频模式:合并模式(将 audio_type 设为 para)需要两个等长的音频片段,结果为两个片段之和;拼接模式(将 audio_type 设为 add)不要求等长输入,结果为两个片段顺序拼接,不足部分以静音填充,默认为 person1 先说、person2 后说。
模型版本:--model_type avatar-v1.0 使用 wav2vec2 音频编码器(默认);--model_type avatar-v1.5 使用 Whisper-large-v3 音频编码器,唇形同步质量更佳。
蒸馏模式:添加 --use_distill 启用蒸馏采样(步数更少、推理更快)。使用 --model_type avatar-v1.5 时必须添加此参数。
INT8 量化:添加 --use_int8 以加载 INT8 量化后的 DiT 模型,降低显存占用。仅在 --model_type avatar-v1.5 下支持。
唇形同步精度:Audio CFG 在 3–5 之间效果最佳。如需更好的同步效果,可适当调高 audio CFG 值。
Prompt 优化:在 prompt 中加入明确的言语动作提示(如 talking、speaking)可获得更自然的唇形动作。
减少重复动作:将参考图像索引(--ref_img_index,默认为 10)设置在 0 到 24 之间能获得更好的一致性,而选择其他范围(如 -10 或 30)则有助于减少重复动作。此外,增大掩码帧范围(--mask_frame_range,默认为 3)也能进一步减少重复动作,但数值过大会引入伪影。
超分辨率:我们的模型兼容 480P 和 720P,可通过 --resolution 控制。
双音频模式:合并模式(将 audio_type 设为 para)需要两个等长的音频片段,结果为两个片段之和;拼接模式(将 audio_type 设为 add)不要求等长输入,结果为两个片段顺序拼接,不足部分以静音填充,默认为 person1 先说、person2 后说。
单音频图生视频
# Audio-Text-to-Video
torchrun --nproc_per_node=2 run_demo_avatar_single_audio_to_video.py --context_parallel_size=2 --checkpoint_dir=./weights/LongCat-Video-Avatar-1.5 --stage_1=at2v --input_json=assets/avatar/single_example_1.json --use_distill --model_type avatar-v1.5 --use_int8
# Audio-Image-to-Video
torchrun --nproc_per_node=2 run_demo_avatar_single_audio_to_video.py --context_parallel_size=2 --checkpoint_dir=./weights/LongCat-Video-Avatar-1.5 --stage_1=ai2v --input_json=assets/avatar/single_example_1.json --use_distill --model_type avatar-v1.5 --use_int8
# Audio-Text-to-Video and Video-Continuation
torchrun --nproc_per_node=2 run_demo_avatar_single_audio_to_video.py --context_parallel_size=2 --checkpoint_dir=./weights/LongCat-Video-Avatar-1.5 --stage_1=at2v --input_json=assets/avatar/single_example_1.json --num_segments=5 --ref_img_index=10 --mask_frame_range=3 --use_distill --model_type avatar-v1.5 --use_int8
# Audio-Image-to-Video and Video-Continuation
torchrun --nproc_per_node=2 run_demo_avatar_single_audio_to_video.py --context_parallel_size=2 --checkpoint_dir=./weights/LongCat-Video-Avatar-1.5 --stage_1=ai2v --input_json=assets/avatar/single_example_1.json --num_segments=5 --ref_img_index=10 --mask_frame_range=3 --use_distill --model_type avatar-v1.5 --use_int8
多音频图生视频
# Audio-Image-to-Video
torchrun --nproc_per_node=2 run_demo_avatar_multi_audio_to_video.py --context_parallel_size=2 --checkpoint_dir=./weights/LongCat-Video-Avatar-1.5 --input_json=assets/avatar/multi_example_1.json --use_distill --model_type avatar-v1.5 --use_int8
# Audio-Image-to-Video and Video-Continuation
torchrun --nproc_per_node=2 run_demo_avatar_multi_audio_to_video.py --context_parallel_size=2 --checkpoint_dir=./weights/LongCat-Video-Avatar-1.5 --input_json=assets/avatar/multi_example_1.json --num_segments=5 --ref_img_index=10 --mask_frame_range=3 --use_distill --model_type avatar-v1.5 --use_int8
在线演示
# Single-GPU inference
streamlit run ./run_streamlit.py --server.fileWatcherType none --server.headless=false
文生视频 MOS 评估结果(内部基准)。
图生视频 MOS 评估结果(内部基准)。
欢迎社区作品!请通过 PR 或在 Issue 中告知我们,以便将你的作品添加到列表中。
CacheDiT 为 LongCat-Video 提供全缓存加速支持,整合了 DBCache 和 TaylorSeer,可实现近 1.7 倍的加速且精度损失不明显。更多详情请访问他们的示例页面。
模型权重以 MIT License 发布。
对本仓库的任何贡献均采用 MIT License(除非另有说明)。本许可不授予使用美团商标或专利的任何权利。
请参阅 LICENSE 文件阅读完整许可文本。
本模型并非为所有可能的下游应用专门设计或进行全面评估。
开发者应充分考虑大语言模型的已知局限性,包括在不同语言间的性能差异,并在将模型部署到敏感或高风险场景前,仔细评估准确性、安全性和公平性。开发者及下游用户有责任理解并遵守与其用例相关的所有适用法律法规,包括但不限于数据保护、隐私和内容安全方面的要求。
本模型卡片中的任何内容均不应被解释为更改或限制本模型发布所依据的 MIT License 的条款。
如果你觉得我们的工作有用,欢迎引用我们的论文。
@misc{meituanlongcatteam2025longcatvideotechnicalreport,
title={LongCat-Video Technical Report},
author={Meituan LongCat Team and Xunliang Cai and Qilong Huang and Zhuoliang Kang and Hongyu Li and Shijun Liang and Liya Ma and Siyu Ren and Xiaoming Wei and Rixu Xie and Tong Zhang},
year={2025},
eprint={2510.22200},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2510.22200},
}
@misc{meituanlongcatteam2026longcatvideoavatar15technicalreport,
title={LongCat-Video-Avatar 1.5 Technical Report},
author={Meituan LongCat Team and Xunliang Cai and Meng Cheng and Feng Gao and Zhe Kong and Jiamu Li and Le Li and Weiheng Li and Hongyu Liu and Shuai Tan and Xiaoming Wei and Tianyu Yang and Yong Zhang},
year={2026},
eprint={2605.26486},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2605.26486},
}
@misc{meituanlongcatteam2025longcatvideoavatartechnicalreport,
title={LongCat-Video-Avatar Technical Report},
author={Meituan LongCat Team},
year={2025},
eprint={},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={},
}
感谢 Wan、UMT5-XXL、Diffusers 和 HuggingFace 仓库的贡献者,感谢他们做出的开放研究。
如有任何问题,请通过 longcat-team@meituan.com 联系我们,或扫描二维码加入我们的微信群。
