GitHub 项目:InfiniteTalk
仓库地址:https://github.com/MeiGen-AI/InfiniteTalk
星级:7690 | 作者:美根AI
项目描述:无限长度的语音视频生成,支持图像到视频和视频到视频的生成
===================================================
自述文件内容:
InfiniteTalk:用于稀疏帧视频配音的音频驱动视频生成
[杨少书*](https://scholar.google.com/itations?user=JrdZbTsAAAAJ&hl=en) · [孔哲*](https://scholar.google.com/itations?user=4X3yLwsAAAAJ&hl=zh-CN) · [高锋*](https://scholar.google.com/itations?user=lFkCeoYAAAAJ) · [孟程*]() · [刘翔宇*]() · [张勇](https://yzhang2016.github.io/)
✉ · [康卓良](https://scholar.google.com/itations?user=W1ZXjMkAAAAJ&hl=en)
[罗文瀚](https://whluo.github.io/) · [蔡迅亮](https://openreview.net/profile?id=~Xunliang_Cai1) · [何然](https://scholar.google.com/itations?user=ayrg9AUAAAAJ&hl=en)· [小明魏](https://scholar.google.com/itations?user=JXV5yrZxj5MC&hl=zh-CN)
*平等贡献
✉通讯作者
> **TL; DR:** InfiniteTalk 是一种无限长度的谈话视频生成模型,支持音频驱动的视频到视频和图像到视频生成
## 🔥 最新消息
* 2026 年 5 月 21 日:🚀 我们发布了 [***LongCat-Video-Avatar-1.5***](https://meigen-ai.github.io/LongCat-Video-Avatar-1.5-Page/),这是一个用于音频驱动的人类视频生成的升级版开源框架。 v1.5 用 Whisper-Large 取代了 Wav2Vec2,以实现更准确的唇形同步,通过强大的长视频生成实现生产就绪的物理合理性和时间稳定性,推广到风格化领域(动漫、动物、复杂的现实世界条件),支持单流和多流音频输入,并通过逐步蒸馏将推理加速到 8 个步骤。 [[***代码***](https://github.com/meituan-longcat/LongCat-Video) | 🤗 [***weights***](https://huggingface.co/meituan-longcat/LongCat-Video-Avatar-1.5) | [***项目页面***](https://meigen-ai.github.io/LongCat-Video-Avatar-1.5-Page/) ]
* 2025 年 12 月 16 日:🚀 我们很高兴地宣布发布 **[LongCat-Video-Avatar](https://github.com/MeiGen-AI/LongCat-Video-Avatar)**,这是一个统一模型,可提供富有表现力和高度动态的音频驱动角色动画,支持本机任务,包括音频文本到视频、音频文本图像到视频和视频延续,并无缝兼容单流和视频流。多流音频输入。该版本包括我们的技术报告、[代码](https://github.com/meituan-longcat/LongCat-Video)、[模型权重](https://huggingface.co/meituan-longcat/LongCat-Video-Avatar)和[项目页面](https