[](https://arxiv.org/abs/2502.04896)
[](https://saiyan-world.github.io/goku/)
> [陈首发](https://www.shoufachen.com)、[葛崇健](https://chongjiane.github.io/)、[张雨绮](https://scholar.google.com/itations?user=7FlkVy8AAAAJ)、[张一达](https://openreview.net/profile?id=~Yida_Zhang2)、[冯达朱](https://www.zhufengda.net/)、[杨浩](https://github.com/haoy945)、[郝鸿翔](https://scholar.google.com/itations?user=173GpBQAAAAJ&hl=zh-CN)、[吴辉](https://github.com/whlook)、[赖志超](https://github.com/lazychao)、[一飞胡](https://openreview.net/profile?id=~Yifei_Hu3)、[林婷车](https://github.com/tcl326)、[张世龙](https://jshilong.github.io/)、[李富](https://scholar.google.com/itations?user=2A7_3hoAAAAJ&hl=en)、[川李](https://www.linkedin.com/in/chuanli1101/)、[王兴](https://www.linkedin.com/in/xing-wang-49369620/)、[彭杨华](https://scholar.google.com/itations?user=Gf9amnoAAAAJ&hl=en)、[孙佩泽](https://peizesun.github.io/)、[Ping罗](http://luoping.me/)、[姜一](https://scholar.google.com/itations?user=6dikuoYAAAAJ&hl=en)、[袁泽欢](https://shallowyuan.github.io/)、[彭冰月](https://www.linkedin.com/in/bingyp)、[Xiaobing刘](https://scholar.google.com/itations?user=1ypDmDwAAAAJ&hl=en) >
香港大学、字节跳动
## 概述 Goku 是基于整流流 Transformer 的新系列图像和视频联合生成模型。它旨在实现工业级性能,集成了高质量视觉生成的先进技术,包括细致的数据管理、模型设计和流程制定。 主要贡献包括: - 📊 高质量细粒度图像和视频数据管理。 - 🔄 开创性地使用整流流来增强视频和图像令牌之间的交互。 - 🌟 在图像和视频生成任务中具有卓越的定性和定量性能。 悟空支持多生成任务: - 🎬 **文本到视频生成** - 🖼️ **图像到视频生成** - 🎨 **文本到图像生成** ## 性能基准🏅 悟空在主要基准测试中取得最高分: - GenEval(文本到图像生成)上的 **0.76** - DPG-Bench 上的 **83.65**(文本到图像生成) - VBench 上的 **84.85**(文本到视频生成) ### VBench 性能🏆 Goku-T2V 在 VBench 中取得了令人印象深刻的 **84.85** 分数,截至 2024 年 10 月 7 日稳居第二,超越了几种领先的商业文本到视频模型。 |我