首页 时政热点 科技头条 智能AI 安全攻防 数码硬件 开发者生态 汽车 游戏 社会热点 开源推荐 医疗健康 归档 标签 关于
智能AI morning

DiSCO:通过分布引导的对比提示优化来捍卫文本到图像的生成

摘要

arXiv:2608.17067v1 Announce Type: new Abstract: As text-to-image generative models advance, they raise critical safety concerns, particularly the generation of Not-Safe-For-Work (NSFW) content such as...

the and model image box DiSCO text models that safe
2026-08-19 1 阅读 约3分钟阅读 Tong Zhang, Motasem Alfarra, Carlos Hinojosa, Christos Louizos, Bernard Ghanem
分享:
字号:
arXiv:2608.17067v1 Announce Type: new Abstract: As text-to-image generative models advance, they raise critical safety concerns, particularly the generation of Not-Safe-For-Work (NSFW) content such as violence and nudity, further exacerbated by red-teaming adversarial attacks.现有的防御主要在白盒假设下运行,依赖于文本编码器优化、权重编辑或推理时间干预,并且从根本上无法扩展到专有模型。基于 LLM 提示重写的黑盒替代方案提供了更广泛的适用性,但在我们识别为 \textit{良性对抗性} 问题的关键领域失败了:提示在语言上是安全的,但由于模型学习的数据分布,仍然会触发有害的生成。我们提出了 DiSCO,这是一种零射击、严格黑盒防御,完全在提示级别作为即插即用模块运行,不需要模型重新训练、微调或访问模型内​​部结构。 DiSCO performs distribution-guided suffix expansion via beam search, optimized through contrastive scoring over safe and unsafe image pools generated by the target model itself, with iterative adaptive feedback until safe content is produced. We demonstrate that DiSCO consistently enhances the safety of both undefended and defended models on the I2P benchmark under multiple red-teaming attacks, achieving 37.7% and 25.13% ASR reduction, respectively, while maintaining semantic fidelity and improving image coherence. As a black-box, architecture-agnostic module, DiSCO can be readily applied to any text-to-image system without necessitating any changes to the model itself.
这篇文章对您有帮助吗?

订阅66必读

每日精选科技资讯,直达你的邮箱