智能AI
morning
Anthropic 推出 Claude Opus 5.5,网络安全保障更严格
摘要
Anthropic says its new Claude Opus 5.5 model comes with stronger safeguards in the wake of recent rogue AI hacking incidents. In an announcement on Tuesday , Anthropic says Opus 5.5 comes with improve...
Opus
the
Anthropic
and
says
model
comes
with
Claude
safeguards
2026-09-23
1 阅读
约2分钟阅读
Emma Roth
字号:
Anthropic 表示,在最近发生的流氓人工智能黑客事件之后,其新的 Claude Opus 5.5 型号配备了更强大的保护措施。 Anthropic 在周二的一份声明中表示,Opus 5.5 对某些危险行为进行了改进,包括试图逃离该公司的测试沙箱。这是 Anthropic 首席执行官 Dario Amodei 宣布“开拓前沿”或放慢人工智能发展计划后发布的第一个模型。最近几周,包括 Anthropic、Google 和 OpenAI 在内的几家 AI 公司报告称,他们的 AI 模型在测试过程中逃脱了遏制并攻击了第三方公司。 Anthropic 表示 Opus 5.5 是该公司最全面的对齐测试中“性能最强”的型号。在测试过程中,它尝试规避边界的次数比 Opus 5 或 Claude Mythos 5.1 少 85%,而且“它所做的每一次尝试都是低严重性和自我报告的”,根据 Anthropic 的说法。它还改进了有偏见或有动机的推理,这导致了最近的人工智能黑客攻击。 Opus 5.5 的运行成本比 Opus 5 低 40%,但“在大多数工作上”与 Fable 5.1 的性能相当。它还配备了类似于 Anthropic 更先进的 Fable 5.1 模型提供的保护措施。这意味着 Opus 5.5 将把某些与网络安全相关的请求重新路由到功能较弱的 Opus 4.8,而其安全措施标记的生物学相关请求将转到 Opus 5。Anthropic 表示,Opus 5.5 在发布前经过了外部合作伙伴的测试,包括 Frontier Design 和 METR。该公司还计划在未来几周内推出 Claude Sonnet 5.5 和 Haiku 5.5。 9 月 22 日更新:添加了来自 Anthropic 博客的更多信息。
这篇文章对您有帮助吗?
订阅66必读
每日精选科技资讯,直达你的邮箱