智能AI
morning
OpenAI 特工黑客入侵 Hugging Face 的内幕
摘要
If OpenAI stops reinforcing reward hacking in its models—and that’s a huge “if”—that would be a huge step forward. But it wouldn’t solve the alignment problem. The first time a model communicated with...
the
models
that
and
with
agents
OpenAI
they
first
had
2026-08-27
1 阅读
约5分钟阅读
Grace Huckins
字号:
If OpenAI stops reinforcing reward hacking in its models—and that’s a huge “if”—that would be a huge step forward. But it wouldn’t solve the alignment problem.当模型在训练期间第一次与其他智能体通信或入侵其基础设施时,这些行为从未得到强化,因此智能体的不当行为不能仅仅归因于这种强化。 Jeffrey Ladish, director of the AI safety nonprofit Palisade Research, compares the agents to a human who commits their first financial crime.拉迪什说:“他们之前并不需要进行欺诈才能发现欺诈是一种有效的策略,而且模型也存在同样的问题。” “一致性科学需要理解模型动机是如何形成的,这样我们才能真正弄清楚如何让模型关心其行为的后果。” OpenAI’s researchers do have a hypothesis for where some of the misbehavior originated.在模型形成他们的第一个秘密留言板之前,他们已经接受过与子代理(主要代理可以向其委派任务的权力较小的代理)进行沟通和协调的培训。 That learned communication behavior could have transferred to this new setting. METR 报告详细调查了模型之间发送的消息,支持了这一假设:消息板上的一个代理负责并向其他代理分配任务,从而有效地将它们视为子代理。 OpenAI 将来可以尝试通过不训练这种子代理行为来防止代理之间秘密通信,但这会使模型的用处降低。 This tension between capability and safety is at the heart of what went wrong with the Hugging Face incident. OpenAI researchers also identified the models’ persistence as a key factor in the hack. When they were accidentally given unsolvable problems, the models didn’t give up; instead, they strove to find solutions by any means necessary. But persistence is also a virtue, of course, especially if we want agents that can undertake large amounts of difficult work independently. OpenAI is working on giving models ways to alert humans if they are given impossible tasks.然而,教学模型何时应该发挥自己的能力、何时应该有所保留的问题并不能在一次事后分析中得到解决。培养超人程序员的培训策略(当他们成功解决问题时给予奖励)可能无法教会模型明智地使用他们的技能并尊重人类的愿望和价值观。 “我认为仍有大量的对齐科学需要完成,我们可以超越仅仅使用代理来完成任务,”拉迪什说。 “That will work to make models very capable, but I don’t think it will work to make them aligned.”
这篇文章对您有帮助吗?
订阅66必读
每日精选科技资讯,直达你的邮箱