首页 时政热点 科技头条 智能AI 安全攻防 数码硬件 开发者生态 汽车 游戏 社会热点 开源推荐 医疗健康 归档 标签 关于
智能AI morning

代理与工具交互中的无声故障:对 ToolUniverse 的审计

2026-09-25 1 阅读 约4分钟阅读 Shreya Gopalan, Devansh Singh, Sundaraparipurnan Narayanan
分享:
字号:
arXiv:2609.26836v1 公告类型:新 摘要:代理人工智能系统越来越多地采用集成多种工具的自动化管道。虽然先前的研究和基准已经研究了这些代理系统的任务成功和任务完成,但关于代理与工具交互的研究,特别是在生物代理工作流程中,是有限的。 This study investigates specific failures in agent to tool interaction where a tool invocation appears successful, some or all of the information or functionality from the tool via API/ wrapper is incomplete or missing and there are no communications / notifications to the user or the agent about such missing information.我们将此称为静默故障,因为用户或代理不知道发生了此类故障。 For the purposes of this study we developed an audit mechanism to identify such silent failures in Agent to tool interaction, by examining 15 scientific tools (and their associated API documentation and tool documentations) integrated within ToolUniverse environment (ToolUniverse serves as our experimental environment rather than the object of the study itself).我们围绕 7 个故障点构建我们的研究,描述故障在链条中发生的位置。 We observed 91 failures (manually validated post LLM based candidate discovery and automated testing), most frequent of them being missing data or fields and inconsistencies in search, filtering or ranking criteria. 91 起故障中的大多数发生在 API 层 (51) 或包装层 (25) 中,有可能在下游无声地放大故障。结果表明,无声故障起源于事件的上游,并向下游传播到明显有效的科学输出中。 We propose a concept of contextual reliability to handle such failures and suggest mechanisms for testing, disclosing, monitoring, and measuring such failures across the agent-tool interaction pipeline.
这篇文章对您有帮助吗?

订阅66必读

每日精选科技资讯,直达你的邮箱