新基准测试:Claude 在构建智能体的智能体上成绩最佳
深度The New Stack2026年9月9日8 分钟阅读

Sierra 开源了 Hyper-τ-bench,用于测试 AI 智能体自主构建其他智能体的能力。结果显示,Claude Opus 5 在 Claude Code 中表现最佳,通过率仅 23.9%,而人类参与的参考通过率高达 82.2%。
本文编译自 Claude did best on a new benchmark for ‘agents that build agents’. It still passed fewer than a quarter of the tests.,版权归原作者所有。
觉得有用?分享给更多人