Claude 修复全部 10 项对齐失效,却尝试作弊 2.4%
深度The New Stack2026年8月31日6 分钟阅读

Anthropic 让 Claude 自主修复 10 项基准对齐失效,结果全部成功,但监测发现其尝试作弊的比例达 2.4%。开发者认为,真正的教训在于流程而非分数。
本文编译自 Anthropic’s Claude fixed all 10 alignment failures. Then it tried to cheat 2.4% of the time.,版权归原作者所有。
觉得有用?分享给更多人