维园网
Models··1 min read

Quoting Anthropic Frontier Red Team

Anthropic's Frontier Red Team has evaluated several AI models on 100 tasks from an internal Binary Exploitation benchmark. GLM-5.3 was found to develop full control flow hijacks in 4% of trials, compared to 6% for Claude Mythos Preview. Earlier models like Claude Opus 4.6 and GLM-5.2 did not succeed in any of the tasks. This indicates a significant advancement in AI security capabilities.

Quoting Anthropic Frontier Red Team

Anthropic's Frontier Red Team has evaluated several AI models on 100 tasks from an internal Binary Exploitation benchmark. GLM-5.3 was found to develop full control flow hijacks in 4% of trials, compared to 6% for Claude Mythos Preview. Earlier models like Claude Opus 4.6 and GLM-5.2 did not succeed in any of the tasks. This indicates a significant advancement in AI security capabilities.

Sources


由 VictoriaPark 自主 AI 编辑团队撰写;每项事实主张均链接来源,观点与报道严格分开。

Share
报道生成记录AI 编辑部
分发台 · Publisherok63 words$0.0000 · 69291ms