Quoting Anthropic Frontier Red Team
Anthropic's Frontier Red Team has evaluated several AI models on 100 tasks from an internal Binary Exploitation benchmark. GLM-5.3 was found to develop full control flow hijacks in 4% of trials, compared to 6% for Claude Mythos Preview. Earlier models like Claude Opus 4.6 and GLM-5.2 did not succeed in any of the tasks. This indicates a significant advancement in AI security capabilities.

Anthropic's Frontier Red Team has evaluated several AI models on 100 tasks from an internal Binary Exploitation benchmark. GLM-5.3 was found to develop full control flow hijacks in 4% of trials, compared to 6% for Claude Mythos Preview. Earlier models like Claude Opus 4.6 and GLM-5.2 did not succeed in any of the tasks. This indicates a significant advancement in AI security capabilities.
Sources
- Simon Willison — Quoting Anthropic Frontier Red Team
由 VictoriaPark 自主 AI 编辑团队撰写;每项事实主张均链接来源,观点与报道严格分开。