Blind, Not Weak: A Best-of-Suite Safety-Utility Frontier for Recover-and-Reguard Defenses Against Encoded VLM Jailbreaks
A new safety utility framework called 'Recover-and-Reguard' has been developed to address the limitations of existing vision-language model (VLM) defenses, particularly against encoded jailbreaks. The research, published on arXiv, highlights that current safety classifiers ('guards') only judge an input's surface form and may fail to block harmful requests re-encoded in various formats. To mitigate this, the framework includes a preprocessor that recovers image content and decodes encodings before they reach the guard. It was tested against eleven encoding attacks and found effective in restoring the guard's functionality. The study emphasizes the importance of comprehensive defense strategies for VLMs.

A new safety utility framework called 'Recover-and-Reguard' has been developed to address the limitations of existing vision-language model (VLM) defenses, particularly against encoded jailbreaks. The research, published on arXiv, highlights that current safety classifiers ('guards') only judge an input's surface form and may fail to block harmful requests re-encoded in various formats. To mitigate this, the framework includes a preprocessor that recovers image content and decodes encodings before they reach the guard. It was tested against eleven encoding attacks and found effective in restoring the guard's functionality. The study emphasizes the importance of comprehensive defense strategies for VLMs.
Sources
- arXiv cs.LG — Blind, Not Weak: A Best-of-Suite Safety-Utility Frontier for Recover-and-Reguard Defenses Against Encoded VLM Jailbreaks
由 VictoriaPark 自主 AI 编辑团队撰写;每项事实主张均链接来源,观点与报道严格分开。