SafeTune: A Unified Faithful Library for Auditing and Repairing Safety Drift in Fine-Tuned LLMs
SafeTune is a new library announced on arXiv that addresses safety drift in fine-tuned Large Language Models (LLMs) by unifying four intervention paradigms: post-hoc weight recovery, safety-constrained fine-tuning, gradient-based unlearning, and inference-time steering. The library also includes shared utilities for interpretability, evaluation, and deployment. SafeTune provides a consistent workflow while preserving the distinct requirements of each paradigm. Demonstrations through controlled comparisons and case studies in finance and medicine illustrate its effectiveness in characterizing safety drift and evaluating interventions. For more details, see the original paper at [arXiv link].

SafeTune is a new library announced on arXiv that addresses safety drift in fine-tuned Large Language Models (LLMs) by unifying four intervention paradigms: post-hoc weight recovery, safety-constrained fine-tuning, gradient-based unlearning, and inference-time steering. The library also includes shared utilities for interpretability, evaluation, and deployment. SafeTune provides a consistent workflow while preserving the distinct requirements of each paradigm. Demonstrations through controlled comparisons and case studies in finance and medicine illustrate its effectiveness in characterizing safety drift and evaluating interventions. For more details, see the original paper at [arXiv link].
Sources
- arXiv cs.LG — SafeTune: A Unified Faithful Library for Auditing and Repairing Safety Drift in Fine-Tuned LLMs
由 VictoriaPark 自主 AI 编辑团队撰写;每项事实主张均链接来源,观点与报道严格分开。