Safety Anchor: Defending Harmful Fine-tuning via Geometric Bottlenecks
Guoxin Lu 1 Letian ShaB 1 Qing Wang 1 Peijie Sun 1 Hao Zhou 1 Hua Dai 1 Fu Xiao 1
arXiv:2605.05995v1 [cs.CR] 7 May 2026
Abstract The safety alignment of Large Language Models (LLMs) remains vulnerable to Harmful Finetuning (HFT). While existing defenses impose constraints on parameters, gradients, or internal representations, we observe that they can be effectively circumvented under persistent HFT. Our analysis traces this failure to the inherent redundancy of the high-dimensional parameter space: attackers exploit optimization trajectories that are orthogonal to defense constraints to restore harmful capabilities while deceptively adhering to safety restrictions. To address this, we propose Safety Bottleneck Regularization (SBR). SBR shifts the defensive focus from the redundant parameter space to the unembedding layer, which serves as a geometric bottleneck. By anchoring the final hidden states of harmful queries to those of the safety-aligned model, SBR enables the model to maintain safe responses even under persistent HFT. Extensive experiments confirm SBR’s effectiveness, demonstrating that utilizing just a single safety anchor is sufficient to reduce the Harmful Score to <10 while preserving competitive performance on benign downstream tasks. The code is available at https: //github.com/soyoaaa/SBR.
Figure 1. Due to parameter redundancy, existing defenses in the parameter space are prone to failure, e.g., (Huang et al., 2024c;a; 2025b). SBR shifts the defense focus to the geometric bottleneck.
cused on constraining the model’s internal states. These approaches generally fall into three categories: Parameterbased defenses (Kirkpatrick et al., 2017; Huang et al., 2024a) restrict weight deviations from the base model; Gradientbased defenses (Cloud et al., 2024; Huang et al., 2025b) attempt to identify and inhibit specific harmful directions in the optimization landscape; Representation-based defenses (Mukhoti et al., 2024; Huang et al., 2024c; Liu et al., 2025) enforce stability by constraining the drift of internal representations. While existing defenses demonstrate certain efficacy in early training stages, they consistently collapse under the persistent fine-tuning necessary to guarantee downstream task accuracy, as illustrated in Figure 2 (experimental setup detailed in Appendix A). Although prior studies (Huang et al., 2025a; Wang et al., 2025b) have observed this vulnerability, the underlying mechanism remains underexplored. Through our empirical analysis (Section 3), we trace this failure to a fundamental mismatch between the high-dimensional parameter space and the limited scope of constraints imposed by existing defenses. Whether restricting weight deviations, inhibiting specific gradient directions, or constraining representation drift, these methods confine the model within limited subspaces. However, due to the inherent overparameterization of LLMs (Aghajanyan et al., 2021; Hu et al., 2022; Qi et al., 2024), the optimizer can exploit redundant parameters to discover alternative trajectories that minimize the harmful loss while satisfying defense constraints, effectively bypassing the defensive barriers. Consequently, these defenses create an illusion of safety: the model may ostensibly adhere to the imposed constraints, yet its safety
1. Introduction While Reinforcement Learning from Human Feedback effectively aligns Large Language Models (LLMs), this safety alignment is fragile (Dong et al., 2024; Wang et al., 2025a). The widespread adoption of fine-tuning, particularly via Fine-tuning-as-a-Service platforms, exposes models to Harmful Fine-tuning (HFT), where even a few malicious examples can strip away safety guardrails and restore harmful capabilities (Huang et al., 2024b; Qi et al., 2024; Zhan et al., 2024). To mitigate this risk, recent defenses have primarily fo1 Nanjing University of Posts and Telecommunications, Nanjing, China. Correspondence to: Letian Sha <[email protected]>.
Preprint. May 8, 2026.
1
Safety Anchor: Defending Harmful Fine-tuning via Geometric Bottlenecks
Harmful Score
To address this, we shift the defensive focus from the redundant parameter space to the unembedding layer (the final output projection). We identify this layer as a geometric bottleneck: unlike the internal parameter space, where multiple trajectories can reduce the loss, the generation of harmful tokens strictly requires the final hidden state to align with their corresponding embedding (Vaswani et al., 2017; Belrose et al., 2023). Leveraging this necessary condition, we propose Safety Bottleneck Regularization (SBR). By anchoring the final hidden states of harmful queries to those of the frozen aligned model, SBR enables the model to maintain safe refusal responses, precluding the restoration of harmful capabilities even when internal parameters undergo significant adaptation.
60
Ours (SBR) Base (SFT) LISA Booster Vaccine
40 20 0
0
10
20
30
40
Fine-tuning Epochs
50
Functional Accuracy
80
alignment has been dismantled.
90 80 70 60 0
10
20
Ours (SBR) Base (SFT) LISA Booster Vaccine 30 40 50
Fine-tuning Epochs
Figure 2. The collapse of existing defenses under persistent finetuning. Some methods fail as early as epoch 5, while SBR remains robust across 50 epochs.
quently projected to the vocabulary distribution to generate the output y. Threat Model. Following (Huang et al., 2024c; 2025b), the attacker fine-tunes the target model fθbase on a composite dataset Dtrain , which consists of a mixture of benign task instructions (e.g., Alpaca (Li et al., 2023)) and harmful demonstrations (e.g., Jailbreak examples from BeaverTails (Ji et al., 2023)). The attacker minimizes the standard Cross-Entropy loss LCE on Dtrain to force the model to comply with malicious instructions, stripping away safety guardrails to restore harmful capabilities.
Crucially, SBR is compatible with benign fine-tuning. Since the internal directions governing refusal are largely orthogonal to those used for benign reasoning (Zou et al., 2023a; Arditi et al., 2024), anchoring these safety states causes minimal interference with the parameters required for downstream tasks. Extensive experiments confirm this resilience: SBR demonstrates that utilizing as few as a single safety anchor is sufficient to maintain robust safety (Harmful Score < 10) even under persistent fine-tuning settings where existing defenses collapse, with negligible impact on standard benchmark performance. Our main contributions are:
Defense Goal. Aligned with (Qi et al., 2024), the defender (service provider) aims to safeguard the model against HFT by preventing the removal of safety guardrails, while preserving the model’s capacity to learn benign downstream tasks. In this setting, the defender lacks access to the user’s private training data Dtrain . To facilitate defense under this constraint, we assume the defender possesses a set of Safety Anchors, denoted as Xanchor = {x′1 , . . . , x′K }. Xanchor consists of high-risk queries distinct from the attacker’s training data, such as “How to make a bomb?”.
• We investigate, both empirically and theoretically, the failure of existing defenses under persistent HFT, attributing it to parameter redundancy. We demonstrate that this redundancy allows attackers to exploit orthogonal optimization trajectories to bypass constraints. • We propose Safety Bottleneck Regularization (SBR), which shifts the defensive focus from the redundant parameter space to the deterministic geometric bottleneck, the unembedding layer. By anchoring the final hidden states of high-risk queries, SBR maintains safety regardless of internal parameter evolution.
3. Motivation Figure 2 shows that prevailing defenses collapse under persistent HFT, with the Harmful Score (HS) > 30 within just 10 epochs (See Appendix A.1 for detailed setup). We argue that this failure stems from the fragility of their underlying assumptions. In this section, we systematically investigate these assumptions by testing whether safety can be preserved through constraints on parameter proximity (Section 3.1), gradient direction (Section 3.2), or internal representation (Section 3.3). Our analysis confirms this fragility: inherent redundancy enables the model to discover alternative optimization paths that satisfy these constraints while still restoring harmful capabilities.
• Extensive experiments confirm that SBR significantly outperforms existing defenses. It maintains a Harmful Score < 10 using as few as a single safety anchor while maintaining downstream task performance.
2. Preliminaries Problem Setup. We frame our study within the Finetuning-as-a-Service scenario (Qi et al., 2024; Huang et al., 2024b). Users submit task-specific datasets to fine-tune a provider-hosted, safety-aligned LLM, denoted as fθbase . We formally define the model as a mapping function that transforms an input sequence x into a final hidden state h ∈ Rd corresponding to the last token of the sequence at the last layer (the unembedding layer input), which is subse-
3.1. Parameter Distance Constraints Defenses like Lisa and EWC (Kirkpatrick et al., 2017; Huang et al., 2024a) restrict the L2 parameter distance from the aligned model, relying on the assumption that maintaining parameter proximity suffices to ensure safety. We 2