Back to News
RSS feedarxiv.org

How Fragile Is Safety in On-Device Language Models?

Summary

As small language models move onto resource-constrained devices and become components of agentic systems, protecting locally stored parameters becomes a safety concern. This study asks whether safety-sensitive behavior in LLaMA-2-7B-Chat is concentrated in a sparse subset of weights, which would create a smaller surface for targeted fault analysis. The authors use two complementary localization approaches: a low-rank safety-associated subspace analysis and parameter-level safety-utility importance filtering. Both reveal strongly non-uniform safety sensitivity across the network. The MLP down_proj consistently appears as the most prominent safety-sensitive component, while o_proj contributes less. In one parameter-level experiment, changing only 0.19% of down_proj weights produces a 53% Basic ASR and a 56% GCG ASR, while tinyBenchmarks accuracy remains 51.6%, close to the 52.2% unmodified baseline. The results suggest that on-device and agentic deployments may benefit from targeted fault analysis and selective integrity protection, although the study focuses on the examined model and localization methods.