When Pruning Meets Interpretability: Preserving Sparse Autoencoder Robustness in LLMs

Published in The Conference on Language Modeling (COLM), 2026

Recommended citation: S. Gupte, X. Zhang, M. Khalili, "When Pruning Meets Interpretability: Preserving Sparse Autoencoder Robustness in LLMs," The Conference on Language Modeling (COLM), 2026. https://arxiv.org/abs/2608.25941

Suchit Gupte, Xueru Zhang, Mohammad Mahdi Khalili

Paper (arXiv)  ·  Code  ·  BibTeX

Layer sensitivity profile of SAEs under pruning Layer sensitivity of Gemma Scope SAEs on gemma-2-2b under pruning, averaged over pruning methods and SAEBench metrics. Middle layers (9–17) degrade the most.

Abstract

Sparse autoencoders (SAEs) are widely used to interpret the internal representations of large language models (LLMs), yet their reliability under post-hoc model compression remains poorly understood. We present a systematic study of how pruning affects SAE behavior and theoretically show that, for a fixed SAE, its impact is governed by perturbation energy, a covariance-weighted norm. This perspective exposes a key limitation of magnitude pruning: by ignoring activation geometry, it distorts the learned representation space and degrades SAE functionality. Activation-aware methods such as Wanda and SparseGPT, in contrast, implicitly control perturbation energy and are therefore substantially more robust at preserving SAE behavior. We further reveal a consistent structural vulnerability across all pruning methods: middle layers are significantly more sensitive to pruning than early or late layers. Guided by this insight, we propose a layer-wise sparsity allocation strategy, achieving lower perplexity under the same average pruning sparsity. Experiments across four model architectures validate our theoretical findings.

Key findings

  • Perturbation energy governs SAE degradation. For a fixed SAE, the damage from pruning a weight matrix W is controlled by ε² = tr(ΔW Σ ΔWᵀ), where Σ is the input activation covariance. Magnitude pruning, Wanda, and SparseGPT minimize successively better approximations of this quantity (Σ replaced by I, then diag(Σ), then the full Σ), which explains the ordering SparseGPT ≻ Wanda ≻ Magnitude.
  • Aggregate metrics hide damage. Magnitude-pruned middle layers keep KL-divergence scores ≥ 0.96 while intervention metrics (SCR, TPP) collapse, so reconstruction fidelity alone cannot certify that an SAE is still reliable.
  • Middle layers are the most fragile across all pruning methods, metrics, and sparsity levels, driven by perturbation accumulated from upstream layers.
  • Layer-wise sparsity allocation. A ramp schedule that protects early layers (20% sparsity rising to 60% at the middle layers) lowers perplexity at the same 50% average sparsity.

Results

SAEBench metrics under pruning, per layer Percent change from the dense baseline for eight SAEBench metrics across all 26 layers of gemma-2-2b at 50% sparsity. Magnitude pruning (red) degrades SAEs far more than the activation-aware methods.

Perplexity of gemma-2-2b at 50% average sparsity:

MethodUniform sparsityLayer-wise sparsity
Wanda212148
SparseGPT14198

The findings hold across four model and SAE pairs: pythia-70m (SAELens), gemma-2-2b and gemma-2-9b (Gemma Scope), and mistral-7b.

BibTeX

@inproceedings{gupte2026pruning,
  title     = {When Pruning Meets Interpretability: Preserving Sparse Autoencoder Robustness in {LLM}s},
  author    = {Gupte, Suchit and Zhang, Xueru and Khalili, Mohammad Mahdi},
  booktitle = {Conference on Language Modeling (COLM)},
  year      = {2026}
}