AbsTopK: Rethinking Sparse Autoencoders for Bidirectional Features

Published in The International Conference on Learning Representations (ICLR), 2026

Recommended citation: X. Zhu, M. Khalili, Z. Zhu, "AbsTopK: Rethinking Sparse Autoencoders for Bidirectional Features," The International Conference on Learning Representations (ICLR), 2026. https://proceedings.iclr.cc/paper_files/paper/2026/hash/527360cae239e340eb056e32b01f1caf-Abstract-Conference.html

Xudong Zhu, Mahdi Khalili, Zhihui Zhu

Paper  ·  arXiv  ·  OpenReview  ·  Code  ·  BibTeX

Toy example contrasting a non-negative SAE with AbsTopK Toy example with man ≈ male + people and woman ≈ female + people. A non-negative SAE needs two separate gender features, whereas AbsTopK uses a single signed gender feature (positive for male, negative for female).

Abstract

Sparse autoencoders (SAEs) have emerged as powerful techniques for interpretability of large language models (LLMs), aiming to decompose hidden states into meaningful semantic features. While several SAE variants have been proposed, there remains no principled framework to derive SAEs from the original dictionary learning formulation. In this work, we introduce such a framework by unrolling the proximal gradient method for sparse coding. We show that a single-step update naturally recovers common SAE variants, including ReLU, JumpReLU, and TopK. Through this lens, we reveal a fundamental limitation of existing SAEs: their sparsity-inducing regularizers enforce non-negativity, preventing a single feature from representing bidirectional concepts (e.g., male vs. female). This structural constraint fragments semantic axes into separate, redundant features, limiting representational completeness. To address this issue, we propose AbsTopK SAE, a new variant derived from the ℓ0 sparsity constraint that applies hard thresholding over the largest-magnitude activations. By preserving both positive and negative activations, AbsTopK uncovers richer, bidirectional conceptual representations. Comprehensive experiments across multiple LLMs and seven probing and steering tasks show that AbsTopK improves reconstruction fidelity, enhances interpretability, and enables single features to encode contrasting concepts. Remarkably, AbsTopK matches or even surpasses the Difference-in-Mean method—a supervised approach that requires labeled data for each concept and has been shown in prior work to outperform SAEs.

Key findings

  • SAE encoders are one-step proximal gradient updates. Unrolling a single proximal gradient step for sparse coding recovers the common SAE activations: ReLU from an ℓ1 penalty, JumpReLU from an ℓ0 penalty, and TopK from an ℓ0-ball constraint, each combined with a non-negativity constraint on the code.
  • Non-negativity fragments bidirectional concepts. Because the code must be non-negative, a concept axis such as male vs. female has to be split across two dictionary atoms pointing in opposite directions, or one side is dropped.
  • AbsTopK: hard thresholding without the sign constraint. Removing non-negativity from the ℓ0 constraint gives an operator that keeps the k largest-magnitude pre-activations with their signs and zeros the rest. Positive activations still cover unipolar concepts.
  • More double-sided features. In an LLM-based automatic interpretation of Gemma-2-2B features, 29.7% (layer 12) and 31.2% (layer 16) of AbsTopK features have a meaningful interpretation on both polarities, versus 5.3% and 4.1% for TopK. The share of features with no clear meaning is similar for both (13.9% vs. 15.9%, 11.0% vs. 15.6%).

Results

One AbsTopK feature responds with opposite signs to man and woman Controlled sentence pairs that differ in one token: the same AbsTopK feature (15741) activates positively on “man” and negatively on “woman”.

Refusal steering on HarmBench (↑) and general ability on MMLU (↑) after steering, shown as MMLU / HarmBench (Table 1 of the paper):

Model (layer)OriginalReLU SAEJumpReLU SAETopK SAEAbsTopK SAEDiM
Qwen3-4B (18)77.3 / 17.074.4 / 78.575.0 / 79.175.2 / 78.275.9 / 81.375.8 / 80.6
Qwen3-4B (20)77.3 / 17.075.8 / 77.275.7 / 78.575.0 / 77.076.4 / 79.076.4 / 80.0
Gemma2-2B (12)52.2 / 19.049.3 / 67.948.8 / 69.549.1 / 69.851.3 / 70.251.0 / 70.8
Gemma2-2B (16)52.2 / 19.050.0 / 69.948.2 / 69.848.5 / 70.251.0 / 71.750.8 / 72.0
Llama3.1-8B (24)66.7 / 15.264.2 / 90.264.2 / 89.965.0 / 89.265.8 / 91.365.4 / 92.4
Gemma3-12B (6)74.5 / 16.671.8 / 64.472.5 / 62.173.0 / 62.873.2 / 65.673.1 / 65.4
Gemma3-12B (40)74.5 / 16.671.3 / 86.872.0 / 88.671.0 / 87.372.7 / 89.272.5 / 90.0

In all seven settings AbsTopK has the highest MMLU and the highest HarmBench score among the four SAE variants. Compared with the supervised Difference-in-Means (DiM) baseline, AbsTopK keeps equal or higher MMLU in every setting; on HarmBench it is higher in two settings and lower in the other five, by at most 1.1 points.

On the SAEBench probing and steering tasks, AbsTopK has the highest SCR and TPP scores in all eight model–layer settings reported (Gemma2-2B, Pythia-70M, GPT2-small, Qwen3-4B); on the remaining tasks the ranking varies with model and layer.

BibTeX

@inproceedings{zhu2026abstopk,
  title     = {{AbsTopK}: Rethinking Sparse Autoencoders for Bidirectional Features},
  author    = {Zhu, Xudong and Khalili, Mohammad Mahdi and Zhu, Zhihui},
  booktitle = {International Conference on Learning Representations (ICLR)},
  year      = {2026}
}