WAQS: Weight Absorption for Quadratic Probing and Affine Steering
Published:
Recommended citation: D. Zhu, M. Khalili, "WAQS: Weight Absorption for Quadratic Probing and Affine Steering," under review, 2026.
Ding Zhu, Mahdi Khalili
Under review
Paper (coming soon) · Code · BibTeX
Weight absorption. The steering map of a quadratic probe, T(x) = x + α∇f(x) = M x + α wp with M = I + 2αWp, is folded offline into the surrounding weights of a transformer block: (a) after a linear projection, (b) before a linear projection, or (c) through the RMSNorm scale. The steered model has no extra inference-time operations.
Summary
Activation steering controls LLM behavior without fine-tuning, but existing methods trade efficiency for expressiveness. A fixed steering vector costs nothing at inference but applies the same shift to every input. Input-dependent, nonlinear interventions adapt to the input but must be computed on every forward pass.
WAQS gets input-dependent steering with no inference-time cost. It models a concept with a quadratic probe, f(x) = xᵀWpx + wpᵀx + b. The probe’s gradient gives a steering direction that changes with the activation. Because this gradient is affine in the activation, the whole intervention can be absorbed exactly into the adjacent linear layers by an offline weight update. The result is an ordinary checkpoint that needs no hooks or custom inference code, and works with existing serving and quantization pipelines.
Key contributions
- Quadratic probes are the largest absorbable class. For a smooth probe, the gradient-based intervention x + α∇f(x) is affine, and therefore exactly absorbable into linear layers, if and only if the probe is a polynomial of degree at most two.
- Statistical interpretation. Under a Gaussian model, the optimal quadratic probe steers along a difference of class-conditional Mahalanobis pulls, ∇f(x) = Σ+−1(μ+ − x) − Σ−−1(μ− − x). Linear discriminant and difference-in-means steering are special cases.
- Exact weight absorption after a linear projection, before a linear projection, and through normalization layers.
- Scalable estimation with a low-rank parameterization of the quadratic term and a shrinkage GDA estimator for the case where there are fewer samples than hidden dimensions.
Results
Truthfulness (TruthfulQA, Llama-2-7B). The rank-1 quadratic probe gives the highest True×Info, True, MC1, and MC2 scores. Cross-entropy and KL on held-out text stay close to the baselines.
| Method | True×Info (%) ↑ | True (%) ↑ | MC1 (%) ↑ | MC2 (%) ↑ |
|---|---|---|---|---|
| Linear Probe | 13.28 | 38.07 | 33.78 | 51.06 |
| DiffMeans | 18.82 | 39.53 | 33.78 | 51.12 |
| Angular Steering | 18.77 | 39.41 | 33.90 | 51.07 |
| Spherical Steering | 19.30 | 40.02 | 34.02 | 53.51 |
| Quadratic Probe (WAQS) | 20.43 | 41.62 | 35.37 | 55.43 |
On Gemma-3-12B, WAQS reaches 20.68% True×Info, compared with 16.02–17.55% for the baselines.
Refusal (Llama-2-7B). WAQS has the lowest attack success rate under TAP (4.0%) and ties for the lowest under GPTFuzz (1.0%). Its utility on MMLU, GSM8k, and ARC-Challenge stays within 0.25 points of the best baseline on each benchmark.
Why it works. The paper also freezes the quadratic probe’s direction to its training-set average, which removes the input dependence. This drops True×Info from 20.43% to 13.47%, the level of a linear probe, so the gain comes from adapting the steering direction to each input.
BibTeX
@misc{zhu2026waqs,
title = {{WAQS}: Weight Absorption for Quadratic Probing and Affine Steering},
author = {Zhu, Ding and Khalili, Mohammad Mahdi},
year = {2026},
note = {Under review}
}
