SAE features are often found to be interpretable but not useful for steering. SAEs are inherently trained for local re-construction, and most existing works study the local geometry and local effect of these features.
In our latest work, we study the downstream geometry of SAE features and try to understand why most often they aren’t useful as stable steering directions. We introduce an analysis framework, FEGA, to study this. In the process, we were also able to distinguish between two classes of features: value-like ones encoding concepts, and pointer-like features encoding functions operating on context supplied values.
We observe that consistent one-dimensional effects are rare across SAE variants. Value-like features more often produce structured low-dimensional effects, but usually across several directions, while pointer-like features predominantly produce diffuse effects.
This means that a feature can be interpretable and causally relevant without behaving like one reusable steering vector.
Check out our paper to learn more. Our project page also provides an interactive way to explore the nature of different features.
🌐 Project page:
ukplab.github.io/FEGA/
📄 Paper:
arxiv.org/abs/2607.24645
💻 Code:
github.com/UKPLab/FEGA
It was a great collaboration with
@UKPLab at
@TUDarmstadt . Kudos to my amazing co-authors: Phu Gia Hoang,
@Tanmoy_Chak ,
@IGurevych , and Subhabrata Dutta.