AI & LLMs

Feature (SAE)

also: sparse feature · monosemantic feature

An interpretable direction in a model’s activations, pulled out by a sparse autoencoder as a single human-meaningful concept.

A feature, in mechanistic interpretability, is a direction in activation space that stands for one human-meaningful concept. A sparse autoencoder (SAE) decomposes a layer’s dense, polysemantic activations into a large dictionary of sparse features, each firing for a specific idea — the building blocks an attribution graph then wires together.