Dictionaries Without Ontologies: What a Sparse Autoencoder Latent Is Evidence For, Why Decodability, Use and Reproducibility Do Not Add Up to Canonicity, and the Identifiability Premise No Real-Model Result Has Checked
Abstract
Sparse autoencoders are now the default instrument for decomposing language model activations, and their latents are routinely reported as the model's features: units the network itself uses, which a wide enough dictionary has pulled out of superposition. This paper asks what a latent is evidence for. It separates four claims that the word "feature" runs together - that a property is linearly readable at an activation site, that the model's downstream computation depends on a direction, that the same unit recurs across training choices, and that the recurring unit is canonical, meaning unique, complete and atomic - and reads the published record against each. The record supports the first broadly, though the natural-language label attached to a latent is a weaker link than the latent itself; it supports the second unevenly, with causal roles that differ between SAE families trained on the same model; it supports the third at the grain of subspaces more readily than at the grain of individual latents; and it supports the fourth on no real language model. The reason is structural. Classical results make sparse dictionaries unique only when the data are generated by a sparse linear code meeting stated conditions; that premise has been tested only piecemeal, concept by concept, and the comparisons with a ground truth for the full dictionary that could reach canonicity have been run only on toy, synthetic, formal-language and board-game models; natural-language models have only task-specific, approximate ground truths. A short rotation argument shows how a multi-dimensional feature would produce the pattern of seed-dependent latents inside reproducible subspaces that has been reported. The paper ends with a use-indexed licensing procedure and five studies that would settle the open part. Throughout, it works from published measurements, each credited to the source that reported it.
Questions about this paper
Who wrote "Dictionaries Without Ontologies"?
Pranay Mahendrakar wrote "Dictionaries Without Ontologies: What a Sparse Autoencoder Latent Is Evidence For, Why Decodability, Use and Reproducibility Do Not Add Up to Canonicity, and the Identifiability Premise No Real-Model Result Has Checked", published 9 Oct 2026. Pranay Mahendrakar is an Indian AI specialist and LLM engineer based in Bengaluru, India. He is the Managing Director of SonyTech, Nodal Coordinator at IIRS-ISRO, and an instructor at Tutorials Point. His work covers large language models, natural language processing, computer vision and retrieval-augmented generation. He publishes open-access research papers and is the author of three books: Just AI With Pranay, Multiverse of AI and It's Me LLM.
Is "Dictionaries Without Ontologies" free to read?
Yes. "Dictionaries Without Ontologies" by Pranay Mahendrakar is open access under a Creative Commons Attribution 4.0 licence, with the full PDF available from Zenodo at https://doi.org/10.5281/zenodo.23250481. There is no paywall and no account required.
How do I cite "Dictionaries Without Ontologies"?
Cite the DOI: Mahendrakar, P. (2026). Dictionaries Without Ontologies: What a Sparse Autoencoder Latent Is Evidence For, Why Decodability, Use and Reproducibility Do Not Add Up to Canonicity, and the Identifiability Premise No Real-Model Result Has Checked. Zenodo. https://doi.org/10.5281/zenodo.23250481 A BibTeX entry is provided on this page.