article open access

Circuits Are Estimates, Not Mechanisms: A Recovered Subgraph Is One Draw From a Class Set by the Ablation, the Dataset and the Grain, the Results That Defend Circuits Hold at the Level of That Class, and What Ablation Evidence Licenses at Each Level

Abstract

A circuit in a language model is usually presented as the mechanism by which the model performs a task: a sparse subgraph, found by ablation or patching, that reproduces the behaviour when everything else is knocked out. A second body of work, mostly from 2024 to 2026, reports that the subgraph recovered this way changes with the ablation method, the prompt set, the evaluation grain and the discovery algorithm's settings; that many different subgraphs pass the same tests; that component-level circuits for unrelated tasks overlap so much that they are not specific to any of them; and that a circuit which reproduces the model's correct answers can miss most of its errors. In this analysis the author reviews and synthesises published results to show that the two bodies of work are not in contradiction, because they measure different objects. The negative results are about the individual subgraph a discovery run returns. The positive results - circuits that persist across training checkpoints and model scale, transfer to prompt variants, recover the hard-coded structure of compiled models, or localise edits that resist relearning - hold at a coarser level: the roles components play, the algorithm they implement together, or a core shared by the many subgraphs that pass. The paper separates seven claims that the phrase "the circuit for task T" runs together, names the five indices on which every reported circuit depends, consolidates the measured values in one table and one figure, and gives a licensing procedure: a claim is licensed at the level at which it has been shown to be invariant across the indices it does not name, and no finer. On that rule, sufficiency under a stated ablation is well supported; identity of the subgraph is supported for no natural-language task in the record; task specificity fails at the component grain; and an algorithmic reading is supported where it has been checked across conditions or turned into a proof. Five studies would settle the open part.

Pranay Mahendrakar, AI specialist

About the author

Pranay Mahendrakar is an ai specialist and large language model engineer based in Bengaluru, India. He builds production artificial intelligence systems and publishes open-access research on how those systems fail. See all 84 papers by Pranay Mahendrakar, or his ORCID record.

Questions about this paper

Who wrote "Circuits Are Estimates, Not Mechanisms"?

Pranay Mahendrakar wrote "Circuits Are Estimates, Not Mechanisms: A Recovered Subgraph Is One Draw From a Class Set by the Ablation, the Dataset and the Grain, the Results That Defend Circuits Hold at the Level of That Class, and What Ablation Evidence Licenses at Each Level", published 11 Oct 2026. Pranay Mahendrakar is an Indian AI specialist and LLM engineer based in Bengaluru, India. He is the Managing Director of SonyTech, Nodal Coordinator at IIRS-ISRO, and an instructor at Tutorials Point. His work covers large language models, natural language processing, computer vision and retrieval-augmented generation. He publishes open-access research papers and is the author of three books: Just AI With Pranay, Multiverse of AI and It's Me LLM.

Is "Circuits Are Estimates, Not Mechanisms" free to read?

Yes. "Circuits Are Estimates, Not Mechanisms" by Pranay Mahendrakar is open access under a Creative Commons Attribution 4.0 licence, with the full PDF available from Zenodo at https://doi.org/10.5281/zenodo.23289446. There is no paywall and no account required.

How do I cite "Circuits Are Estimates, Not Mechanisms"?

Cite the DOI: Mahendrakar, P. (2026). Circuits Are Estimates, Not Mechanisms: A Recovered Subgraph Is One Draw From a Class Set by the Ablation, the Dataset and the Grain, the Results That Defend Circuits Hold at the Level of That Class, and What Ablation Evidence Licenses at Each Level. Zenodo. https://doi.org/10.5281/zenodo.23289446 A BibTeX entry is provided on this page.

Related research by Pranay Mahendrakar

← All papers by Pranay Mahendrakar