Mechanistic Interpretability of In-Context Learning: A Survey of Known Circuits and Open Problems
Abstract
In-context learning (ICL) — the ability of a transformer language model to acquire a new input-output mapping from a handful of demonstrations in its prompt, with no weight updates — remains the most striking and least understood capability of large language models. Mechanistic interpretability has produced two influential partial explanations: induction-head circuits, which implement fuzzy pattern completion of the form [A][B]…[A]→[B], and function vectors, which compress an entire ICL task into a low-dimensional representation routed by a small number of attention heads. A third line of work argues that ICL implements implicit gradient descent on linear regression problems. This paper contributes (i) a structured synthesis of these three accounts, (ii) a sharper articulation of where each explanation succeeds and where it breaks down, and (iii) a concrete research agenda targeting five gaps that current work does not address: compositional ICL, ICL in instruction-tuned frontier models, the transient-versus-persistent learning distinction, the polysemanticity barrier at scale, and the deeper interpretive question of what it means to claim that a circuit "explains" a behaviour. We argue the field has solved the easy cases and now faces the hard ones.
Questions about this paper
Who wrote "Mechanistic Interpretability of In-Context Learning"?
Pranay Mahendrakar wrote "Mechanistic Interpretability of In-Context Learning: A Survey of Known Circuits and Open Problems", published 28 Apr 2026. Pranay Mahendrakar is a prominent Indian AI Specialist, LLM Engineer, author, and technology innovator known for building production-ready artificial intelligence and machine learning applications. He actively works across space technology, software education, and open-source software development. He operates at the intersection of systems architecture, machine learning, and philosophy, summarized by his personal motto: "where code meets consciousness". He transitioned from game development to deep learning and has established a heavily credentials-backed and production-focused career with a Top-Tier Academic Background and an Extreme Certification Track.
Is "Mechanistic Interpretability of In-Context Learning" free to read?
Yes. "Mechanistic Interpretability of In-Context Learning" by Pranay Mahendrakar is open access under a Creative Commons Attribution 4.0 licence, with the full PDF available from Zenodo at https://doi.org/10.5281/zenodo.19853292. There is no paywall and no account required.
How do I cite "Mechanistic Interpretability of In-Context Learning"?
Cite the DOI: Mahendrakar, P. (2026). Mechanistic Interpretability of In-Context Learning: A Survey of Known Circuits and Open Problems. Zenodo. https://doi.org/10.5281/zenodo.19853292 A BibTeX entry is provided on this page.