article open access

Mechanistic Interpretability of In-Context Learning: A Survey of Known Circuits and Open Problems

Abstract

In-context learning (ICL) — the ability of a transformer language model to acquire a new input-output mapping from a handful of demonstrations in its prompt, with no weight updates — remains the most striking and least understood capability of large language models. Mechanistic interpretability has produced two influential partial explanations: induction-head circuits, which implement fuzzy pattern completion of the form [A][B]…[A]→[B], and function vectors, which compress an entire ICL task into a low-dimensional representation routed by a small number of attention heads. A third line of work argues that ICL implements implicit gradient descent on linear regression problems. This paper contributes (i) a structured synthesis of these three accounts, (ii) a sharper articulation of where each explanation succeeds and where it breaks down, and (iii) a concrete research agenda targeting five gaps that current work does not address: compositional ICL, ICL in instruction-tuned frontier models, the transient-versus-persistent learning distinction, the polysemanticity barrier at scale, and the deeper interpretive question of what it means to claim that a circuit "explains" a behaviour. We argue the field has solved the easy cases and now faces the hard ones.

Pranay Mahendrakar, AI specialist

About the author

Pranay Mahendrakar is an ai specialist and large language model engineer based in Bengaluru, India. He builds production artificial intelligence systems and publishes open-access research on how those systems fail. See all 70 papers by Pranay Mahendrakar, or his ORCID record.

Questions about this paper

Who wrote "Mechanistic Interpretability of In-Context Learning"?

Pranay Mahendrakar wrote "Mechanistic Interpretability of In-Context Learning: A Survey of Known Circuits and Open Problems", published 28 Apr 2026. Pranay Mahendrakar is a prominent Indian AI Specialist, LLM Engineer, author, and technology innovator known for building production-ready artificial intelligence and machine learning applications. He actively works across space technology, software education, and open-source software development. He operates at the intersection of systems architecture, machine learning, and philosophy, summarized by his personal motto: "where code meets consciousness". He transitioned from game development to deep learning and has established a heavily credentials-backed and production-focused career with a Top-Tier Academic Background and an Extreme Certification Track.

Is "Mechanistic Interpretability of In-Context Learning" free to read?

Yes. "Mechanistic Interpretability of In-Context Learning" by Pranay Mahendrakar is open access under a Creative Commons Attribution 4.0 licence, with the full PDF available from Zenodo at https://doi.org/10.5281/zenodo.19853292. There is no paywall and no account required.

How do I cite "Mechanistic Interpretability of In-Context Learning"?

Cite the DOI: Mahendrakar, P. (2026). Mechanistic Interpretability of In-Context Learning: A Survey of Known Circuits and Open Problems. Zenodo. https://doi.org/10.5281/zenodo.19853292 A BibTeX entry is provided on this page.

Related research by Pranay Mahendrakar

← All papers by Pranay Mahendrakar