5 / papers

Interpretability

Opening the model up — circuits, features, and what "explaining a behaviour" actually means. 5 open-access papers by Pranay Mahendrakar, each with a permanent DOI.

About this research line

What has Pranay Mahendrakar published on Interpretability?

Pranay Mahendrakar has published 5 open-access papers on Interpretability: "Predicting Your Own Failure Is Not Predicting Another Model's Advantage", "The Self-Verification Gap", "Memory Architectures Beyond Attention", "Mechanistic Interpretability of In-Context Learning" and "Sheaf-Theoretic Semantics and Quantum Contextuality in Large Language Models". Each is deposited on Zenodo with a permanent DOI and is free to read.

← All research topics