article open access

Three Rounds on Emergent Analogy: What the Webb, Hodel-West and Lewis-Mitchell Exchange Settled, How Its Third Round Tested a System While Still Claiming the Model, and Whether 'Counting' Names the Step That Analogy Theory Calls Inference

Abstract

In 2023 a large language model was reported to solve text-based analogy problems zero-shot at or above the level of college students. Two critiques followed. They showed that performance on letter-string analogies collapses when the alphabet is permuted or replaced by symbols, while human performance does not. The original authors replied that the failures come from an auxiliary difficulty with counting, and that GPT-4 solves the permuted problems at a human level once it can write and execute code. This paper reads all three rounds in full, including the preprint and published versions of the reply, and sorts their claims into three groups. Settled: the replications agree, and the unaided model fails the counterfactual variants under answer-only prompts. Moved: the evidence. The third round's decisive result is for a model-plus-interpreter system, while its title claim remains about language models. Unsettled: which operations count as auxiliary, and whether the models represent the new alphabet at all. Published checks from both sides, run in different studies with different alphabets, suggest splitting "counting" into two operations: GPT-4 names the one-step successor of a letter in a permuted alphabet almost perfectly, but identifies the interval (of up to two steps) between two given letters about one time in ten. A study of four newer models, however, finds every model worse at naming items two steps away, and its authors conclude that the models do not build representations of novel alphabets on the fly. If the difficulty lies in recognising the relation in the source pair, it lies in what componential theories of analogy call inference; if it lies in representing the order, it lies in encoding, which the same theories also count as part of analogy. On that decomposition, either way, the reply's "auxiliary" label is not yet earned. The paper states a procedure for reading capacity claims from counterfactual tasks, in which the decomposition of the task is fixed before data are seen, and proposes five experiments on existing materials that could decide between the readings. Every number in it is taken, or summed, from the published sources.

The literature search, drafting and citation verification for this paper were carried out with AI assistance under the author's direction. Every arXiv citation was machine-verified against its live arXiv Atom API record, and every other citation against its Crossref or DOI record, during drafting (title and author list checked against the record returned). The three rounds of the exchange were read in full: the arXiv HTML renderings of Webb, Holyoak and Lu (2023), Hodel and West (2023), Lewis and Mitchell (2024a, 2024b) and Webb, Holyoak and Lu (2024), and the Europe PMC full text of the published PNAS Nexus version (Webb, Holyoak and Lu, 2025). Every quantitative claim is taken from the abstract, main text or a table of the source credited with it. Two numbers (GPT-4 with code execution solving 40 of 60 and 30 of 60 problems) were computed by the author by summing the error counts in Tables 1 and 2 of Webb et al. (2024); the text says so where they appear. No experiment was run and no number in this paper was measured by its author. Table 1 re-presents published numbers, each named on its row; Figure 1 re-plots published values with no transformation. The three-unit account, the production/recognition split, Algorithm 1 and the proposed tests in Section 12 are conceptual synthesis by the author, not empirical results, and are presented as such. The supplementary material of the PNAS Nexus version was not read.

Pranay Mahendrakar, AI specialist

About the author

Pranay Mahendrakar is an ai specialist and large language model engineer based in Bengaluru, India. He builds production artificial intelligence systems and publishes open-access research on how those systems fail. See all 66 papers by Pranay Mahendrakar, or his ORCID record.

Questions about this paper

Who wrote "Three Rounds on Emergent Analogy"?

Pranay Mahendrakar wrote "Three Rounds on Emergent Analogy: What the Webb, Hodel-West and Lewis-Mitchell Exchange Settled, How Its Third Round Tested a System While Still Claiming the Model, and Whether 'Counting' Names the Step That Analogy Theory Calls Inference", published 30 Sep 2026. Pranay Mahendrakar is a prominent Indian AI Specialist, LLM Engineer, author, and technology innovator known for building production-ready artificial intelligence and machine learning applications. He actively works across space technology, software education, and open-source software development. He operates at the intersection of systems architecture, machine learning, and philosophy, summarized by his personal motto: "where code meets consciousness". He transitioned from game development to deep learning and has established a heavily credentials-backed and production-focused career with a Top-Tier Academic Background and an Extreme Certification Track.

Is "Three Rounds on Emergent Analogy" free to read?

Yes. "Three Rounds on Emergent Analogy" by Pranay Mahendrakar is open access under a Creative Commons Attribution 4.0 licence, with the full PDF available from Zenodo at https://doi.org/10.5281/zenodo.23049431. There is no paywall and no account required.

How do I cite "Three Rounds on Emergent Analogy"?

Cite the DOI: Mahendrakar, P. (2026). Three Rounds on Emergent Analogy: What the Webb, Hodel-West and Lewis-Mitchell Exchange Settled, How Its Third Round Tested a System While Still Claiming the Model, and Whether 'Counting' Names the Step That Analogy Theory Calls Inference. Zenodo. https://doi.org/10.5281/zenodo.23049431 A BibTeX entry is provided on this page.

Related research by Pranay Mahendrakar

← All papers by Pranay Mahendrakar