Did the Model Change, or Did Your Measurement? A Drift Claim Is a Test Against a Declared Null, the Careful Longitudinal Studies Mostly Caught Announced Version Moves, and a Capability Change Under a Pinned Identifier Is the Claim the Evidence Supports Least
Abstract
Two bodies of work describe the same deployed language models in opposite terms. One reports that models behind a fixed product name change over time and calls for continuous monitoring. The other reports that a single model, queried twice under settings meant to be deterministic, returns different answers, and that benchmark scores move by several points under seeds, hardware, batch size and prompt format alone. Read together, they raise a methodological question that neither answers on its own: when is an observed difference between two dates evidence that the provider changed something, rather than a draw from the noise of the measurement? This paper argues that the common framing of that question - drift claims ignore the noise floor - is too strong in one direction and too weak in another. It is too strong because the most careful longitudinal studies did estimate a floor: the original ChatGPT drift study measured same-version disagreement on one task, a daily study of an unpinned endpoint ran repeated queries and z-tests, and a clinical study separated same-day from across-day variation. It is too weak because a noise floor is not a property of the model that can simply be measured; it is a declaration of what counts as "the same", and published monitors declare three different things. Tests that treat any change in the output distribution as change detect hardware swaps and serving-stack updates readily; tests of task capability need far larger samples and are rarely run. The confirmed changes located are mostly announced re-pointings of a name or provider infrastructure events. A capability change under a pinned identifier, with no announced event and a calibrated floor, was not established by any study found. The paper separates the three claims, collects the published magnitudes of change and of noise in one table and one figure, states the evidence standard as an algorithm, and lists the studies that would settle the rest.
Questions about this paper
Who wrote "Did the Model Change, or Did Your Measurement? A Drift Claim Is a Test Against a Declared Null, the Careful Longitudinal Studies Mostly Caught Announced Version Moves, and a Capability Change Under a Pinned Identifier Is the Claim the Evidence Supports Least"?
Pranay Mahendrakar wrote "Did the Model Change, or Did Your Measurement? A Drift Claim Is a Test Against a Declared Null, the Careful Longitudinal Studies Mostly Caught Announced Version Moves, and a Capability Change Under a Pinned Identifier Is the Claim the Evidence Supports Least", published 5 Oct 2026. Pranay Mahendrakar is an Indian AI specialist and LLM engineer based in Bengaluru, India. He is the Managing Director of SonyTech, Nodal Coordinator at IIRS-ISRO, and an instructor at Tutorials Point. His work covers large language models, natural language processing, computer vision and retrieval-augmented generation. He publishes open-access research papers and is the author of three books: Just AI With Pranay, Multiverse of AI and It's Me LLM.
Is "Did the Model Change, or Did Your Measurement? A Drift Claim Is a Test Against a Declared Null, the Careful Longitudinal Studies Mostly Caught Announced Version Moves, and a Capability Change Under a Pinned Identifier Is the Claim the Evidence Supports Least" free to read?
Yes. "Did the Model Change, or Did Your Measurement? A Drift Claim Is a Test Against a Declared Null, the Careful Longitudinal Studies Mostly Caught Announced Version Moves, and a Capability Change Under a Pinned Identifier Is the Claim the Evidence Supports Least" by Pranay Mahendrakar is open access under a Creative Commons Attribution 4.0 licence, with the full PDF available from Zenodo at https://doi.org/10.5281/zenodo.23168094. There is no paywall and no account required.
How do I cite "Did the Model Change, or Did Your Measurement? A Drift Claim Is a Test Against a Declared Null, the Careful Longitudinal Studies Mostly Caught Announced Version Moves, and a Capability Change Under a Pinned Identifier Is the Claim the Evidence Supports Least"?
Cite the DOI: Mahendrakar, P. (2026). Did the Model Change, or Did Your Measurement? A Drift Claim Is a Test Against a Declared Null, the Careful Longitudinal Studies Mostly Caught Announced Version Moves, and a Capability Change Under a Pinned Identifier Is the Claim the Evidence Supports Least. Zenodo. https://doi.org/10.5281/zenodo.23168094 A BibTeX entry is provided on this page.