article open access

How Many Classes Are Out There? In Category Discovery the Class Count Is Granularity Carried Over From the Labelled Set, Supplying It Can Supply the Taxonomy, and Without It the Count Is Set by a Hyperparameter

Abstract

Category discovery asks a model to sort an unlabelled image collection into classes, some of which it has never been shown, using a labelled subset of other classes as its guide. Almost every method needs one number before it can produce an answer: how many classes the unlabelled data contains. The literature holds two positions about that number without setting them against each other. One treats it as a property of the data that can be estimated, and reports estimators that recover it exactly on generic benchmarks. The other, voiced by the setting's own originators and by the clustering literature, holds that the number of classes is not intrinsic to the images at all but is fixed by the labelling convention. This paper argues that both positions are correct about different objects, and that the reconciliation has consequences the field's evaluation practice has not absorbed. Every category-discovery class-count estimator examined here calibrates against the labelled classes, so what it estimates is the count at the granularity the labelled set exhibits; it succeeds when the novel classes share that granularity, and should fail in a predictable direction when they do not, a prediction no published benchmark has tested. Supplying the ground-truth count, as is common in headline results, can therefore supply the taxonomy level along with it, and no standard benchmark separates the two. Where no count and no batch estimator are available, as in on-the-fly discovery, the count becomes a hyperparameter: published predicted-class counts on a 200-class benchmark range from 153 to 2,910 depending on the method and a hash length. The paper states what a protocol would have to measure to tell a granularity-transfer failure from a representation failure, and what is not known.

The literature search, drafting and citation verification for this paper were carried out with AI assistance under the author's direction. Every citation was machine-verified against its live arXiv Atom API record or its Crossref DOI record during drafting (title and author list checked against the record returned), and every quantitative claim is taken from the abstract, full text or a table of the source credited with it; the full-text numbers were read from the sources' own HTML renderings rather than from summaries. No experiment was run and no number in this paper was measured by its author. Table 1 re-presents numbers published by the cited papers, each named on its row; Figure 1 plots published class-count estimates from two cited tables with no transformation beyond a log axis. The few relative-error figures this paper computes itself from published counts are labelled as such where they appear. Algorithm 1 is original conceptual synthesis by the author, not an empirical result, and is presented as such.

Pranay Mahendrakar, AI specialist

About the author

Pranay Mahendrakar is an ai specialist and large language model engineer based in Bengaluru, India. He builds production artificial intelligence systems and publishes open-access research on how those systems fail. See all 66 papers by Pranay Mahendrakar, or his ORCID record.

Questions about this paper

Who wrote "How Many Classes Are Out There? In Category Discovery the Class Count Is Granularity Carried Over From the Labelled Set, Supplying It Can Supply the Taxonomy, and Without It the Count Is Set by a Hyperparameter"?

Pranay Mahendrakar wrote "How Many Classes Are Out There? In Category Discovery the Class Count Is Granularity Carried Over From the Labelled Set, Supplying It Can Supply the Taxonomy, and Without It the Count Is Set by a Hyperparameter", published 26 Sep 2026. Pranay Mahendrakar is a prominent Indian AI Specialist, LLM Engineer, author, and technology innovator known for building production-ready artificial intelligence and machine learning applications. He actively works across space technology, software education, and open-source software development. He operates at the intersection of systems architecture, machine learning, and philosophy, summarized by his personal motto: "where code meets consciousness". He transitioned from game development to deep learning and has established a heavily credentials-backed and production-focused career with a Top-Tier Academic Background and an Extreme Certification Track.

Is "How Many Classes Are Out There? In Category Discovery the Class Count Is Granularity Carried Over From the Labelled Set, Supplying It Can Supply the Taxonomy, and Without It the Count Is Set by a Hyperparameter" free to read?

Yes. "How Many Classes Are Out There? In Category Discovery the Class Count Is Granularity Carried Over From the Labelled Set, Supplying It Can Supply the Taxonomy, and Without It the Count Is Set by a Hyperparameter" by Pranay Mahendrakar is open access under a Creative Commons Attribution 4.0 licence, with the full PDF available from Zenodo at https://doi.org/10.5281/zenodo.22980799. There is no paywall and no account required.

How do I cite "How Many Classes Are Out There? In Category Discovery the Class Count Is Granularity Carried Over From the Labelled Set, Supplying It Can Supply the Taxonomy, and Without It the Count Is Set by a Hyperparameter"?

Cite the DOI: Mahendrakar, P. (2026). How Many Classes Are Out There? In Category Discovery the Class Count Is Granularity Carried Over From the Labelled Set, Supplying It Can Supply the Taxonomy, and Without It the Count Is Set by a Hyperparameter. Zenodo. https://doi.org/10.5281/zenodo.22980799 A BibTeX entry is provided on this page.

Related research by Pranay Mahendrakar

← All papers by Pranay Mahendrakar