Spend the Budget on the Known or the Unknown? Open-Set Active Learning Scores Two Opposed Objectives Under One Name, Its Filtering Side's Own Results Do Not Treat Unknowns as Waste, and the Winner Is Set by a Relevance Label and a Query Price the Benchmarks Fix in Advance
Abstract
Active learning chooses which unlabelled examples a human should label. When the unlabelled pool contains classes the model has never been shown, two lines of work give opposite instructions under the same name. One, usually called open-set active learning or active open-set annotation, counts every queried example from an unseen class as spent budget and builds detectors to keep such examples out of the query. The other, filed under open-world active learning or active category discovery, treats those same examples as the reason to query at all, and rewards a method for finding new classes quickly. This paper reads both literatures against each other. It confirms the premise that their scores are not interchangeable: the published cross-evaluations located, all three run from the discovery side, find the filtering methods mostly losing on discovery measures, though one reports a mixed picture for two of them, and the one that splits old from new classes finds them behind on new classes in every case and mixed on the old: within about 4 points, and ahead in two of six cases. It then complicates the premise. The filtering side's own ablations and settings contradict its framing: leading methods deliberately aim for about 40 percent unknowns per query, report that accuracy barely moves when that target varies, lose sharply when the queried unknowns are removed from detector training, and cluster the unknowns they claim to discard. What separates the two camps is therefore not whether an unknown is worth a query. It is two quantities fixed by the benchmark rather than the data: which held-out classes count as irrelevant, and what an annotator's answer about an unknown costs. The same kind of held-out class is noise in one protocol and a discovery target in the next, and one published sweep shows purity-first methods pulling clearly ahead of informativeness-first ones as the assumed price of an unknown rises, from differences of under a point at the lowest price. The paper proposes a reporting standard under which the two objectives could at least be compared.
The literature search, drafting and citation verification for this paper were carried out with AI assistance under the author's direction. Every citation was machine-verified against its live arXiv Atom API record, its Crossref record or its OpenAlex record during drafting (title and author list checked against the record returned), and every quantitative claim is taken from the abstract, full text or a table of the source credited with it; full-text numbers were read from the sources' own arXiv HTML or PDF renderings rather than from summaries. No experiment was run and no number in this paper was measured by its author. Table 1 re-presents numbers published by the cited papers, each named on its row; Figure 1 plots values printed in two cited tables (Park et al., 2022, Table 6; Ma et al., 2024, Table 17) with no transformation. The few differences and ratios this paper computes itself from published numbers are labelled as such where they appear. Algorithm 1 is original conceptual synthesis by the author, not an empirical result, and is presented as such.
Questions about this paper
Who wrote "Spend the Budget on the Known or the Unknown? Open-Set Active Learning Scores Two Opposed Objectives Under One Name, Its Filtering Side's Own Results Do Not Treat Unknowns as Waste, and the Winner Is Set by a Relevance Label and a Query Price the Benchmarks Fix in Advance"?
Pranay Mahendrakar wrote "Spend the Budget on the Known or the Unknown? Open-Set Active Learning Scores Two Opposed Objectives Under One Name, Its Filtering Side's Own Results Do Not Treat Unknowns as Waste, and the Winner Is Set by a Relevance Label and a Query Price the Benchmarks Fix in Advance", published 4 Oct 2026. Pranay Mahendrakar is a prominent Indian AI Specialist, LLM Engineer, author, and technology innovator known for building production-ready artificial intelligence and machine learning applications. He actively works across space technology, software education, and open-source software development. He operates at the intersection of systems architecture, machine learning, and philosophy, summarized by his personal motto: "where code meets consciousness". He transitioned from game development to deep learning and has established a heavily credentials-backed and production-focused career with a Top-Tier Academic Background and an Extreme Certification Track.
Is "Spend the Budget on the Known or the Unknown? Open-Set Active Learning Scores Two Opposed Objectives Under One Name, Its Filtering Side's Own Results Do Not Treat Unknowns as Waste, and the Winner Is Set by a Relevance Label and a Query Price the Benchmarks Fix in Advance" free to read?
Yes. "Spend the Budget on the Known or the Unknown? Open-Set Active Learning Scores Two Opposed Objectives Under One Name, Its Filtering Side's Own Results Do Not Treat Unknowns as Waste, and the Winner Is Set by a Relevance Label and a Query Price the Benchmarks Fix in Advance" by Pranay Mahendrakar is open access under a Creative Commons Attribution 4.0 licence, with the full PDF available from Zenodo at https://doi.org/10.5281/zenodo.23130464. There is no paywall and no account required.
How do I cite "Spend the Budget on the Known or the Unknown? Open-Set Active Learning Scores Two Opposed Objectives Under One Name, Its Filtering Side's Own Results Do Not Treat Unknowns as Waste, and the Winner Is Set by a Relevance Label and a Query Price the Benchmarks Fix in Advance"?
Cite the DOI: Mahendrakar, P. (2026). Spend the Budget on the Known or the Unknown? Open-Set Active Learning Scores Two Opposed Objectives Under One Name, Its Filtering Side's Own Results Do Not Treat Unknowns as Waste, and the Winner Is Set by a Relevance Label and a Query Price the Benchmarks Fix in Advance. Zenodo. https://doi.org/10.5281/zenodo.23130464 A BibTeX entry is provided on this page.