article open access

Problem Choice Without a Referee: An Automated Novelty Check Certifies an Empty Search, a Rediscovery Benchmark Credits the Hypothesis Humans Found Next, and Every Located Referee of Worth Independent of Proposer and Field Needs the Objective Given

Abstract

Pipelines that claim to automate scientific discovery gate their search on a judgement that a proposed idea or problem is novel and worth pursuing. The best-controlled evidence on that judgement points two ways at once: in a blind study involving more than one hundred NLP researchers, expert reviewers rated language-model research ideas as more novel than expert-written ones, and when forty-three researchers executed randomly assigned ideas from the same study, the model ideas lost that advantage and fell further than expert ideas on every metric, novelty included. This paper asks what, if anything, can referee a proposed research problem before the field has worked on it. It distinguishes four kinds of referee: an absence check (was a match retrieved?), a judgement before execution (does a reader rate it novel or promising?), uptake (did humans later publish it?), and outcome (did executing it work against an objective fixed in advance?). Assembling published measurements for each, it argues that the absence check fails mainly at retrieval, with comparison a smaller and less-measured source of error, and is benchmarkable without labels; that model and expert judges disagree about which questions are non-obvious, and human judges have documented biases of their own, including lower merit scores for highly novel proposals; that uptake benchmarks supply the research background and score the rediscovered hypothesis, so by construction they credit only what humans went on to publish, as several of their authors concede; and that, of the referees located, the only one that scores worth and is independent of both the proposer and the field's present and future judgement is the outcome referee, which requires the objective to be supplied. Choosing the objective is what problem discovery means, so at that step the only referees of worth on offer are the field's own judgement, now or later. That conclusion is partly definitional, and the observation that current systems take the research question as given has been made before; the contribution is the decomposition, the measurements that locate each referee's failure, the finding that the absence check can be audited without labels, and a refutation condition that a domain-general criterion of problem worth could meet. The claim is scoped to research-idea and research-question generation; it does not say that such systems cannot discover anything.

The literature search, drafting and citation verification for this paper were carried out with AI assistance under the author's direction. Every arXiv citation was machine-verified against its live arXiv Atom API record, and every other citation against its Crossref record (with PubMed, OpenAlex or Semantic Scholar used to read abstracts Crossref does not carry), during drafting; title and author list were checked against the record returned. The full texts (arXiv PDF renderings) of the load-bearing sources were read for the passages and numbers attributed to them, including Si, Yang and Hashimoto (2024), Si, Hashimoto and Yang (2025), Gupta and Pruthi (2025), Beel et al. (2025), Lu et al. (2024), Yamada et al. (2025), Sinhahajari et al. (2026), Liu and Zhai (2026), Wen et al. (2025), Sourati and Evans (2023), Krenn et al. (2023), Luo et al. (2025), Yang et al. (MOOSE-Chem), Kumar et al. (2025), Wang et al. (RND) and Bianchi et al. (2025). Every quantitative claim is taken from the abstract, main text or a table of the source credited with it; where published numbers are summed, the text says so. No experiment was run and no number in this paper was measured by its author. Table 1 and Figure 1 re-present published values, each named with its source. The four-referee decomposition, Table 2, Algorithm 1 and the proposed tests in Section 12 are conceptual synthesis by the author, not empirical results, and are presented as such.

Pranay Mahendrakar, AI specialist

About the author

Pranay Mahendrakar is an ai specialist and large language model engineer based in Bengaluru, India. He builds production artificial intelligence systems and publishes open-access research on how those systems fail. See all 66 papers by Pranay Mahendrakar, or his ORCID record.

Questions about this paper

Who wrote "Problem Choice Without a Referee"?

Pranay Mahendrakar wrote "Problem Choice Without a Referee: An Automated Novelty Check Certifies an Empty Search, a Rediscovery Benchmark Credits the Hypothesis Humans Found Next, and Every Located Referee of Worth Independent of Proposer and Field Needs the Objective Given", published 2 Oct 2026. Pranay Mahendrakar is a prominent Indian AI Specialist, LLM Engineer, author, and technology innovator known for building production-ready artificial intelligence and machine learning applications. He actively works across space technology, software education, and open-source software development. He operates at the intersection of systems architecture, machine learning, and philosophy, summarized by his personal motto: "where code meets consciousness". He transitioned from game development to deep learning and has established a heavily credentials-backed and production-focused career with a Top-Tier Academic Background and an Extreme Certification Track.

Is "Problem Choice Without a Referee" free to read?

Yes. "Problem Choice Without a Referee" by Pranay Mahendrakar is open access under a Creative Commons Attribution 4.0 licence, with the full PDF available from Zenodo at https://doi.org/10.5281/zenodo.23093043. There is no paywall and no account required.

How do I cite "Problem Choice Without a Referee"?

Cite the DOI: Mahendrakar, P. (2026). Problem Choice Without a Referee: An Automated Novelty Check Certifies an Empty Search, a Rediscovery Benchmark Credits the Hypothesis Humans Found Next, and Every Located Referee of Worth Independent of Proposer and Field Needs the Objective Given. Zenodo. https://doi.org/10.5281/zenodo.23093043 A BibTeX entry is provided on this page.

Related research by Pranay Mahendrakar

← All papers by Pranay Mahendrakar