Three Things Called Budget Awareness: Observability, Forecasting and Allocation in LLM Agents, Why Every Published Allocation Gain Is Keyed to a Signal Measured After the Fact, and the Run-to-Run Variance No Forecast Is Scored Against
Abstract
Two 2024-2026 literatures make claims about resource use in language-model agents that look incompatible. One reports that allocating test-time compute according to problem difficulty beats spending it uniformly, by margins up to a factor of four in compute at equal accuracy. The other reports that agents cannot estimate their own resource consumption, that the estimates are systematically optimistic, and that the ability to produce them is only weakly related to the ability to do the task. If intelligent allocation were the decisive lever, the most capable models should be the best allocators, and several independent measurements say they are not. This paper argues that the incompatibility is apparent and dissolves under a partition. The phrase budget awareness is carrying at least four distinct decisions: adherence to a budget someone else set, observability of the remaining budget inside the agent's context, forecasting of one's own future consumption, and allocation of a known budget across competing demands. The four are measured on different instruments and do not have the same answer. Adherence is largely solved in the tested ranges: Aggarwal and Welleck (2025) report output lengths that closely track requested targets at 512, 1024, 2048 and 3600 tokens. Observability is not a competence at all, and it accounts for a share of the reported failure large enough to move headline numbers. Liu et al. (2025b) report that a standard ReAct agent issues 13.77 search calls when given a budget of 30 and 14.24 when given 100, and that a plug-in which does nothing but keep the remaining budget in context reaches 12.8 percent on BrowseComp at a budget of 10 against 12.6 percent for the unmodified agent at a budget of 100, using 40.4 percent fewer search calls at 31.3 percent lower cost. Wen et al. (2025) reach the same conclusion at the token level and state the sharper form of it: naming the budget once in the prompt is insufficient, and the remaining budget has to be re-inserted periodically during generation. Forecasting is poor and structurally biased. Lin et al. (2026) find that across twenty model-environment pairs, optimistic misses outnumber conservative ones at every rollout-progress bin, that weaker models are more optimistic rather than less, and that on failed trajectories models predict feasibility above 70 percent after 60 percent of the budget is spent. Bai et al. (2026) find that eight frontier models correlate with their own realized token usage at up to 0.39 and underestimate it systematically. Allocation fails on a different axis again: Fan et al. (2026) find that under a shared budget, solving order tracks presentation position at rho = +0.68 while effort-value correlations run from 0.00 to +0.11, and that the set of questions receiving substantive work overlaps the highest-value-density set at 0.59 against a chance reference of 0.59. The central claim is about what the allocation results license. In every case located by this run's search, the per-instance signal that tells an allocator which question deserves more compute is a quantity obtained by running the model on that question and measuring the outcome: Snell et al. (2024) define the compute-optimal strategy as an argmax indexed by the ground-truth answer and bin difficulty from 2048 samples per question, stating in their own text that this "assumes oracle access to a ground-truth correctness checking function, which is of course not available upon deployment"; Damani et al. (2024) train a separate predictor on eight sampled responses per query scored by an 8B reward model, and state that Monte Carlo estimation of the quantity "may in general require more computation than we eventually wish to allocate"; Fan et al. (2026) compute value density from an independent 40,960-token reference run; Zhou et al. (2026b) gate budget on rollout-derived solvability and write that solvability "is a posterior signal, while budget investment must be committed before reasoning begins"; the two methods that advertise self-assessed difficulty, Huang et al. (2025) and Singh et al. (2025), train the self-assessment against a label computed from the pass rate over sixteen or N sampled rollouts; Nazi and Dipta (2026) score plans against an oracle that knows solvability and cost for every item; and the ground-truth target in the founding self-knowledge result, Kadavath et al. (2022), is the fraction of sampled answers that are correct. Two located designs are genuine partial exceptions and are treated as such: Zuo and Zhu (2025) pay for difficulty estimation out of the same budget they allocate, and Nogueira et al. (2025) read certainty off the forward pass and fit only a single global threshold. The two literatures are therefore not in conflict, because the out-of-band allocation gains are produced by an estimator fitted to measured outcomes at a compute cost that the same papers decline to put on the axis, and the budget-awareness results ask whether the agent can supply that quantity itself. What survives the partition is narrower and less comfortable than either headline. Feasibility prediction trains easily: Lin et al. report Qwen-7B moving from 25.5 percent to roughly 90 percent with supervised fine-tuning alone, which they read as a formatting and calibration problem rather than a capability gap. What does not train is the interval, and what does not transfer is anything, with cross-task reward retention of 17 to 36 percent. Against this, two results cut the other way and are given their weight here: Sun et al. (2026) remove a budget-state scaffold after training and find the behaviour persists, with re-adding it at inference buying nothing, and Chen et al. (2026) find that exposing resource metadata is not sufficient and that the benefit is model-dependent, concluding that resource-aware orchestration is a distinct capability rather than an automatic consequence of exposure. The paper's main open problem is a missing denominator. Bai et al. report that runs of the same agent on the same SWE-bench Verified task differ by up to 30 times in total tokens. No budget forecaster located here is scored against the spread of the quantity it forecasts, so a reported 47 percent interval coverage cannot presently be read as either near or far from what any forecaster could achieve, and the share of forecast error that belongs to the forecaster rather than to the variance of the trajectory is unmeasured. One incidental finding is recorded because the claim it affects is repeated everywhere: in Snell et al., the running text and the caption of the figure it describes state the pretraining-versus-test-time tradeoff in opposite directions on the difficulty axis, the text putting pretraining ahead on the hardest bins and the caption putting it ahead on the easy questions, and the caption's gloss of its own ratio is inverted against the definition given two paragraphs earlier.
The literature search, drafting and citation verification for this paper were carried out with AI assistance under the author's direction. Every citation was machine-verified against the DataCite record for its DOI before inclusion, and every quantitative claim was read back against the cited source's own table or text before it was written down. No experiment was run and no number in this paper was measured or recomputed by its author; every figure is quoted from the paper credited with it, and where two figures from one source are set side by side to make a point that the source does not itself make, the juxtaposition is flagged in the text. Section 2 states the search procedure and its limits, including the fact that one of the two usual metadata routes was unavailable on the day of the search, so that the coverage claims in Sections 7, 13 and 14 can be checked and, if wrong, corrected. The author is responsible for the final text and for all claims made in it.
Questions about this paper
Who wrote "Three Things Called Budget Awareness"?
Pranay Mahendrakar wrote "Three Things Called Budget Awareness: Observability, Forecasting and Allocation in LLM Agents, Why Every Published Allocation Gain Is Keyed to a Signal Measured After the Fact, and the Run-to-Run Variance No Forecast Is Scored Against", published 22 Sep 2026. Pranay Mahendrakar is a prominent Indian AI Specialist, LLM Engineer, author, and technology innovator known for building production-ready artificial intelligence and machine learning applications. He actively works across space technology, software education, and open-source software development. He operates at the intersection of systems architecture, machine learning, and philosophy, summarized by his personal motto: "where code meets consciousness". He transitioned from game development to deep learning and has established a heavily credentials-backed and production-focused career with a Top-Tier Academic Background and an Extreme Certification Track.
Is "Three Things Called Budget Awareness" free to read?
Yes. "Three Things Called Budget Awareness" by Pranay Mahendrakar is open access under a Creative Commons Attribution 4.0 licence, with the full PDF available from Zenodo at https://doi.org/10.5281/zenodo.22884949. There is no paywall and no account required.
How do I cite "Three Things Called Budget Awareness"?
Cite the DOI: Mahendrakar, P. (2026). Three Things Called Budget Awareness: Observability, Forecasting and Allocation in LLM Agents, Why Every Published Allocation Gain Is Keyed to a Signal Measured After the Fact, and the Run-to-Run Variance No Forecast Is Scored Against. Zenodo. https://doi.org/10.5281/zenodo.22884949 A BibTeX entry is provided on this page.