Project 2026 Summer

Forecasting Sufficient Output-Token Caps for Adaptive Inference in Black-Box Large Language Models

Project Image
Student Ryan Chung Kam Chung
Supervisor Sriram Subramanian
Abstract

Adaptive inference can improve large language model cost and quality trade-offs by assigning more computation to difficult inputs. This project defines Operational Sufficient Budget (OSB) as the smallest tested refinement cap satisfying a repeated-correctness criterion under a fixed solve policy and measurement protocol, and evaluates whether this correctness-conditioned output-cap target can be forecast from pre-generation information available in black-box hosted settings. Across 600 tasks, three hosted models, and two reasoning settings, the study compares historical summaries, prompt features, prompt-based forecasts of generated demand, and detached numerical requests. On the 2,805-row primary common cohort, controllers using prompt features achieved 87.4 to 87.9 percent measured-OSB coverage at pooled mean caps of 286 to 288 native output tokens, compared with 87.2 percent at 374 for request-only calibration. Raw requests were weakly associated with measured OSB, with a Spearman rank correlation of 0.146, and contained detectable incremental predictive association beyond prompt features, but no clear incremental allocation benefit. Results varied across hosted-model conditions and task families. Measured-OSB coverage is retrospective target attainment, not prospective correctness or realized cost. Overall, substantial structure in measured OSB was forecastable from accessible black-box signals, while detached numerical requests required condition-specific validation before serving as allocation controls.