Introducing Predictive Scheduling. Can we predict how much reasoning a query needs before generating a single token? Blog: aneeshers.github.io/predicti… Paper: arxiv.org/abs/2602.01237 Co-led with @katrinarbrown and @rana_shahout
2
4
14
1,438
LLMs use a fixed reasoning budget at inference time. This wastes compute on easy queries and truncates hard ones. Can we allocate tokens more intelligently?
1
1
137
We call this Predictive Scheduling. The idea is to predict per-query reasoning needs before generation, and allocate a fixed token budget where it actually matters. This is the same total compute. Higher accuracy.

Feb 4, 2026 · 1:14 AM UTC

1
1
84
We train lightweight predictors on transformer hidden states to forecast reasoning difficulty, with middle layers (12–17)being most informative (peak r = 0.742).
1
1
51
We also fine-tune DeepSeek-R1-Distill-Qwen-1.5B (via LoRA) to predict query difficulty directly from the forward pass on the question text. This gives a fast, pre-generation easy/medium/hard label that we use to assign per-query token budgets.
1
1
125
Sort replies: Relevant Recent Liked