Introducing Predictive Scheduling.
Can we predict how much reasoning a query needs before generating a single token?
Blog: aneeshers.github.io/predicti…
Paper: arxiv.org/abs/2602.01237
Co-led with @katrinarbrown and @rana_shahout
2
4
14
1,438
We call this Predictive Scheduling.
The idea is to predict per-query reasoning needs before generation, and allocate a fixed token budget where it actually matters. This is the same total compute. Higher accuracy.
Feb 4, 2026 · 1:14 AM UTC
1
1
84
We train lightweight predictors on transformer hidden states to forecast reasoning difficulty, with middle layers (12–17)being most informative (peak r = 0.742).
1
1
51
We also fine-tune DeepSeek-R1-Distill-Qwen-1.5B (via LoRA) to predict query difficulty directly from the forward pass on the question text. This gives a fast, pre-generation easy/medium/hard label that we use to assign per-query token budgets.
1
1
125
On GSM8K, predictive scheduling yields +7.9% absolute accuracy at identical token cost, closing over 50% of the oracle gap. No model changes, just smarter inference-time scheduling!
Paper: arxiv.org/abs/2602.01237
Code: github.com/brownkat6/reasoni…
Blog: aneeshers.github.io/predicti…
1
93


