NeurIPS 2025 Efficient Reasoning Workshop · Spotlight
A GRPO-compatible adaptive length penalty reduced DeepScaleR-1.5B generation length by more than 50% at matched accuracy across MATH-500, AIME, and OlympiadBench, while allocating 5.35× more tokens to hard prompts than easy ones.
