COLM 2026
RL-based mid-training with dense outcome- and process-level rewards from a reference-guided LLM judge. On held-out AIME 2026 with Qwen3-4B, ExpRL improved pass@1 over the strongest baseline at each stage by 3.6 points after priming and 4.7 points after identical sparse-GRPO post-training.
