ICML 2026
Identified persistent length, uncertainty, position, sycophancy, and model-style biases in language reward models. Null-space probe projections reduced three of these biases without retraining or loss on RewardBench-2.
ICML 2026
Identified persistent length, uncertainty, position, sycophancy, and model-style biases in language reward models. Null-space probe projections reduced three of these biases without retraining or loss on RewardBench-2.