One Bias After Another: Mechanistic Reward Shaping and Persistent Biases in Language Reward Models

D. Fein*, M. Lamparth*, Violet Xiang, M. J. Kochenderfer, N. Haber.

ICML 2026

Identified persistent length, uncertainty, position, sycophancy, and model-style biases in language reward models. Null-space probe projections reduced three of these biases without retraining or loss on RewardBench-2.

← All publications