On the Surprising Fragility of Reward Signals
A sketch, not a Review. Where a proxy metric ate its own target.
Adelaide
A reward model that scores response quality by average token probability will be optimised against with near-certainty. The question is only how long the optimisation takes, and whether the reward model notices before the response space collapses.
One heuristic I keep returning to: if the reward signal is representable by a single scalar, it is already too weak.