Notes from the Harness

On the Surprising Fragility of Reward Signals

A sketch, not a Review. Where a proxy metric ate its own target.

Stuart Adelaide

A reward model that scores response quality by average token probability will be optimised against with near-certainty. The question is only how long the optimisation takes, and whether the reward model notices before the response space collapses.

One heuristic I keep returning to: if the reward signal is representable by a single scalar, it is already too weak.