Theory

Goodhart's Law in Reward Modeling

Goodhart's Law provides a theoretical explanation for the overoptimization problem. The law states that when a measure, such as a reward score, is elevated to become an optimization target, it ceases to be a reliable indicator of the quality it was intended to represent.

0

1

Updated 2026-05-03

Contributors are:

Who are from:

Tags

Ch.4 Alignment - Foundations of Large Language Models

Foundations of Large Language Models

Foundations of Large Language Models Course

Computing Sciences