When Reward Scores Favor the Wrong Outcome
If a model assigns a higher reward to a worse outcome than to a better one, the scoring rule is working as intended.
0
1
Tags
Machine Learning
Deep Learning
Supervised Learning
Dive into Deep Learning @ D2L
Data Science
Machine Learning Strategy
Machine Learning Yearning @ DeepLearning.AI
Related
Diagnosing an RL problem with verified rewards
When Reward Scores Favor the Wrong Outcome
Interpreting _____ in reward verification
Match the trajectories and reward comparison to their meanings.
Reward Comparison for a Reinforcement Learning Check
What It Means When the Human Trajectory Outperforms the Learned One
Autonomous Drone Landing Reward Diagnosis
Deciding the RL component to fix
What does it mean if a human plan scores lower than an automated plan on the reward model?
What Optimization Verification Checks