Concept icon
Concept

Checking Reward Quality in Reinforcement Learning

When an RL agent produces a trajectory that performs worse than an expert trajectory, compare the reward function's score on the expert behavior with its score on the agent's behavior. If the expert behavior receives the higher score, focus on improving the RL algorithm; if it does not, the reward function is likely the problem and should be revised.

0

1

Concept icon
Updated 2026-08-12

Contributors are:

Who are from:

Tags

Machine Learning

Deep Learning

Supervised Learning

Dive into Deep Learning @ D2L

Data Science

Machine Learning Strategy

Machine Learning Yearning @ DeepLearning.AI