1Cademy - True or False: In a basic policy gradient method, if an agent completes a trajectory with a high positive total reward, the learning algorithm will reinforce every action taken during that trajectory, even those that were suboptimal or did not directly contribute to the final outcome.

Learn Before

High Variance in Policy Gradient Estimates

True/False

True or False: In a basic policy gradient method, if an agent completes a trajectory with a high positive total reward, the learning algorithm will reinforce every action taken during that trajectory, even those that were suboptimal or did not directly contribute to the final outcome.

Updated 2025-10-06

Contributors are:

Who are from:

Tags

Ch.4 Alignment - Foundations of Large Language Models

Foundations of Large Language Models

Foundations of Large Language Models Course

Computing Sciences