1Cademy - An autonomous agent is navigating a maze. At a particular state, the agents value function estimates the value of its current state to be 10. The agent decides to move to an adjacent state, receiving an immediate reward of -1 for the move. The value function estimates the value of the new state to be 15. Assuming a discount factor of 0.9, calculate the one-step advantage estimate for the action taken and determine its implication for future action selection.

Learn Before

Temporal Difference (TD) Error as an Advantage Function Estimator

Multiple Choice

An autonomous agent is navigating a maze. At a particular state, the agent's value function estimates the value of its current state to be 10. The agent decides to move to an adjacent state, receiving an immediate reward of -1 for the move. The value function estimates the value of the new state to be 15. Assuming a discount factor of 0.9, calculate the one-step advantage estimate for the action taken and determine its implication for future action selection.

Updated 2025-10-02

Contributors are:

Who are from:

Learn Before

Related