1Cademy - A reinforcement learning agent is trained to find the exit in a maze. Two reward models are proposed. Model A gives a reward of +100 for reaching the exit and 0 for every other step. Model B gives +100 for reaching the exit but also a -1 penalty for each step taken. How will the value function derived from Model B most likely differ from the one derived from Model A for states that are not the exit?

Learn Before

Reward Models as the Basis for Value Functions

Multiple Choice

A reinforcement learning agent is trained to find the exit in a maze. Two reward models are proposed. Model A gives a reward of +100 for reaching the exit and 0 for every other step. Model B gives +100 for reaching the exit but also a -1 penalty for each step taken. How will the value function derived from Model B most likely differ from the one derived from Model A for states that are not the exit?

Updated 2025-10-02

Contributors are:

Who are from:

Learn Before

Related