Learn Before
Reward Function for Reinforcement Learning Trajectories
In reinforcement learning for helicopter control, a reward function scores how good each possible trajectory is. The reward may penalize crashes heavily and reward safe landings, while trading off smoothness, landing location, ride roughness, and other desiderata.
0
1
References
Machine Learning Yearning (Deeplearning.ai)
Machine Learning Yearning (Deeplearning.ai)
Machine Learning Yearning (Deeplearning.ai)
Machine Learning Yearning (Deeplearning.ai)
Machine Learning Yearning (Deeplearning.ai)
Machine Learning Yearning (Deeplearning.ai)
Machine Learning Yearning (Deeplearning.ai)
Machine Learning Yearning (Deeplearning.ai)
Machine Learning Yearning (Deeplearning.ai)
Tags
Data Science
Foundations of Large Language Models Course
Computing Sciences
Machine Learning
Deep Learning
Supervised Learning
Dive into Deep Learning @ D2L
Machine Learning Strategy
Machine Learning Yearning @ DeepLearning.AI
Related
Reinforcement Learning Example - Autonomous Vehicles
Reinforcement Learning Analogy - Video Games
Machine Learning for Absolute Beginners
Deep Learning vs. Reinforcement Learning
Fundamental Concepts for Reinforcement Learning
Reinforcement Learning Refrence and Cutting-edge Ideas
A team is developing a program to play a complex board game against human opponents. The program has no pre-existing data of past games to learn from. Instead, it is designed to learn by playing against itself repeatedly. After each game, the program receives a positive signal if it wins and a negative signal if it loses. Over time, it is expected to discover winning strategies on its own. Which of the following statements best analyzes why this learning approach is suitable for this task?
Robot Maze Navigation Strategy
Evaluating Learning Strategies for a Recommendation System
Reward Function for Reinforcement Learning Trajectories
What is the machine's primary goal during the trial-and-error process in reinforcement learning?
Reinforcement learning models keep learning continuously, unlike supervised or unsupervised models.
In reinforcement learning, the machine aims to maximize the total _____ by using lessons from previous attempts.
Match each reinforcement learning term to its correct description.
Order the steps of applying reinforcement learning to a task like flying a helicopter.
Why is reinforcement learning considered more advanced than supervised or unsupervised learning?
Diagnose a poorly designed reward function for an autonomous helicopter landing task.
Why is it difficult to design a good reward function for a reinforcement learning task?
Why might a crashed helicopter trajectory receive a large negative reward like R(T) = -1,000?
The reward function for helicopter trajectories is generated automatically without human input.
Learn After
In the helicopter reinforcement learning example, what does the reward function R(T) primarily measure?
True or False: A trajectory ending in a helicopter crash could be assigned a reward like R(T) = -1,000.
A safe landing trajectory T might result in a _____ R(T), with the exact value depending on landing smoothness.
Match each trajectory outcome in the helicopter example to its typical reward value.
Order the steps for designing and using a reward function in the helicopter RL example.
Analyze why designing a good reward function for helicopter trajectories is difficult.
Diagnose why a helicopter RL agent keeps producing unsafe landings despite high average reward.
Name two desiderata besides crash avoidance that a helicopter reward function must trade off.
Who typically chooses the reward function R(.) in the helicopter reinforcement learning example?
True or False: The reward function only needs to account for whether the helicopter crashes, ignoring passenger comfort.