Concept icon
Concept

Learning from Rewards and Penalties

This learning approach trains an agent by letting it interact with an environment and observe the consequences of its actions. After each action, the agent receives feedback in the form of rewards or penalties, then adjusts its behavior to improve future outcomes. The objective is to choose actions that maximize long-term reward, not just the next immediate payoff. This is especially useful for tasks where decisions affect later results, such as game playing, robot control, or scheduling.

0

5

Concept icon
Updated 2026-08-12

Tags

Data Science

Foundations of Large Language Models Course

Computing Sciences

Machine Learning

Deep Learning

Supervised Learning

Dive into Deep Learning @ D2L

Machine Learning Yearning @ DeepLearning.AI

Learn After