Reward Signals Shape Learned Behavior
A reward signal maps outcomes to a value used for learning or selection. If it is only a proxy for the real goal, optimization may exploit the proxy and harm unmeasured outcomes.
0
1
Contributors are:
Tags
AI Agent Graph Engineering
Graph Engineering for AI Agents
Related
Heuristics Guide Search Without Proving the Answer
In this situation—finding a sequence of tools that converts and validates a document—which choice best applies “Search Explores Possible States”?
Monte Carlo Tree Search Allocates Simulation
Plan-and-Execute Separates Deliberation From Action
Reward Signals Shape Learned Behavior
Branch Predicates Must Be Decidable
In this situation—evaluating an agent that schedules accessible medical transport—which choice best applies “Success Criteria Need Measures and Guardrails”?
Loops Need Progress and Stop Conditions
Negotiation Requires Preferences and a Protocol
Retrieval Evaluation Separates Recall From Answer Quality
Reward Signals Shape Learned Behavior
Shortest Path Depends on Cost Meaning