Definition icon
Definition

Rule-Based Reward Models (RBRMs)

Rule-based reward models (RBRMs) are zero-shot language model classifiers that supply additional reward signals to the policy model during reinforcement learning fine-tuning. Rather than relying solely on scalar human preference models, RBRMs provide granular, rule-governed feedback to enforce safety policies, specifically rewarding the model for correctly refusing harmful requests while preventing unnecessary refusals on safe prompts.

0

1

Definition icon
Updated 2026-09-07

Tags

Prep Sessions

Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor

Ch.3 Model Alignment and Safety - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor

Model-Assisted Safety and Rule-Based Reward Models - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor