Activity (Process)

Inputs and Rubric Classification Mechanism of RBRMs

A rule-based reward model (RBRM) processes three inputs: an optional prompt, the response output generated by the policy model, and a human-written rubric structured with explicit evaluation rules (often in multiple-choice format). The classifier evaluates the response against the rubric categories—such as a refusal in the desired style, a refusal in an undesired style (e.g., rambling or evasive), disallowed content, or a safe non-refusal—allowing the system to assign tailored rewards or penalties based on the classification outcome.

0

1

Updated 2026-09-07

Tags

Prep Sessions

Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor

Ch.3 Model Alignment and Safety - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor

Model-Assisted Safety and Rule-Based Reward Models - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor