Learn Before
How are rule-based reward models (RBRMs) categorized in terms of classifier type?
0
1
Contributors are:
Who are from:
Tags
Prep Sessions
Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Ch.3 Model Alignment and Safety - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Model-Assisted Safety and Rule-Based Reward Models - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Frontier Foundation Models, Capability Evaluation, and Just-In-Time Agent Harnesses @ University of Michigan - Ann Arbor
Ch.2 Post-Training Alignment and Calibration - Frontier Foundation Models, Capability Evaluation, and Just-In-Time Agent Harnesses @ University of Michigan - Ann Arbor
Safety Alignment and Rule-Based Reward Models - Frontier Foundation Models, Capability Evaluation, and Just-In-Time Agent Harnesses @ University of Michigan - Ann Arbor
Related
How are rule-based reward models (RBRMs) categorized in terms of classifier type?
Rule-based reward models (RBRMs) rely solely on scalar human preference models to enforce safety policies.
During which training phase do rule-based reward models (RBRMs) supply additional reward signals to the policy model?
Explain the two key safety objectives that rule-based reward models (RBRMs) are designed to balance regarding model refusals.
Inputs and Rubric Classification Mechanism of RBRMs