Learn Before
Model-Assisted Safety and Rule-Based Reward Models - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Model-Assisted Safety Pipeline
Safety Alignment and Rule-Based Reward Models - Frontier Foundation Models, Capability Evaluation, and Just-In-Time Agent Harnesses @ University of Michigan - Ann Arbor
Rule-Based Reward Models (RBRMs)
Rule-based reward models (RBRMs) are zero-shot language model classifiers that supply additional reward signals to the policy model during reinforcement learning fine-tuning. Rather than relying solely on scalar human preference models, RBRMs provide granular, rule-governed feedback to enforce safety policies, specifically rewarding the model for correctly refusing harmful requests while preventing unnecessary refusals on safe prompts.
0
1
Contributors are:
Who are from:
Tags
Prep Sessions
Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Ch.3 Model Alignment and Safety - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Model-Assisted Safety and Rule-Based Reward Models - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Frontier Foundation Models, Capability Evaluation, and Just-In-Time Agent Harnesses @ University of Michigan - Ann Arbor
Ch.2 Post-Training Alignment and Calibration - Frontier Foundation Models, Capability Evaluation, and Just-In-Time Agent Harnesses @ University of Michigan - Ann Arbor
Safety Alignment and Rule-Based Reward Models - Frontier Foundation Models, Capability Evaluation, and Just-In-Time Agent Harnesses @ University of Michigan - Ann Arbor
Related
Model-Assisted Safety Pipeline
Rule-Based Reward Models (RBRMs)
Inputs and Rubric Classification Mechanism of RBRMs
Function and Inputs of the RLHF Reward Model
Rule-Based Reward Models for Reasoning
Which set of failure modes in standard reinforcement learning from human feedback (RLHF) is the model-assisted safety pipeline designed to address?
Rule-Based Reward Models (RBRMs)
Within this alignment framework, ___ themselves are used as tools to steer model behavior.
What specific role do rule-based reward models (RBRMs) play within the model-assisted safety pipeline?
Model-Assisted Safety Pipeline
Rule-Based Reward Models (RBRMs)
Inputs and Rubric Classification Mechanism of RBRMs
Learn After
How are rule-based reward models (RBRMs) categorized in terms of classifier type?
Rule-based reward models (RBRMs) rely solely on scalar human preference models to enforce safety policies.
During which training phase do rule-based reward models (RBRMs) supply additional reward signals to the policy model?
Explain the two key safety objectives that rule-based reward models (RBRMs) are designed to balance regarding model refusals.
Inputs and Rubric Classification Mechanism of RBRMs