1Cademy - An AI development team is fine-tuning a language model using a reinforcement learning process guided by a reward model. They observe that the models outputs, while receiving high scores from the reward model, are becoming stylistically unnatural and deviating significantly from the helpful tone established during its initial supervised training. Which of the following adjustments to the training process is most specifically designed to counteract this behavioral drift?

Learn Before

Reference Policy in RLHF

Multiple Choice

An AI development team is fine-tuning a language model using a reinforcement learning process guided by a reward model. They observe that the model's outputs, while receiving high scores from the reward model, are becoming stylistically unnatural and deviating significantly from the helpful tone established during its initial supervised training. Which of the following adjustments to the training process is most specifically designed to counteract this behavioral drift?

Updated 2025-09-26

Contributors are:

Who are from:

Learn Before

Related