1Cademy - An engineer is implementing a reward model by adapting a pre-trained language model. After feeding a concatenated prompt and response sequence into the model, they have access to the final layers hidden state vector for each token in the sequence. To derive a single scalar reward score from these vectors, which of the following procedures should they implement?

Learn Before

Reward Model Implementation using a Pre-trained LLM

Multiple Choice

An engineer is implementing a reward model by adapting a pre-trained language model. After feeding a concatenated prompt and response sequence into the model, they have access to the final layer's hidden state vector for each token in the sequence. To derive a single scalar reward score from these vectors, which of the following procedures should they implement?

Updated 2025-09-29

Contributors are:

Who are from:

Learn Before

Related