True or False: Free-response answer generation and sampling methodologies rely heavily on a model's capacity to follow open-ended instructions accurately.
0
1
Tags
Prep Sessions
Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Ch.3 Model Alignment and Safety - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Impact of RLHF on Model Capability - Transformer Architecture and Large Language Model Capabilities @ University of Michigan - Ann Arbor
Frontier Foundation Models, Capability Evaluation, and Just-In-Time Agent Harnesses @ University of Michigan - Ann Arbor
Ch.2 Post-Training Alignment and Calibration - Frontier Foundation Models, Capability Evaluation, and Just-In-Time Agent Harnesses @ University of Michigan - Ann Arbor
RLHF Effects on Capability and Calibration - Frontier Foundation Models, Capability Evaluation, and Just-In-Time Agent Harnesses @ University of Michigan - Ann Arbor
Related
When attempting to evaluate pre-trained base models against post-RLHF models on an equal footing, which format presents fewer evaluation difficulties than free-response tasks?
Explain why answer generation and sampling methodologies create an evaluation asymmetry when comparing pre-trained base models to post-RLHF models on free-response tasks.
True or False: Free-response answer generation and sampling methodologies rely heavily on a model's capacity to follow open-ended instructions accurately.
Which two specific post-training stages develop a model's ability to accurately follow open-ended instructions?