1Cademy - A research team is training a language model to act as a programming assistant that writes complex, multi-step code functions. The training method rewards the model only if the final generated code executes without errors and produces the correct output. Despite extensive training, the model frequently generates code that is logically flawed, even if it sometimes produces the correct final result for the training examples. Which of the following statements best analyzes the fundamental weakness

Learn Before

Insufficiency of Outcome-Based Rewards for Complex Reasoning

Multiple Choice

A research team is training a language model to act as a programming assistant that writes complex, multi-step code functions. The training method rewards the model only if the final generated code executes without errors and produces the correct output. Despite extensive training, the model frequently generates code that is logically flawed, even if it sometimes produces the correct final result for the training examples. Which of the following statements best analyzes the fundamental weakness

Updated 2025-10-07

Contributors are:

Who are from:

Learn Before

Related