Stage I Joint Customization and Preference Loss
The training objective for Stage I in JIT-Agent couples supervised imitation learning with reference-anchored preference optimization into a combined loss with .
The supervised fine-tuning loss maximizes the log-likelihood of validated teacher tokens to teach protocol-valid scaffold generation:
The preference loss biases candidate generation toward scaffolds that are simultaneously effective and resource-efficient by anchoring against the frozen Stage-I SFT model , weighted by the value margin :
where is the logistic sigmoid, controls preference sharpness, and is length-normalized sequence log-likelihood.
0
1
Tags
Prep Sessions
Dynamic Agent Scaffolding: Synthesis, Diagnostic Repair, and Evolutionary Optimization @ University of Michigan - Ann Arbor
Ch.2 Multi-Stage Harness Optimization - Dynamic Agent Scaffolding: Synthesis, Diagnostic Repair, and Evolutionary Optimization @ University of Michigan - Ann Arbor
Stage I: Task-Conditioned Customization and Preference Learning - Dynamic Agent Scaffolding: Synthesis, Diagnostic Repair, and Evolutionary Optimization @ University of Michigan - Ann Arbor
Related
Stage I Task-Conditioned Harness Synthesis
Stage I Joint Customization and Preference Loss
Harness Protocol Spaces and Syntactic Subsets
Protocol-Compatible Harness Seed Bank
Candidate Harness Preference Learning in JIT-Agent
What type of model is utilized in Stage I to synthesize task-adapted scaffolds?
How many reference scaffolds are sampled to form the exemplar set for a given task, and from which partition of the seed bank are they selected?
Match each component of the teacher generation context to its corresponding definition.
Stage I task-conditioned harness synthesis operates under the fixed ___ protocol.
Order the steps involved in synthesizing and admitting a task-conditioned harness during Stage I.
Determine whether this candidate harness will be admitted into the Stage I imitation training corpus and explain the operational rule governing this decision.
Stage I Joint Customization and Preference Loss
Candidate Harness Preference Learning in JIT-Agent
Learn After
In the Stage I preference loss formulation , what is the function of the hyperparameter ?
In the Stage I preference loss , the reference model is actively updated alongside the policy parameters .
What specific generation capability does the supervised fine-tuning loss instill in the model?
Describe the two primary criteria that candidate scaffolds are biased toward by the preference loss , and explain the function of the weighting factor .
The training objective for Stage I in JIT-Agent couples reference-anchored preference optimization with supervised ___ learning into a combined loss.
Arrange the mathematical operations in the order they are evaluated within the preference loss expectation to compute the weighted argument of the loss.
Explain the structural impact of the configuration in Experiment A on the overall loss function, and describe the flaw introduced in Experiment B by fixing .