Learn Before
Stage I Harness Customization and Preference Learning - Frontier Foundation Models, Capability Evaluation, and Just-In-Time Agent Harnesses @ University of Michigan - Ann Arbor
Stage I: Task-Conditioned Customization and Preference Learning - Dynamic Agent Scaffolding: Synthesis, Diagnostic Repair, and Evolutionary Optimization @ University of Michigan - Ann Arbor
Stage I Task-Conditioned Harness Synthesis
Candidate Harness Preference Learning in JIT-Agent
In JIT-Agent, candidate agent harnesses are compared under identical backbone models and evaluation random seeds to ensure valid comparison. Because protocol-compliant harnesses may still exhibit divergent performance and resource consumption, preference data is constructed using a strict Pareto-efficiency condition. A candidate harness is strictly preferred over () if and only if:
where represent rollout averages for reward, execution latency, and monetary operational cost, respectively. The magnitude of preference is weighted by a multi-objective margin:
where and alpha_r, alpha_ell, alpha_kappa ge 0 balance the value dimensions, enforcing that higher task rewards are favored while strictly penalizing excess latency and monetary expense.
0
1
Tags
Prep Sessions
Frontier Foundation Models, Capability Evaluation, and Just-In-Time Agent Harnesses @ University of Michigan - Ann Arbor
Ch.3 Adaptive Agent Harness Design - Frontier Foundation Models, Capability Evaluation, and Just-In-Time Agent Harnesses @ University of Michigan - Ann Arbor
Stage I Harness Customization and Preference Learning - Frontier Foundation Models, Capability Evaluation, and Just-In-Time Agent Harnesses @ University of Michigan - Ann Arbor
Dynamic Agent Scaffolding: Synthesis, Diagnostic Repair, and Evolutionary Optimization @ University of Michigan - Ann Arbor
Ch.2 Multi-Stage Harness Optimization - Dynamic Agent Scaffolding: Synthesis, Diagnostic Repair, and Evolutionary Optimization @ University of Michigan - Ann Arbor
Stage I: Task-Conditioned Customization and Preference Learning - Dynamic Agent Scaffolding: Synthesis, Diagnostic Repair, and Evolutionary Optimization @ University of Michigan - Ann Arbor
Related
Candidate Harness Preference Learning in JIT-Agent
Stage I Task-Conditioned Harness Synthesis
Stage I Joint Customization and Preference Loss
Harness Protocol Spaces and Syntactic Subsets
Protocol-Compatible Harness Seed Bank
Candidate Harness Preference Learning in JIT-Agent
What type of model is utilized in Stage I to synthesize task-adapted scaffolds?
How many reference scaffolds are sampled to form the exemplar set for a given task, and from which partition of the seed bank are they selected?
Match each component of the teacher generation context to its corresponding definition.
Stage I task-conditioned harness synthesis operates under the fixed ___ protocol.
Order the steps involved in synthesizing and admitting a task-conditioned harness during Stage I.
Determine whether this candidate harness will be admitted into the Stage I imitation training corpus and explain the operational rule governing this decision.
Stage I Joint Customization and Preference Loss
Candidate Harness Preference Learning in JIT-Agent
Learn After
Under what condition are candidate agent harnesses evaluated and compared in JIT-Agent?
In JIT-Agent, candidate agent harness selection is trained using preference learning.
Beyond achieving high task rewards, what two operational factors does JIT-Agent seek to preserve or reduce when selecting candidate harnesses?
Detail the full mathematical criteria required for candidate harness h+ to be strictly preferred over h- (h+ ≻_τ h-) during Stage I preference data construction in JIT-Agent.
Match each mathematical notation from JIT-Agent Stage I preference learning to its operational role.
Order the steps taken in JIT-Agent to establish and quantify harness preferences for preference dataset construction.
Based on the Pareto-efficiency condition defined in JIT-Agent, determine whether Harness A is strictly preferred over Harness B (h_A ≻_τ h_B) and explain the reason.