Concept icon
Concept

Candidate Harness Preference Learning in JIT-Agent

In JIT-Agent, candidate agent harnesses are compared under identical backbone models πψ\pi_\psi and evaluation random seeds to ensure valid comparison. Because protocol-compliant harnesses may still exhibit divergent performance and resource consumption, preference data DIpref\mathcal{D}_{\text{I}}^{\text{pref}} is constructed using a strict Pareto-efficiency condition. A candidate harness h+h^+ is strictly preferred over h−h^- (h+≻τh−h^+ \succ_\tau h^-) if and only if:

h+≻τh−  ⟺  r+>r−∧ℓ+≤ℓ−∧κ+≤κ−∧(ℓ+<ℓ−∨κ+<κ−)h^+ \succ_\tau h^- \iff r^+ > r^- \land \ell^+ \le \ell^- \land \kappa^+ \le \kappa^- \land (\ell^+ < \ell^- \lor \kappa^+ < \kappa^-)

where (r±,ℓ±,κ±)(r^\pm, \ell^\pm, \kappa^\pm) represent rollout averages for reward, execution latency, and monetary operational cost, respectively. The magnitude of preference is weighted by a multi-objective margin:

Δval(τ;h+,h−)=αr(r+−r−)+αℓ[ℓ−−ℓ+]++ακ[κ−−κ+]+\Delta_{\text{val}}(\tau; h^+, h^-) = \alpha_r (r^+ - r^-) + \alpha_\ell [\ell^- - \ell^+]_+ + \alpha_\kappa [\kappa^- - \kappa^+]_+

where [x]+=max⁡(x,0)[x]_+ = \max(x, 0) and alpha_r, alpha_ell, alpha_kappa ge 0 balance the value dimensions, enforcing that higher task rewards are favored while strictly penalizing excess latency and monetary expense.

0

1

Concept icon
Updated 2026-10-02

Tags

Prep Sessions

Frontier Foundation Models, Capability Evaluation, and Just-In-Time Agent Harnesses @ University of Michigan - Ann Arbor

Ch.3 Adaptive Agent Harness Design - Frontier Foundation Models, Capability Evaluation, and Just-In-Time Agent Harnesses @ University of Michigan - Ann Arbor

Stage I Harness Customization and Preference Learning - Frontier Foundation Models, Capability Evaluation, and Just-In-Time Agent Harnesses @ University of Michigan - Ann Arbor

Dynamic Agent Scaffolding: Synthesis, Diagnostic Repair, and Evolutionary Optimization @ University of Michigan - Ann Arbor

Ch.2 Multi-Stage Harness Optimization - Dynamic Agent Scaffolding: Synthesis, Diagnostic Repair, and Evolutionary Optimization @ University of Michigan - Ann Arbor

Stage I: Task-Conditioned Customization and Preference Learning - Dynamic Agent Scaffolding: Synthesis, Diagnostic Repair, and Evolutionary Optimization @ University of Michigan - Ann Arbor