Activity (Process)

Stage I Task-Conditioned Harness Synthesis

Stage I of JIT-Agent training constructs an imitation training corpus DI\mathcal{D}_{\text{I}} by leveraging a frozen, stronger teacher model qϕq_\phi to synthesize task-adapted scaffolds under the fixed four-module protocol Π\Pi. For each task τ\tau, three reference scaffolds are sampled from the task-type-matched partition of the seed bank: Eτ={h(1),h(2),h(3)}∼Sample3(B0(d(τ)))\mathcal{E}_\tau = \{h^{(1)}, h^{(2)}, h^{(3)}\} \sim \text{Sample}_3(\mathcal{B}_0^{(d(\tau))}), where d(τ)d(\tau) specifies the task type. The teacher receives the generation context c_tau = (tau, Pi, C_tau, mathcal{E}_tau), consisting of the task specification, protocol schemas, available capability registry CτC_\tau, and sampled reference scaffolds. A generated harness hteach∼qϕ(⋅∣cτ)h^{\text{teach}} \sim q_\phi(\cdot \mid c_\tau) is admitted into DI\mathcal{D}_{\text{I}} if and only if it satisfies static protocol rules and executes successfully under environment validation checks, formalized as ValidΠ(hteach;τ,πψ,Cτ)=1\text{Valid}_\Pi(h^{\text{teach}}; \tau, \pi_\psi, C_\tau) = 1.

0

1

Updated 2026-10-02

Tags

Prep Sessions

Dynamic Agent Scaffolding: Synthesis, Diagnostic Repair, and Evolutionary Optimization @ University of Michigan - Ann Arbor

Ch.2 Multi-Stage Harness Optimization - Dynamic Agent Scaffolding: Synthesis, Diagnostic Repair, and Evolutionary Optimization @ University of Michigan - Ann Arbor

Stage I: Task-Conditioned Customization and Preference Learning - Dynamic Agent Scaffolding: Synthesis, Diagnostic Repair, and Evolutionary Optimization @ University of Michigan - Ann Arbor

Related