Classification

Three Empirical Tests of Harness Scaling

Harness scaling investigates whether external execution controls can convert a model's latent capabilities into reliable work through three progressively stronger empirical evaluations:

  1. Fixed-Model Improvement (Within-Model Lift): Assesses whether a refined execution harness can improve the end-to-end task completion rate of a fixed model without altering its underlying weights.
  2. Cross-Generation Transfer (Frozen Transfer): Assesses whether an execution control profile developed on an earlier model generation can transfer directly to a newer, more capable model without retuning or modifying control logic.
  3. Beyond-Benchmark Generalization (Held-Out Task Generalization): Assesses whether the resulting control principles and runtime mechanisms can improve performance on completely different, held-out task distributions beyond the benchmark on which they were originally developed.

0

1

Updated 2026-09-21

Tags

Prep Sessions

Long-Horizon Agent Reliability: Stateful Scaffolding and Runtime Verification @ University of Michigan - Ann Arbor

Ch.1 Foundations and Operational Challenges - Long-Horizon Agent Reliability: Stateful Scaffolding and Runtime Verification @ University of Michigan - Ann Arbor

Harness Scaling Foundations and Control Design Space - Long-Horizon Agent Reliability: Stateful Scaffolding and Runtime Verification @ University of Michigan - Ann Arbor

Related