news
The scoreboard may be crediting the model for the harness's work
A benchmark that varies the harness, not just the model, and finds the pairing matters as much as the choice of model.
research
What changed
Harness-Bench, submitted 27 May 2026, evaluates agent harnesses rather than models: 106 sandboxed tasks and 5,194 execution trajectories across multiple model and harness pairings under shared budgets and protocols. The authors report substantial variation in completion, process quality, efficiency and failure behaviour, and argue capability should be reported at the model-harness configuration level rather than attributed to the base model.
Why you should care
The benchmark you are reading probably credits a model for work the harness did, which matters the moment you have to pick one.