The Sentence That Wasn't There
I almost quoted a sentence that wasn’t in the paper.
Cobus Greyling’s write-up of a new benchmark this summer had a line sharp enough to build an essay around: the harness is a crutch, and its value decays as the model improves. It was quotable, it agreed with something I’d already argued, and it was ready to go straight into a case study as external validation. Then I read the paper instead of the write-up. The sentence isn’t in it. It’s Greyling’s own compression of the finding — a sharper line than the section it came from, doing the job a write-up is supposed to do, but not one I could cite as the paper’s conclusion.
What the Benchmark Actually Found
Back in March, in an essay called “Governance as Scaffolding,” I argued that governance — the constraint files, the decision logs, the adversarial review passes I run on my own work — is transitional scaffolding, not permanent infrastructure. The goal was for the system to build itself out of a job. That was a stated position with no measurement behind it, and I said so at the time.
The paper Greyling was writing about didn’t measure that — it got close to it. Harness-Bench (arXiv 2605.27922) ran 106 sandboxed tasks across 8 model backends and 6 different harnesses and found a 23.8-point spread between the weakest and strongest scaffolding on the same tasks with the same models. A light harness scored 76.2 in 7.3 turns; a heavy one scored 71.2 across 22.6 turns and 139.7K tokens — a result that undercuts the assumption that a heavier execution stack necessarily buys better performance, without the paper claiming the extra scaffolding caused the worse score. Separately, the paper found that spread narrows as the backend model gets stronger: cross-harness variance declines toward the high end of model capability. Its own section on the question — “Do stronger models make harnesses less important?” — answers it directly: “stronger models may reduce the need for prompt-level scaffolding and simple procedural guidance. However, they still require reliable execution substrates: permission boundaries, persistent state, interpretable traces, evidence records, and objective verification.”
That’s not the paper confirming “scaffolding decays.” It’s adjacent evidence — performance varies by harness, and that variance narrows for stronger models — without the paper decomposing which harness mechanisms compensate and which record. The two-part split is mine, extracted from the collision between that finding and my own process, not something Harness-Bench set out to measure. It was sharpened into two claims where I’d only made one.
Two Kinds of Scaffolding
The distinction the paper draws is between what a harness does for a model and what it does about a model. Procedural pre-checks exist because a weak model sometimes skips a step it was told to run first. That compensates for something the model can’t reliably do on its own, and it loses its reason to exist the moment the model can.
The forbidden-list pre-check that caught the Greyling sentence above is exactly that kind of scaffolding. It exists because the step of reading my own content rules before recommending a topic is one I have, on at least one occasion, simply skipped. That’s the compensating half, working as designed, catching a mistake a smarter version of the same process still might have made — but built to become unnecessary the day it stops finding anything.
The other half doesn’t compensate for anything. It’s the record of what got decided and why: that a claim sourced from a write-up gets provisional status until the primary paper is read, logged the same afternoon this happened; that a case study missing its third structural leg gets reclassified rather than published with the gap disclosed, a ruling made in one specific case and now binding on every one since. A stronger model arriving tomorrow doesn’t get to skip either of those. It wasn’t slow at inferring them — it never had access to them. They aren’t a capability gap. They’re private history, and nothing about a model release supplies history it wasn’t there for.
Which means the real question isn’t whether a harness will still matter in a year. Some of it won’t, and that’s not a loss — it’s the scaffolding doing its job. The question is which half of what you’re building is compensation and which half is record, because only one of them has an expiration date.
The Honest Part
I don’t have this cleanly separated in my own files, and saying the distinction exists is not the same as having applied it. My constraint files mix both kinds freely — a formatting rule and a decision ruling can sit three lines apart with no marker distinguishing which is which. I’ve never gone through and sorted them. Until I do, the claim that “the accumulating half survives” is true of the idea and unverified against my own corpus.
There’s a harder case against it, too, and it deserves to be stated at full strength rather than pre-defeated. Hugo Bowne-Anderson’s version holds that every harness feature is a bet on something the model can’t do — and that once you frame it that way, nothing is exempt, because even a decision log is a bet that the model won’t rediscover the reasoning on its own. I don’t think that reaches the record-keeping half: a bet resolves when the model gains the capability being bet against, and a stronger model gains capability, not access to a specific afternoon it wasn’t present for. But the argument is real and it’s the strongest one currently in circulation.
And one piece of my own evidence cuts the other way, and I don’t have a clean answer to it. A concept I named in March — the tax of starting over from zero every session — is measurably less true today than when I coined it, because memory features shipped into the tools themselves absorbed part of what I was pointing at. That looks like exactly the counterexample Bowne-Anderson’s argument predicts: a piece of my own record, retired by a capability gain. I don’t know that it is one. It’s just as possible that what moved was the harness underneath me, not the model — memory shifting down into the product itself rather than the model getting smarter — and right now I can’t tell those apart. But I named the concept as if it were permanent, and it wasn’t, and that’s a real miss, not a technicality. I have one instance, and it’s close enough to the boundary that I’m not fully confident where it sits.
The Brace and the Record
A harness that’s just remembering for you was never the part that expires.
The brace comes off. The record doesn’t. I haven’t sorted my own files into which is which yet.
How this was made: drafted in working sessions with Claude, revised across multiple rounds I read and scored myself. The judgment — what’s true, what’s cut, what ships — is mine throughout, including this line.
Robert Ford builds products, writes stories and essays, and publishes The Intelligence Engine — a practitioner research publication about AI systems that compound. His other writing lives at Brittle Views.


