On a Friday afternoon in August, someone sent me a link and four words: Investigate what this means for me?
The link was a product page. A lab had open-sourced an agent harness — the layer that sits around a model and manages context, tools, state, and recovery. The page led with a phrase: Agent = Model + Harness.
My first read was that a major lab had handed me the vocabulary for something I’d been building for months. I run a research practice on exactly that layer.
I was wrong within the hour. Not about whether the layer matters — about the idea that anyone had handed me anything.
The Friction
I run a landscape scanner: a structured sweep of the practitioners I track — who published, who’s converging, who’s challenging me. It has been adversarially rebuilt eight times, and it gates my own publishing while unresolved threats sit open. As of that Friday it had produced 29 dated scans since March, tracking 262 publications, logging 115 obligations.
I ran it. Forty minutes, three facts, ascending order of how bad they were.
First: “Agent = Model + Harness” isn’t the lab’s phrase. It’s the stated equation of a benchmark paper — 106 sandboxed tasks, 8 model backends, 6 harnesses, 5,194 recorded runs. Cobus Greyling had published a full write-up under that exact title on June 2, roughly ten weeks earlier. He isn’t obscure — he writes in an enterprise conversational-AI lane I don’t read.
Second: Sam Thomas Davies published “Everyone’s Arguing About Models. The 30x Was the Harness.” on August 3. He’s been on my watchlist since April at WATCH tier — the second-highest — and my own notes from August 14 call him “active, strong convergence.” He was translating the idea for knowledge workers: not my job, but the same layer. Eighteen days passed. My scanner ran during that window and did not register him.
Third, a different problem than the first two: Ben Dickson published “The art of AI harness engineering” on April 7. I didn’t have to discover it — it was already in my watchlist. I’d read it, noted it, and filed it SKIP: an established analyst publication that doesn’t attempt what I’m doing, so not a candidate for engagement.
The Build
What I built that afternoon was a failure model of my own instrument: a four-stage chain running search vocabulary → candidate space → disposition schema → retained signal. Locating each failure on that chain separated two problems I’d been treating as one.
Stages one and two — where Greyling and Davies were lost. The scanner already knew about this failure mode. Version 8 had added a vocabulary sweep, for a reason preserved in the file: Davies had been “an accidental discovery because his vocabulary was invisible to TIE-vocabulary searches.” The system diagnosed itself correctly, months earlier, and built a countermeasure.
That countermeasure searches seven terms. Every one descends from a concept I had already named: sessions forgetting what came before, knowledge bases that accumulate without compounding, governance files, decision logs.
So the repair for my search terms are too close to my own vocabulary was a list of search terms written in my own vocabulary. Greyling and Davies were never rejected. They never entered the candidate space, because nothing in the query could reach them.
Stage three — where Dickson was lost, and it isn’t the same failure. He made it all the way through: found, read, assessed, recorded. The loss happened at the disposition schema.
SKIP was the only field. The record had one axis: is this person worth engaging? The judgment was correct on that axis — his publication doesn’t attempt the operator-layer work I’d be commenting into. But there was nowhere to put the sentence that mattered: not worth engaging, and using a word you don’t have. A one-dimensional schema was doing a job that needed two, so a signal that survived detection did not survive filing.
One failure prevented candidates from being seen. The other discarded a signal from a candidate already seen. They compounded — Dickson’s lexical signal was exactly the one that would have widened the term list — but they aren’t the same defect and don’t have the same fix.
One more thing belongs here. I ran an independent verification pass over the scan’s own findings. It returned three errors, two in claims the scan had stated with confidence. One was not a misreading but an invention: the scan characterized a practitioner’s argument as “better models shift harness complexity rather than eliminating it,” when reading him showed the argument runs the other way — and considerably harder on me.
The scan fabricated it. I read it and didn’t catch it. The verification pass did.
That sequence is the accurate one and I want it on the record in that order, because the middle step is mine. The report contained a section titled “Unverified — do not repeat as fact.” I wrote it in the same sitting. It didn’t catch the fabrication, because by the time I reread it, the fabrication no longer felt unverified. It felt like something I knew.
The Insight
I’ll name the first mechanism, because it’s the one the architecture demonstrates.
Vocabulary Lock: the failure mechanism created when a discovery system derives its search vocabulary primarily from the concepts already represented inside it.
In this scanner, Vocabulary Lock produced Detection Debt — the condition I named in July, where a consequential gap exists and nothing flagged it. Detection Debt can arise many ways: missing instrumentation, wrong thresholds, coverage never built. Vocabulary Lock is one route to it, and it’s the route this scanner took. A demonstrated instance, not a general theory.
What makes it worth naming is that it fails without a symptom. A filter that rejects too aggressively gives you something to argue with: you see what was rejected and can overrule yourself. Vocabulary Lock produces no rejections at all. The unfamiliar practitioner never enters the search space, so they never appear in any log as a decision I made. Results come back every time, and they look complete. A bounded search and an exhaustive one return the same shape of output, and from inside there is nothing to tell them apart.
The evidence isn’t the miss — one miss proves little. It’s that the scanner had already identified vocabulary invisibility as its problem, built a remedy, and generated that remedy’s terms from the same concept index it was supposed to escape. The countermeasure inherited the boundary it was designed to cross. That’s visible in the file, not inferred from an outcome.
There’s a harder version I have to sit with. I published a piece in April called “Accumulation Is Not Compounding.” Between March and August the watchlist grew to 262 entries. Its ability to detect an unfamiliar vocabulary did not improve, because none of that growth touched the mechanism that decides what gets looked for. I wrote the essay about that failure mode, then spent five months running an instrument that preserved it.
The Honest Part
What I’ve shown is one mechanism in one instrument. Not that Vocabulary Lock is common, or that it’s what usually goes wrong in monitoring systems. Only what went wrong in mine — and I can point at the line of the file where it did.
141 of the 262 entries are marked SKIP. That number describes how much material passed through a one-axis schema. It does not describe how many signals were lost — I haven’t audited them, and until I do, the 141 is exposure, not failure count. It would be convenient to let it imply 141 near-misses. It doesn’t.
The counterfactual isn’t available to me. Catching Dickson in April might have changed nothing — I might have read him, filed him, and drawn no conclusion. The version of me who was four months ahead exists only because I now know what I was supposed to notice.
The correction came from outside. The scanner caught this because someone pasted a link into a conversation. I have no evidence it self-corrects on this axis — only that it performs well once an external prompt has already breached the boundary, which is the one condition under which the mechanism I just named doesn’t apply.
The repairs are unbuilt — both of them, because there were two defects. Widening the source of the search vocabulary addresses Vocabulary Lock. Separating engagement disposition from lexical signal addresses the schema loss that swallowed Dickson. Both are ideas I had on a Friday; neither has been built or tested. This is a diagnosis. The fixes are claims about fixes.
What This Is Actually About
The transfer is a hypothesis, so I’ll state it as one: any system whose discovery vocabulary is generated only or primarily from its existing concept set may have this shape — compliance sweeps, competitive monitoring, hiring filters. I haven’t examined any of them. What I can say is what would settle it — the question I should have been asking about my own scanner eight versions ago.
Not did the monitoring return results. It always does.
Does this system have any mechanism by which a word it does not use can become a word it searches for?
Mine had one. I wrote it myself, out of words I already used.
Case Study Insight: A discovery system that generates its search terms from its own concept index inherits its own boundary — and does so silently, because a bounded search and an exhaustive one return the same shape of result. The test isn’t whether monitoring returns something. It’s whether the system has any route by which unfamiliar vocabulary can enter the next query.
Robert Ford builds products, writes stories and essays, and publishes The Intelligence Engine — a practitioner research publication about AI systems that compound. His other writing lives at Brittle Views.
How this was made: drafted in working sessions with Claude, revised across multiple rounds I read and scored myself. The judgment — what’s true, what’s cut, what ships — is mine throughout, including this line.


