Tashika
2026-09-02 · 6 min read

06 — An Instrument, Not an Opinion

Ask the room whether a system is understood and you sample the very faculty whose absence you're trying to detect. An instrument has to read the system instead.

06 — An Instrument, Not an Opinion cover

There is a version of this that requires nothing to be built. Walk into the room and ask: does anyone here still understand this system? People will answer. Often they will answer well — a good senior engineer has a real sense of which parts of a codebase are solid ground and which are fog, and that sense is not worthless. It is the best thing most organizations currently have. It is also an opinion, and the distance between an opinion and a reading is the distance between feeling something and knowing it.

The usual complaint about opinion is that people are biased, or political, or reluctant to admit what they don't know. All true, and all beside the point. The fatal problem here is more specific, and it is structural: asking whether a system is understood samples the very faculty whose absence you are trying to detect. Understanding that has quietly drained does not leave behind a person who knows it has drained. It leaves behind a person who has not lately been asked. They will tell you the system is fine, and they will be telling the truth as they can see it, because the thing that went missing is precisely the thing that would have told them it was gone. Only the parts that are still there can speak.

So an instrument has to read the system rather than the room. That constraint alone rules out most of what would otherwise be the obvious approach — the survey, the confidence score, the self-assessment, the maturity model. Each of those is opinion with a coat of arithmetic on it. Averaging twelve people's sense of how well they understand something does not convert sense into measurement. It converts one opinion into twelve, and then back into one.

The second requirement is dull and non-negotiable. The same system, read twice, has to give the same reading — and a system that has changed has to give a different one. That sounds like the lowest bar available until you notice how little clears it. A reading that moves when the team is anxious, or when a reorg is pending, or when the person taking it has a stake in the answer, is not reading the system at all. It is reading the conditions around the system. A gauge whose needle responds to the weather in the room is a barometer for the room.

The third requirement costs something to accept. The reading has to be capable of being wrong in a way you find out about. An instrument that can never be contradicted is not cautious, it is empty. So the incident that lands squarely in a region the gauge called well understood is not an embarrassment to be explained away — it is data about the gauge, and the only kind that improves one. A measurement discipline is a thing that keeps a record of its own misses. Most of what gets sold as measurement in this industry has no such record, and could not produce one if asked, because nothing about it was ever specific enough to miss.

The fourth requirement does the most work, because it is where the obvious approach quietly breaks. The natural way to read a system now is to point a model at it and ask. Models are very good at this. They read the whole thing, they never tire, and they produce fluent, confident, specific accounts of what a system does. The trouble is that the account is fluent and confident and specific whether or not it is right — and the same was true of the code the model would have written. Fluency used to be a usable signal, because a person could not sound like that about a system they had not understood. It is not a usable signal any more. A report produced that way tells you what the apparatus emitted. It does not tell you about the system.

Which sets the rule the whole setup is built around, and it is plain enough that anyone can audit it: nothing grades its own homework, and no one artifact is asked to do two jobs that fight each other. A spec that generated the code cannot also be the thing that certifies the code — the code passing it demonstrates that the generator ran, not that the behavior is right. Tests written by reading the implementation cannot tell you the implementation is correct; they can only tell you it still does what it already did. A model asked to assess its own output is being asked to be both the witness and the court. Every one of these is the same mistake wearing a different costume: a loop closed on itself, reporting its own consistency, and being read as evidence about the world. Breaking those loops is unglamorous, and it is most of the engineering.

What comes out the other side is not a grade. That is worth being blunt about, because a grade is what people expect and what they would find easier to act on. It is a location. Three readings that move independently — what is written down where a machine can check it, what survives only in people you can still reach, what survives in neither — taken of a named system at a named moment. The value is not that some number came back high or low. It is that the three come apart, and each one names a different thing to do. Well pinned and unheld means the work is rebuilding understanding of something already captured. Deeply held and unwritten means the work is writing it down before the person holding it walks out. Those are opposite instructions, and no single anxious number has ever been able to tell them apart.

A reading that comes back fine is also a result, and one worth paying for. Most of what an organization dreads is not distributed the way the dread is. The worry is general; the exposure is specific, usually smaller than the worry, and usually somewhere other than where everyone was looking. Being told that the system you were nervous about is in good shape frees attention for the one nobody had flagged. An instrument that only ever confirms alarm is not an instrument either.

It is worth saying plainly what this does not do. It does not tell you whether your team is good. It does not tell you what to do next — a reading is an input to that decision, not a substitute for making it. And it does not convert the unaccounted-for into the accounted-for by naming it: the third quantity gets bounded and cornered, not abolished. The claim is narrow on purpose. The gap has a reading. That is all of it, and it is the thing nobody had.

An opinion about a system tells you about the person holding it. A reading tells you about the system. For fifty years those amounted to the same thing, which is why nobody needed the distinction. They have come apart, and the distinction is now the product.

Which raises a commercial question with an obvious answer, and the obvious answer is wrong. Once you can see which systems have gone illegible, the money looks like it is in fixing them. Everyone who can see the damage sells the repair. We don't — and the reason is not modesty.