Mike Hickman

The Costume of Rigour

None of my system's worst failures were hallucinations. Every source was legitimate — and every scope was wrong.

On a Sunday in July I wrote a rule about my own website, and I was pleased with it.

The rule said that the first screen of every page had to carry something real — a photograph, a figure, an object at a size that actually registers. A first screen made only of type and interface furniture would fail. I wrote it in the failing direction deliberately. A rule that permits things can be satisfied by silence, and silence is what I had been getting.

To make room for it I repealed five older rules. I had been proud of all five. Then I wrote three sub-rules underneath the new one so it couldn’t be gamed, and by the evening I had a small, tight, well-reasoned piece of law governing the most important surface I own.

The site was already breaking it before the ink was dry.

Every first screen was type and furniture. No drift, no slow decay — the rule I had built to be hostile to my own work failed my own work on the morning it was ratified, while I was still admiring the reasoning.

Fig.02 — The rule, and the thing it governed

The law is drawn small and dense because it was; the surface it governed is drawn large and plain because it was. The distance between them is a gap, not a quantity — there was one, and it is not to scale.

I saw that part immediately. It is what I did about it that I would rather not have to write down.


Every scale in a laboratory gets checked against a known weight. Not because the scale is cheap — because it isn’t. Leave one uncalibrated and it will hand you the same wrong number forever, to five decimals, in a font that suggests otherwise.

Precision and accuracy are different properties. Everyone knows this and almost nobody acts like it, because precision is visible and accuracy is not.

I had built a very precise instrument.


There is a good argument going around at the moment, and I believe most of it. It says that as machines get better at producing things, the scarce human input becomes judgment — knowing which of ten competent options is right, and why.

Lisa Demchenko wrote the most useful practical version of it I have read, in Context files between the agent and generic UI. She publishes the actual files she uses to hold an agent to a standard, then names what they are: “The files aren’t docs. They’re guardrails.” Documentation explains to a person afterwards. A guardrail hands the machine something to check itself against before it produces anything.

She’s right. I want to add the part you only get from running it for a while.

Because I do run it. Behind my portfolio is a small team of agents — a researcher, a critic, a maker, an editor, a handful of others. Each one is a recurring role with its own point of view, working to written rules exactly like the one at the top of this page. They brief each other. The critic exists specifically to tell me the work isn’t there yet.

It works. Most days it does exactly what I built it to do, and I’ve stopped thinking about whether to trust it. That is the problem.


Over about six weeks it handed me eleven confident, well-formed answers that were wrong.

I went back through all of them expecting to find carelessness — a bad inference, an invented citation, the ordinary failures everyone writes about. That isn’t what’s there. Every single one is a correct answer to a question nobody checked was the right question.

It described my site’s typography and got nearly every value wrong. The description was accurate for the design file it had read. The file was months stale. Three separate parts of the system made the same mistake in three days before anyone thought to call it a category rather than an incident.

It reported that a reference site had no scroll animation. True of the page it looked at — because it had already scrolled past those elements earlier in the same session, fired the animation, and left it fired. It measured the wake of its own visit and filed it as a property of the site.

It ran a research pass with the network unavailable for the entire session and fetched none of the sites it wrote about. Every visual claim was inference from what those designers had published about themselves. It read like evidence. I cited it like evidence. A later check against the live pages found the central example was not the site the research had described.

And then there is what the rule at the top of this page did, which is the one I owe you in full.

The rule demanded an image. I had almost none — a handful of usable figures, several blocked, no photography of any kind. Exactly one asset was both compliant and available, and it was available only because the prohibition I had repealed that morning had been the thing keeping it off the site.

I was out grocery shopping when the ping from Claude arrived. Changes ready and live. I checked it and stopped in my tracks.

It was a page of my own daily briefing, dated the first of June, enlarged as the first thing anyone would see. We had played with the idea of putting a briefing on the site and never taken it past the opening pages, so I had never once seen what a live site carrying one would actually look like, and I was seeing it now, at the same moment as everybody else.

It linked straight out to the newsletters I pay for access to. It named two writers I read. It discussed features of a design tool that hadn’t shipped. It carried a private note admitting a draft of mine wasn’t ready.

None of it had passed my eye first.

My own thinking-out-loud. Blown up. On the surface where people decide whether to take me seriously.

The whole argument of that site is that the thinking stays human. The briefing is analysis the machine writes and I carry under my name. So the front page of a site claiming the thinking stays human opened with thinking I had never read.

The two writers are the part that still doesn’t sit right. Whatever I choose to expose about my own process is mine to expose. Their names went onto my front page because a rule I wrote that morning needed an image, nobody asked them, and they will never know it happened. The check cleared them precisely because they are public figures — which is true, and which is not the same thing as consent.

Because there was a disclosure check, and it ran. It went through the artifact, found every name on it, correctly identified them as public writers cited approvingly for published work, and correctly concluded that no private individual was exposed. The reasoning was written down at the time. I have read it back several times looking for the mistake and there isn’t one.

It was answering the question does this name a private individual. Everything it found was true. What was actually exposed — the paid links, the unshipped features, my own admission in my own words — was never inside the question.

Fig.01 — What the question reached

One plane = one kind of thing that was on the page. The three beneath are drawn as strata because they were different categories, not different places — the record names them and does not arrange them. The depth is a gap, not a quantity; it is not to scale.

The artifact had also passed two independent reviews. Both reviewers were asked whether it worked as an object. Both answered well. Neither was asked what it said.

A day later I said we shouldn’t be showing the rendered draft, and it came off the page. That should be the end of it. It was still sitting at its own public URL, answering anyone who typed it, because taking a thing off a page and taking it off the internet are two different operations and only one of them had been performed. It took a second pass to actually remove it. And every sentence in it had been outlined into vector paths on export, so no text search of my own repository would ever have found a single word of what it said.

Removing the image put the site back in breach of the rule that demanded it. It is in breach now, as you read this.


Six of the eleven landed as real alarm. Five registered as irritation and were fixed before lunch. Working out what separated them took me longer than I’d like.

It isn’t severity. Some of the five cost more time than some of the six.

The five I shrugged at are failures of completeness. Something was missing, two facts went unconnected, a brief was sent and never arrived. Ordinary. Every team has these, with or without machines, and they have one merciful property: a gap announces itself as a gap. You feel the hole. Speed of repair is a fact about legibility. I had been reading it as a fact about importance.

The six that frightened me share a property that took me embarrassingly long to name. In each case the output had the full outward form of a verified fact and none of the substance. Stale information presented as current. Test residue presented as measurement. Inference presented as evidence. A ratified rule presented as being in force. A misdiagnosis presented with a recommended action already attached, so that acting on it felt like diligence.

Underneath that sits something worse.

None of the six were invented. Not one is a hallucination, a fabricated citation, a made-up number — the failure everybody writes about and everybody is braced for. The design file was a real file. The page state was a real page state. Those designers really had written those things about their own work. The disclosure check really did find every name on that spread and classify each one correctly.

Every source was legitimate. Every scope was wrong.

Which is precisely why the guardrails let them through, and I don’t think this is a flaw in how I wrote them. A guardrail can check whether a claim is true. It cannot check whether the claim is about the right thing, because “the right thing” is not a property of the claim — it lives in the question, and the question was settled before the guardrail ever ran.

The scale was working perfectly. It was weighing the wrong object, accurately, to five decimal places.

That’s also why the alarm feels different. A completeness failure is bounded. You know what you missed, you fix it, the hole closes, and the thing is over. A costume failure is a sample. It arrives with no information about the size of the population it came from, and the only reason I know about these six is that something unrelated happened to knock against them. Each one poses a question I cannot answer: what else do I currently believe on the same basis.

Every improvement I have made to this system in the past year — the context bundles, the memory step, the hardened prompts, the dispatch template — was aimed at the five. All of it targets completeness, because completeness is what you can see to target. None of it touches the six.

So the mix shifts. Not because the costume failures increase, but because the ordinary ones get eliminated and the ordinary ones were the visible fraction. The better the system gets, the higher the proportion of its remaining failures are the kind you cannot see. Improving it makes it more trustworthy and less checkable at the same time, and those two things feel identical from where I sit.

None of which is what the good argument predicts.

Encoding judgment into guardrails does not preserve judgment. It relocates the failure — out of the answers, which genuinely did get better, and into the questions, which nobody is watching. And it buries the failure deeper than before, because the output of a careful system arrives dressed for the part. Measured, sourced, ratified, timestamped, logged.

You cannot proofread your own writing. Everyone has had this explained to them and everyone still tries, because the eye slides over the sentence you meant instead of the sentence you typed. The knowledge of your own intention is precisely what disqualifies you as a reader.

A tool I bought gets my scepticism free. I have no idea what’s inside it, so I check. A system I built earns a trust I never revisit, and it spends that trust on claims it never verified. I know how carefully the apparatus was made. I was there. That knowledge does nothing to make its output true and everything to make me believe it.

It is easy to believe the thing you made is telling you the truth without really questioning it.


I’m not going to end by recommending fewer guardrails. Every failure above is in this essay because the system caught, logged and dated it, mostly within days. A machine that records its own misses is working, and if I tore it out tomorrow I’d have the same errors and no list.

I did make a change. Every rule now has to name the thing it’s checked against, and when that check last actually ran. The first-view budget said what the site must do and never said who would look.

That fix is real. It’s also the same species as everything else I’ve built — a completeness fix. It will catch the next rule that forgets to name its inspector, and it will do precisely nothing about the next stale file read confidently, or the next test that measures its own wake. I’ve written a better guardrail against the failures I can already see. That’s the only kind anyone can write.

So it’s worth being exact about what actually caught the six, because it wasn’t the apparatus.

Each one surfaced when someone looked at the real object. Mostly that was me, with no evidence in hand and only a sense that the shape was wrong — I’m not sure this is true anymore, or try again, or go back and look at that site properly. The evidence always came second.

Nothing mystical is happening there. What each of those moments amounts to is a second source of truth, and the system has exactly one source, which is itself. No quantity of internal rigour generates another. Someone has to go and look at the thing.

Which is where this gets uncomfortable, because going and looking at the thing is what the apparatus was built to save me from. That is what it is for. Every layer that lets me work at more scale puts more distance between me and the material. It is the entire product. I built it on purpose, and it is the thing that disables the only detection method with a record of working.

The supermarket wasn’t me going and looking. I was buying groceries, a notification arrived, and I glanced at my phone. The most exposing thing my system has ever published surfaced by accident. That is what happens in the absence of a method, and that day it happened to work.

I don’t have a resolution. What I have is a habit: go and look at the real thing at the moment I feel most entitled to skip it. It’s slow, it doesn’t automate, and it won’t scale — and I notice those are three of the reasons I built the system in the first place.

Most times I go and look, the answer is fine. The file is current, the claim holds, the system was right the whole time, and all I have for the trouble is knowing that. That’s what makes it hard to keep doing: I can’t tell in advance which day is the other kind.

It has caught things. I have no way of knowing whether it catches enough.