In which the rule outlives the reason.
Version Eleven
"Never Active products don't indicate risk."
One afternoon, I wrote that sentence. An alert, and not a stupid one, which somehow upped my annoyance, had come to a conclusion that was understandable. And it was one of the first times I really felt how today's robot helpers fill in the gaps that we humans routinely throw at them.
Often, I imagine my machine helpers like Thomas the Tank Engine, struggling uphill without enough coal or steam or whatever. In this scenario, the fuel is the stuff I sent at my train: language that looked probably like language around customers at risk. Doing what I'd asked it to do, my Thomas treated that as evidence for a very loud alert. Except this customer had never really started using the product, which I knew and which made sense to me.
I understand intuitively something the machine didn't: you cannot stop doing something you never started doing. Ah, but "Never Active" could mean they should have been active but "Oh No! They have never been active!" Some GPU somewhere had that kind of thought, I guess, and didn't compare against the other options in that field, which would have illuminated to it that its interpretation was wrong.
There is no hill, Thomas.
So: rule. Ta-da! Magically make a problem disappear.
My system is called TARS, after the Interstellar robot. To me, naming internal systems is one of life's great opportunities to amuse yourself. Every fifteen minutes, TARS wakes up, analyzes a handful of accounts, and goes back to sleep. Now, the rule I added that afternoon gets loaded as context over and over and over again, and will potentially survive the death of our Sun and maybe the heat death of the universe.
I remember the conclusion perfectly; I can see it every time I look at that language. But I barely remember getting there. Meanwhile, the sentence is doing fine. Just chugging along. Surviving for another day, another run up a different hill.
This happens enough that the prompt has a filename: 2026-02-21-v11.
Eleven versions. Twelve: on the way already.
All those conversations I don't remember having, false alarms, judgments obvious at the time but now obscure, wrong paths gone down. They're in those versions, but I couldn't easily tell you exactly why.
"I have no recollection of that, Senator."
Should TARS trust most of these versions?
I can open the file, find some instruction I don't really remember writing, read it twice, and think, yes, absolutely, keep doing that TARS.
It's this weird human thing that's always fascinated me: some time, long ago, there was presumably a reason good enough for xyz statement to be written down. That good reason is nowhere to be seen, barely remembered. But things are working great, so keep going Thomas. Which makes deleting anything surprisingly difficult.
Is the rule...clutter? I don't want to find out, so let's leave that in.
Now I have to untangle my space robot and train analogies. In my head, TARS is a factory that just spits out a new Thomas every time there's a hill to climb. Each Thomas runs on different fuel and has different component parts. Oh, and Thomas today is a bullet train.
This is where it gets hard to keep everything straight. Whatever model version TARS uses to build today's Thomas isn't necessarily the model that did whatever stupid thing caused me to write down some rule six months ago. My little sentences and versions hang around like train track signs but we're basically flying spaceships now.
TARS? Please figure it out, and make no mistakes.
I have no easy way to know without removing said rule and seeing what happens, which is precisely what I don't want to do. And I'm not going to do that, probably, or maybe I will because the clutter question is becoming the question.
At some point the responsible thing becomes obvious: go through the file and clean it up.
But I'm lost in my own analogies, like you. Is this...fuel? Is it the track?
What exactly am I auditing against?
Even though the rule is visible, enough of the reasoning behind why it's there rarely is. Which makes deciding what counts as clutter more like gardening, a matter of taste. It's hard to get your mind around this.
I could expose every rule TARS uses, put them on a page, let anybody read them, make the whole thing maximally transparent. That still won't ever recreate the afternoon that produced the rule. Or illuminate with the power of 1000 suns which exact model was behaving badly, what evidence confused it, or why this particular sentence seemed like the right fix.
Visibility is but one step toward understanding.
Knowing that will almost certainly not stop me from making version twelve. TARS will do something dumb or ill-advised or wrong tomorrow or next week. Maybe it'll start making Thomas submarines. I'll stare for a while, figure out what happened, and eventually arrive at some sentence that makes the problem go away. At that moment, I'll have little to no interest in conducting a thorough forensic investigation of versions one through eleven. I will want Thomas to get over the hill.
So I'll add something new, like a sentence. And another, and another, and it will accumulate into something unwieldy. Nobody sits down and decides to create an inscrutable pile of old judgment. Each addition makes complete sense from three inches away. The pileup might be right, even.
And yet, I see someone six months from now, possibly me, opening the file and asking why any of it is there.
Footnotes
This is probably why evals are becoming so important. Traditional software leans on determinism: this input should produce that output. We're not in that world anymore, or not nearly as much as we were. I don't ipso facto care if the answers are always the same; what I care about is whether whatever model is reasoning over or about something understands the thing the old rule was trying to protect, e.g., "Never Active products..."
That means I should probably preserve the failure cases rather than agonize over every sentence I once wrote to fix one of them. Each new Thomas can be tested against an old hill that defeated prior Thomases. The trains are changing, and the hills are changing, but we can at least see how our new and improved speed train Thomases do against some old hill while we throw them at new mountains.
Maintenance of rules is rarely fun. Adding one is a local problem: something went wrong right now, you can see what it was, and you have an idea of how you might address it going forward
Removing a rule is harder, stranger, terrible sometimes. A rule that works well gives off evidence of how useful it is. Nothing bad has happened? Great. But is that the rule doing its job over and over, or because five versions of Thomas ago that rule was needed to get over the hill. New Thomas can just handle the hill. Deletion is a scary bet, because it's so hard to remember the original problem and compare to today.
Old rules have a home-field advantage in every argument.
Some of these rules are possibly just ghosts. The file is haunted. Whatever part of the chain needed them is now gone, and the sentence remains because nobody can prove or has bothered to consider that removing it is now safe.
| Published | 5 April 2026 (6 months ago) |
|---|---|
| Reading time | 7 min |
| Tags | ai, automation |
| Constellation | Deep Current |
Reply
I’d welcome your thoughts on this essay. Send me a note →

