Decisions · Open where closed, closed where open
Mistakes both ways
Exact work handed to a model and open judgment handed to a branch tree are mirror images with one asymmetry between them: the first leaves evidence, and the second passes its own tests. That asymmetry is why teams swing instead of converging on correct placement.
12 min read
Cover for Mistakes both waysTeam A sends portion arithmetic through a language model, because the roadmap said the assistant would handle the whole flow. The evaluation suite reads a hundred percent on those links, and the failures, when they come, are arithmetic wearing the costume of reasoning. A bill arrives every month for work a comparison operator would have done for nothing.
Team B encodes busy-evening practicality as a branch tree: reject the meal when active effort exceeds twenty-five minutes. The exceptions begin arriving within a month, one for meals that finish unattended, one for leftovers, one for the week the spinach has to be used. Every branch was locally correct, nobody can say why the number is twenty-five rather than thirty, and the tests pass on all of it.
Both teams put a decision on the wrong side of the same sort, in opposite directions, and both believed they were being responsible when they did it. What separates the two errors is what each one leaves behind. Closed work left inside cognition announces itself in evidence teams misread as good news, while open work forced into code passes its own tests, and that asymmetry of visibility is what makes organisations swing between the two rather than converge on correct placement.
Helpful context: Open, closed, and closing owns the sort both teams failed. Closed decisions belong in code is the constructive rule when that sort returns closed, and contested decisions covers work that has no correct side to be pushed toward. Judgment terrain (next in this track) explains why some judgments resist specification even when they sound like rules.
One sort, two ways to fail it
The sort behind both failures belongs to open, closed, and closing, and what it returns is an address rather than a difficulty rating. Portion arithmetic comes back closed, so its address is deterministic code. Busy-evening fit comes back closing, so its address is a named cognitive unit carrying a suite. Neither team disputed those answers, because neither team ran the sort.
Each error is what happens when a link lands on the wrong side, and the two costs are comparable in size while being different in kind. Closed work inside a cognitive unit pays money and latency for a comparison, and it takes an operation that was right every time and makes it right almost every time, which is a downgrade bought at a monthly rate. Open work inside a branch tree pays nothing per run and gives up the thing that judgment needed most, because there is nowhere for a correction to land except as another branch.
The two costs are symmetric. What the two errors leave behind is not, and that is where they stop being mirror images.
The mistake that leaves evidence
The over-open error is the cleaner of the two, because it is worse on every axis at once and each axis leaves a trace. The first trace is the evaluation score. A cognitive unit that is right on every case in a properly built suite is reproducing a rule, so the correct response to a hundred percent is to go looking for that rule in the cases you have just graded, rather than to circulate the number.
The second trace sits in the reasoning, and closed decisions belong in code turns it into a working test, because grounds with a fixed structure and only the values changing are a template with a rule inside them. Failures then confirm the reading, since a cognitive unit doing genuinely open work fails by weighing something differently than you would have, while one handed closed work fails by losing a currency or missing a date boundary by one. Those two textures are hard to stop telling apart once seen, which is what makes this the trace an engineer can act on without waiting for a bill.
The cognitive unit · Ch. 5A closed decision placed inside a cognitive unit is worse on every axis at once, which makes it an unusually clean kind of error.
The third trace arrives on the invoice, and it is the only one anybody outside engineering can read. The price of closed work left open does that arithmetic properly. What matters here is that all three traces are legible to somebody, which is why this error gets caught at all, and why teams so often circulate the first two before understanding what they mean.
The defence organisations reach for is a learning narrative, usually some version of the model getting better at dates. Closed work does not improve through repetition inside a template, since there is no gradient to climb once the answer is already determinate. It improves by being endorsed and moved into code, and until somebody does that, the repetition is what the monthly bill has been buying.
The mistake that passes its own tests
The over-closed error is older than this technology and much harder to see, because nothing it produces looks like a defect. The tree is green, coverage is respectable, and the branch added in March did exactly what the ticket asked for. What the tree cannot do is account for itself, since there is no suite that could disagree with it, no name for the judgment it is approximating, and nowhere for a household correction to land.
That absence is the harm, and it is a harm of a different kind from the first one. Team A wasted money on work that stayed correct throughout. Team B produced a decision with no owner, so nobody can say whether busy-evening fit improved or degraded last quarter, and nobody could answer if asked. The threshold is doing the work of a judgment while carrying none of the obligations of one, because it has no declared inputs, no evaluation, and no record of who set it or why.
The tells for this error sit in artefacts the team already has, and where rules stop working reads all three of them: a branch chain that grows every month without converging, a threshold nobody can justify, and complaints in the register of not understanding my case rather than being broken. What is worth adding is where those tells go once they exist. Each one arrives in a channel reporting to somebody other than the architect, which is why this error runs for years inside teams that are genuinely paying attention to the other one.
The repair starts by giving the judgment a name. AssessMealPracticality decides whether a candidate fits one evening, so the endorsed effort budget arrives as an input rather than as a constant buried in nested conditionals, and corrections land against the cognitive unit that owns the call because there is finally something to land them against. A suite can then disagree with it. None of that makes the first answer better, since naming a judgment does not improve it, and what changes is that a wrong answer now has one address and an owner.
Why teams swing between them
Put the two sections together and the pendulum explains itself. The over-open error announces itself in three places, so it gets a hardening sprint. The over-closed error announces itself in a support queue nobody routes to architecture, so it accumulates quietly. A team therefore meets the two errors at very different rates and concludes that one of them is the real risk.
The swing follows from that conclusion. A painful model failure on exact work triggers a hardening push, so thresholds go into code faster than anybody checks whether the decisions behind them were closed. Eighteen months later a rigid tree blocks a legitimate case, and somebody reopens the judgment inside a prompt to add flexibility, which arrives with no name and no suite because nobody was asking for either. Neither move was unreasonable given what the team could see, and the sequence still leaves the system further from correct placement than it started.
Leadership narratives set the amplitude. Being an AI company pushes placement open, needing guardrails pushes it closed, and both are answers to a question about materials rather than about links. The corrective is to put both columns on one review agenda, so portion arithmetic sitting in a model call and busy-evening fit sitting in a branch tree can be named in the same meeting with different remedies attached and no narrative to resolve first.
Closed work in cognition
Money, latency, and variance spent on a comparison. Evaluation reads perfect, and the perfection is the symptom.
Open work in code
A branch tree standing in for a judgment nobody owns. Tests pass, and the complaints land where the tests cannot see them.
Placement is a direction rather than an audit
Correct placement moves, which is why a one-time audit produces a snapshot instead of a habit. Decisions drift one way as knowledge accumulates, from open through closing into code, and how decisions close traces the mechanism. They do not drift back on their own, so when a decision does move back it is because a person classified it wrongly rather than because the system degraded.
The direction has an end, and the end is not maximum code. Contested preferential work may never close, and closing it would be harm even where the tree could be written, because the tree would answer a question the organisation has never decided. Contested decisions covers what a system owes a reader in that state, and the short version is to surface the disagreement rather than let a fluent answer settle it on the organisation's behalf.
Placement per link is also the only actionable version of any of this, since a feature can be making both errors at once on different links of one user journey. Most decisions are several decisions is what makes the per-link push possible, because until a request is cut into links the two mistakes get averaged into a single impression of how the feature is doing, and an average of a loud error and a quiet one reports the loud one.
What to do next
The loud mistake will find you, so the afternoon is better spent hunting the quiet one. Pick the part of the product nobody has escalated, ideally a rule-based feature with a stable suite and no recent incidents, and read its support tickets instead of its tests. Count how many say the system did the wrong thing for this case rather than the system is broken. A constant rate of the first kind, six months after the last exception branch was added, is the signature of a judgment that has been living in a branch tree the whole time.
Both directions of misplacement survive code review for the same reason, which is that the question they answer is about the decision rather than about the implementation, and no diff shows it. Judgment terrain supplies the six conditions that explain why a judgment resists specification even when it sounds like a rule, and reading a link against them is the quickest way to learn which side it belonged on before anybody wrote anything at all.
Read next: Judgment terrain.