Decisions · Six factors around a call
Judgment terrain
Arguments about whether a decision is too fuzzy for software are usually arguments about models. This page covers the six conditions that describe the ground around a decision, why they never combine into a score, and why the reading belongs to one link rather than to a product.
14 min read
Cover for Judgment terrainAn architecture review is deciding which parts of the household dinner planner the system may lead. Somebody writes "fuzzy" beside four links and "exact" beside three, and the meeting moves on. Nothing in that page of notes separates a judgment the household makes fifty times a year from one nobody in the household has ever made, or a wrong answer somebody can shrug off from a wrong answer that cannot be taken back.
Fuzziness describes the question. What determines whether software may lead a decision is the set of conditions around it: how often this household has faced weeks shaped like this one, what the planner has been told about them and can trust, what a wrong answer costs, and whether anybody would find out that it was wrong. Those conditions are the terrain, and they hold still while models improve. Terrain describes the ground a decision sits on rather than the model asked to walk it, which is why a capability answer can never settle a delegation question.
Helpful context: Where rules stop working shows how shifting context, competing goals, and low reversibility break an exact procedure, and hands the rest of the ground forward to this page. Open, closed, and closing asks whether a decision has a determinate answer and whether its rule can be written. Terrain asks a different question: what the ground around that decision will bear while the judgment stays open. Established, mixed, and frontier regions (next in this track) compresses these readings into leadership policy.
What the ground is made of
Six factors do the work, and they are worth grouping before they are worth listing, because the grouping is what stops them from becoming a checklist. Three of them describe what the decision can draw on. Two describe what an error costs. One describes whether an error would ever come to light.
The support pair is pattern density and knowledge coverage. Density asks how often sufficiently similar situations occur, which is a question about repetition rather than volume, since a decision can arrive daily and be different every time while another arrives twice a year in the same shape both times. Coverage asks how much reliable information stands behind the answer, and quality does more work here than quantity, because a large body of outdated, contradictory, or unendorsed material creates the appearance of support while reducing it. Busy Tuesdays in the Meal Companion score well on both, since the household has had dozens of them and the twenty-five-minute effort budget is endorsed knowledge rather than somebody's guess.
Sitting between that support and the case in front of you is novelty, which measures the distance between the two and therefore belongs to the situation rather than to its wording. A familiar request phrased strangely is linguistically novel and operationally routine. A request phrased in the household's usual language can hide a structural difference that puts it outside every strategy the system has, which is what happens when "plan like last Thanksgiving" arrives carrying a different guest list, a different oven, and a different allergy profile. A novelty reading has to ask what changed rather than what rhymes.
The cost pair is consequence and reversibility, and they come apart more often than teams expect. Consequence asks what being wrong costs and to whom, which matters most when the cost falls on somebody other than the operator, since that is the case nobody in the room feels. Reversibility asks how easily the effects can be undone, and it is indifferent to how serious those effects were, so the two readings can point in opposite directions on one link. A heavy meal suggested for Tuesday stays consequential until somebody picks a different recipe, whereas a message already sent or money already spent cannot be recalled by correcting a record.
That leaves evaluability, which sits apart from the other five because it governs whether any of them can be maintained. It asks how clearly the quality of an outcome can be judged, and the answer is usually worse than a team assumes. Where feedback is delayed, contested, or supplied by the system about its own work, a decision can be wrong for a year while every dashboard stays green, which means the terrain reading taken at the start can rot without producing a single symptom.
| Factor | The question it asks | What should raise caution |
|---|---|---|
| Pattern density | How often sufficiently similar situations occur | Frequency without repetition, or precedent belonging to the industry rather than to this product |
| Knowledge coverage | How much reliable information stands behind the answer | Volume that is outdated, contradictory, or unendorsed |
| Novelty | How far this case departs from known patterns | Structural difference behind familiar wording |
| Consequence | What being wrong costs, and to whom | Costs borne by somebody other than the operator |
| Reversibility | How easily the effects can be undone | Anything sent, spent, disclosed, or missed |
| Evaluability | How clearly outcome quality can be judged | Delayed, contested, or self-assessed outcomes |
Why the readings never add up
The grouping also explains the commonest misuse, which is arithmetic. Support, cost of error, and visibility of error are three different quantities, so averaging them produces a number whose components can no longer be recovered. Worse, an average lets a strong reading in one group cancel a fatal reading in another, which is precisely the move the six exist to prevent.
Consequence and reversibility routinely override the support pair on their own. The Meal Companion has assembled shopping lists for forty weeks with dense patterns and strong coverage, and it still may not buy the groceries, because spending cannot be reversed by the system and nobody handed it the household's money. Established, mixed, and frontier regions takes up what follows from that gap between what a system can do and what it may do.
A high consequence reading raises the standard rather than forbidding cognitive work, which is a distinction teams lose in both directions. It asks for more evidence, stronger evaluation, narrower authority, and a person owning the consequential link while software does the analysis around it. A team reading high consequence as a prohibition builds nothing, and a team reading it as a detail builds something a person is nominally responsible for and cannot inspect.
Low evaluability fails differently, and it is the reading worth being frightened of. Every other factor is a claim you would want to check, and evaluability is what checking requires, so a link with poor evaluability can produce a year of confident answers without accumulating any evidence about whether they were good. Teams usually discover the gap in a postmortem, when "how would we have known this was wrong" produces silence instead of a metric.
Introduction to Thoughtware · Ch. 3No single dimension determines delegation. Judgment terrain makes the relevant conditions visible rather than producing an automatic answer from a score.
Terrain belongs to a link, not to a product
Because the readings depend on the surroundings of one particular decision, they change from link to link inside a single feature, and the Meal Companion crosses the full range in an ordinary week.
Busy evening fit reads favourably on all six at once. Dozens of comparable Tuesdays supply density, endorsed knowledge covers what busy means for this household, a typical week carries little novelty, a slightly heavy meal costs little, the plan stays reversible until somebody starts cooking, and afterwards there is a plain verdict about whether the evening worked. AssessMealPracticality can own that judgment inside declared limits, and the reason is the ground rather than anything about the model behind it.
Interpreting an unfamiliar medical diet inverts every one of those readings in the same week. A visitor arrives on a renal restriction the household has never planned around, so precedent is sparse and endorsed coverage is thin, the case is structurally novel rather than novel in phrasing, a misreading is serious, the wrong meal cannot be unserved, and nobody in the household can grade the answer without professional input. A model will produce fluent text about renal diets, and fluency was never the constraint that was binding.
Which is why the question "is our product ready for AI" has no available answer. Readiness moves link by link as knowledge is endorsed and patterns accumulate, and two links can sit at opposite ends of the six while sharing a screen. Allergy validation and meal suggestion appear in the same view and obey different leadership rules, so a roadmap reporting one autonomy phase for the whole product has already thrown away the information a stakeholder needed.
What a single adjective hides
The habit that defeats all of this is compression into one word. Teams call a decision easy, hard, risky, or fuzzy, and each of those words collapses at least two of the six readings into a judgment nobody can inspect. The compression looks like shorthand and behaves like concealment, because the readings it merged usually call for different work.
"Novel" is the clearest case. If it means sparse patterns, the response is to gather cases, or to accept that leadership stays with a person until enough of them have gone past. If it means familiar wording over a structural change, the response is detection, since something in the pipeline has to notice that this Thanksgiving departs from the last one before an established strategy fires. Those are different pieces of engineering, and the word picks neither.
"High stakes" does the same damage to the cost pair, which is why it produces guards that protect the wrong thing. A suggestion the household can ignore and an email the system has already sent both attract the phrase, and only one of them needs an approval gate in front of it, because only one of them cannot be taken back. The same defect runs through "we have a lot of data", since volume speaks to coverage while saying nothing about density or quality. Writing the factors as separate lines costs a minute, and it turns the argument in the room from whether the team is optimistic or cautious into which condition each person was actually looking at.
When the ground moves
None of these readings is permanent, and the change that invalidates one is rarely the change teams watch for. A feature that suggested meals becomes a feature that emails the shopping order to a store. Pattern density is untouched, the model is the same model, and consequence and reversibility have both moved, which means the reading taken at design time is now wrong about the two factors that were holding the design up.
So the trigger for re-reading terrain is a change in what the system does with its answer rather than a change in how the answer is produced. New side effects, a new class of person bearing the cost, and a new integration that makes an output durable are each reason enough to run the six again, and none of them arrives labelled as a model change.
Retrospectives are where that re-reading should happen and frequently does not. When a failure was foreseeable from the consequence and reversibility of the link that produced it, the useful output of the review is a changed guard, a narrowed authority boundary, or a reassigned leadership mode. Reaching for retraining instead treats a terrain problem as a capability problem, which is the same mistake the meeting at the top of this page was making, arriving a quarter later with an incident attached.
What to do next
The exercise worth running starts from a sentence somebody in your organisation has already said about a decision. "This is too risky for AI" and "the model handles this fine" are both terrain claims disguised as verdicts, and the useful question is which of the six the speaker had in mind. The answer is almost always one factor, usually consequence in the first case and pattern density in the second, while everyone else in the room heard a claim about all six at once.
That is the practical value of naming the ground. Once a speaker has to say consequence, the disagreement becomes a design question with a known repair, and it stops being a referendum on whether the team believes in AI. Established, mixed, and frontier regions takes these readings one step further, into the three combinations that recur often enough to carry a leadership policy, and who should lead each judgment turns those into assignments a product can be built against.
Read next: Established, mixed, and frontier regions.