Economics of judgment · Spend where misses hurt
Paying for reliability
Reliability is a line item with a price. Spend on wrappers, graders, and guards proportional to consequence, not to maximise cheap fluency.
11 min read
Cover for Paying for reliabilityWrappers turn reliability into something with a price. The question that follows is whether the price is worth paying, and how anybody would know. Teams that treat reliability as taste buy wrappers they cannot defend. Teams that treat it as free skip wrappers they should have bought on consequential paths. Both errors have the same root: reliability treated as an adjective rather than a budget line with measurable return.
Helpful context: Reliability means measured behaviour over time, not a confidence score from one call. Two budgets frame money and latency. Evaluation as engineering supplies the suite movement wrappers must produce. Tuning by wrapping is how spend attaches without rewriting identity.
Tiered spend by consequence
Not every judgment deserves the same wrapper stack. Risk-tiered spending matches spend to consequence and to eval coverage on each class. A variety judgment on a familiar week where the household has accepted similar plans for months carries low consequence. A first busy-week plan of the season where timing matters carries medium consequence. A medical deferral or unfamiliar dietary restriction carries high consequence because the failure cost extends beyond the product into the household's safety.
| Path | Wrapper policy | Meal Companion example |
|---|---|---|
| Low consequence | Single run, mid-tier model | JudgeWeeklyVariety on a familiar week |
| Medium consequence | Sample or critic | AssessMealPracticality on first plan version |
| High consequence | Double-run judge, human bridge | Medical deferral, unfamiliar dietary restriction |
Double-run JudgeHouseholdFit when a plan touches endorsed allergy knowledge. Single-run variety judgment when the household has accepted similar plans for months. The spend is visible on the cost per decision spreadsheet. The tiering is arguable in leadership review because the consequence categories are named and the wrapper policies are documented.
- 01Wrapper spendSampling, critics, cascades
- 02Suite movementMeasured accuracy on the edge
- 03Autonomy returnFewer human exceptions
- 04CeilingDeterminacy limit for the class
Publish spend and effect on the same page.
The ceiling still holds
Open decisions have a determinacy ceiling: no cognitive unit exceeds the human agreement rate for its class. Money cannot buy past it. Spending up to the ceiling where consequence justifies it is rational investment. Spending beyond it is reliability theatre, where the bill rises and accuracy flatlines because the task itself admits disagreement that no amount of sampling resolves.
See determinacy ceiling. What wrappers buy is often autonomy: fewer cases reaching a clerk or household approval step, not a prettier template. A small monthly spend that removes a thousand human exceptions can outweigh token shaving on a rare prestige edge, because the clerk time and household wait time carry cost the provider bill never shows.
The cognitive unit, Ch. 14Reliability has stopped being a matter of professional judgement exercised privately and become a line item.
Eval spend is reliability spend
Exact checks, human-labeled calibration sets, and regression suites cost money and time. That cost is part of paying for reliability, not overhead to cut first. Skipping eval saves money until failures externalize to support, incidents, and trust erosion. Evaluation replaces certainty names the operating norm this spend supports.
Grader calibration on busy weeks is how operators detect drift before households do. The calibration cost appears as reliability spend the same way wrapper cost does. A team that refuses eval spend on a consequential path cannot know whether any wrapper spend on that path works. The two investments are coupled: wrappers without measurement are theatre, and measurement without action is documentation.
What this looks like in the household planner
Return to a Meal Companion feature costed by named cognitive units. Sampling the consequential ones three times with consensus raises the cognition bill. The rise is visible to finance without requiring them to understand templates.
What the spend buys must be measured concretely: points of accuracy on the suite that convert into autonomy (fewer approval steps the household must complete), avoided failure cost when allergy check would have failed open (trust protection), and reduced loop iterations when a critic wrapper prevents shallow compose passes (time budget savings). Publishing spend, suite movement, and autonomy or failure effect on the same page prevents private taste from driving investment. "Making it more reliable" without numbers is how teams buy theatre.
The tiering plays out across the full range: variety on a familiar week gets a single run with a mid-tier model and light suite. First busy-week plan of the season gets sample or critic on AssessMealPracticality with full busy-week regression. Endorsed allergy touched gets double-run JudgeHouseholdFit and the exact allergy guard must pass regardless. Medical diet deferral gets a human bridge with no autonomy expansion until coverage exists. The tier table is policy, not improvisation in incident response.
Guards as cheap reliability
Deterministic guards before cognition often deliver the highest reliability return per dollar: exact allergy enforcement, schema validation, permission denial. They belong in the reliability budget alongside wrappers because they are reliability investments with near-zero token cost. Teams that triple-sample a variety judge while leaving allergy inside a cognitive unit pay twice: for wrappers on a low-consequence edge and for variance on a high-consequence predicate that a guard would eliminate entirely.
The distinction becomes visible in incident postmortems. An allergy violation traces back to closed work left in cognition. A variety complaint traces back to a judgment call on open terrain. The first failure was avoidable with a guard. The second may warrant wrapper spend. Treating both as "reliability problems" and applying the same solution is how budgets grow without outcomes improving.
Reliability without coverage is theatre
Spending on wrappers for a decision class with no suite coverage is theatre. Without a suite, you cannot know whether double-run JudgeHouseholdFit moved accuracy. The wrapper exists for comfort rather than for measured outcome. Adding cases before adding wrappers, or adding wrappers only as part of suite expansion with before-and-after measurement, keeps the investment honest.
Coverage and spend should rise together on consequential classes. Low-consequence classes may ship with lighter instrumentation until volume justifies investment. This asymmetry is deliberate: consequence determines where the organisation allocates reliability budget first, and volume determines where it allocates second.
Finance-friendly reliability review
Once per quarter, publishing a reliability spend table makes the investment arguable outside engineering: decision class, wrapper policy, monthly spend on that policy, suite accuracy or autonomy delta since last quarter, note on determinacy ceiling. Finance can approve or challenge spend without reading templates. Engineering can defend or retire wrappers with evidence instead of adjectives.
Including guard spend (near zero tokens) beside wrapper spend (visible tokens) ensures cheap exact work gets credit for reliability return alongside expensive sampling. Including incident cost (illustrative failure cost beside wrapper cost) ensures consequential edges get defended with full arithmetic rather than with the assumption that low token spend means low importance.
Under-spend is also a choice
Teams that refuse all wrapper spend on consequential paths externalize cost to incidents and manual review. Teams that refuse eval spend cannot know whether any spend works. Paying for reliability is the middle path: measured spend on guards, suites, and wrappers proportional to consequence, stopped at the determinacy ceiling, published in numbers finance and product can read.
Autonomy return accounting makes this visible. Fewer clerk exceptions, fewer household approval steps, and fewer support tickets on covered classes are the return on reliability spend. Double-run JudgeHouseholdFit when endorsed allergy knowledge is touched costs roughly $0.045 per plan. If that spend removes one human review per ten plans at fifteen minutes each, the wrapper pays back in clerk time even when token math looks expensive. Documenting both sides in quarterly reliability review means finance approves spend with evidence, not adjectives.
Long-horizon trust
Reliability assembles over time from sufficiency, stability, determinacy, coverage, and accuracy. Paying for reliability today buys lower variance tomorrow on covered classes. Households learn what the system will never do when allergy guard never fails open. That learning compounds into trust that the product can draw on when introducing new features or expanding autonomy.
Cheap fluency without guards teaches the opposite lesson: the system sounds confident but occasionally violates commitments. One missed allergy after months of correct plans can collapse trust faster than a year of creative variety delighted. The economic cost of trust erosion appears in churn, support, and manual override rate even when token spend looks stable. Reliability spend is therefore a retention investment alongside an inference optimisation.
Tier review cadence
Reviewing tier tables each quarter against incident history keeps spend aligned with consequence rather than with last year's assumptions. A path that never triggered failure may drop wrapper tier. A path that triggered trust incidents may rise tier even when mean grader score looks stable. The cadence prevents both over-spend on quiet paths and under-spend on paths that changed character since the last review.
Updating the table when authority rises or suites expand keeps the tiering current. A cognitive unit that gained autonomy since last quarter may warrant reduced human bridge spend if eval supports the expansion. A cognitive unit that showed regression on busy-week cases may warrant increased wrapper spend even when monthly totals look flat.
What to do next
Tiering decision classes by consequence and documenting wrapper policy per tier is the entry point. Adding one consequential edge to double-run or critic wrap and measuring before and after on the existing suite makes the economics concrete. Moving one exact check from cognitive unit to guard when eval shows closed work left open delivers cheap reliability return. Reviewing monthly whether reliability spend moved suite outcomes or only the bill prevents theatre from accumulating.
See tuning by wrapping, cost per decision, and decision cost falls with history.
Read next: Decision cost falls with history.