Measuring the cost of a verified AI outcome

An AI workflow becomes cheaper per request after a model change. The team celebrates, but users spend longer correcting its outputs, more attempts are retried, and fewer tasks finish successfully. The request price improved. The economics of delivering usable work may have worsened. Those observations can coexist because they measure different things.

I prefer to start an evaluation with a named outcome and an explicit accounting boundary. What work is the system supposed to deliver? Which costs belong to delivering it? What independent evidence makes an outcome count as successful? Without those definitions, a lower cost can reflect a narrower denominator, omitted review, or an easier workload rather than a better system.

Ethen Research Lab's cost-per-verified-outcome note is a company-authored research synthesis, not a report of new measured Ethen savings. This article uses that topic as a starting point for an original synthetic worksheet. The numbers below are invented to expose accounting and interpretation choices; they do not describe an actual product, provider bill, or experiment I conducted.

Define a task that has a checkable end state

Suppose an assistant prepares structured incident summaries from a set of internal reports. A usable outcome must identify the incident, preserve the report's critical facts, distinguish missing information, and produce a record accepted by the receiving system. Generating a fluent paragraph is an attempt, not necessarily an accepted outcome.

For this illustration, the verification rule combines deterministic checks on record structure with a separate review of factual consistency. The output is successful only when both pass. Rejected, incomplete, and pending records retain their own states. A reviewer cannot improve the success rate by making difficult records disappear from the cohort after seeing the results.

Task identity also needs to survive retries. If an incident summary fails twice and succeeds on a third attempt, there is one successful task with three attempts. Counting the third attempt as a success while discarding the first two costs understates delivery cost. Counting each successful intermediate response as a separate task overstates the useful work delivered.

The verification rule should be written before comparing configurations. If a model change produces a new output style, the team may revise its acceptance rule for good reasons. It should not quietly grade one configuration against a stricter standard and another against a looser one, then report the difference as model quality. Versioning the task and verifier makes that change visible.

Set the cohort and observation window

The worksheet covers one thousand incident-summary tasks admitted under the same declared task contract. It follows their costs and terminal states through a chosen observation window. This boundary determines which work appears in both the numerator and the outcome count.

A fixed admission cohort is easier to interpret than a mixture of work started, completed, and billed during different periods. A task can begin near the end of a reporting month and finish in the next one. If its inference bill appears in the first month's costs but its verified result appears in the second month's denominator, each month's ratio becomes difficult to compare without reconciliation.

Retain pending work explicitly at the reporting cutoff. It has already consumed resources, but its final outcome is not yet known. A report can show provisional cohort cost, current verified successes, and the number still pending. Later, it can publish a reconciled view without pretending the earlier estimate was a final delivery measure.

Workload composition belongs beside the cohort definition. Simple incidents, contradictory reports, and missing attachments can require different resources. If one configuration handles a more difficult mix, an aggregate ratio may reflect that assignment. A fair comparison can use matched tasks or report clearly defined strata. It should not imply that a better overall ratio automatically means better performance on every kind of incident.

Include the work required to deliver and verify

For the synthetic baseline, initial inference costs six hundred dollars, retries cost one hundred and twenty, tools and storage cost eighty, verification costs one hundred and fifty, and allocated human review costs two hundred and fifty. The eligible delivery total is twelve hundred dollars. These are illustrative amounts, not current provider prices.

Eight hundred tasks finish with independently verified successes. One hundred and twenty are rejected, fifty remain pending, and thirty are cancelled. The provisional cost per verified success is twelve hundred divided by eight hundred, or one dollar and fifty cents. The rejected and cancelled attempts remain in the accounting boundary because they consumed resources while the cohort was being served.

This worksheet includes human review to make a substitution visible: a system can transfer effort from automated processing to people while appearing cheaper on its inference bill. Allocating human time requires a stated rule. The rate, included roles, and treatment of shared review work should be recorded so another reader can understand what the number means.

Not every expense must appear in every ratio. Research development, fixed infrastructure, and long-term maintenance might be reported separately from marginal delivery cost. The report should name the chosen boundary rather than claim to represent all economics. A transparent narrower measure is more useful than a supposedly comprehensive figure assembled from inconsistent allocations.

Compare a cheaper request path with a more expensive delivery

Now consider a second synthetic configuration serving the same one thousand tasks. Initial inference costs three hundred dollars, retries cost two hundred, tools and storage cost eighty, verification costs one hundred and eighty, and human review costs five hundred and forty. Total eligible cost is thirteen hundred dollars. Six hundred and fifty tasks finish successfully.

The request path is cheaper at the first attempt, but delivery costs two dollars per verified success. Against the baseline's one dollar and fifty cents, that is a worse result under this worksheet's definition. The example does not prove that a cheaper model generally increases review. It demonstrates why an evaluation must account for the work after the first model response.

The second configuration might still have a defensible use. It could serve a low-risk task category where a different contract allows simpler verification. It could be valuable for generating candidates that a later system evaluates. Those are separate workflows, and their economics need separate outcome definitions. Reclassifying the incident-summary workflow after observing its result would weaken the original comparison.

An arithmetic check reproduces the two totals and ratios from the specified synthetic inputs. It checks that the reported numbers follow the worksheet. It cannot validate the chosen allocations, the reliability of a real verifier, or the relationship between cost and quality in deployed traffic. Numerical correctness is necessary for a defensible result, but it is only one part of the evidence.

The complete synthetic cost worksheet

All amounts below are invented US-dollar accounting amounts. Each column covers the same fixed cohort of 1,000 tasks at a reporting cutoff. No invoices, provider prices or measured task traces are represented. Human review is an allocated cost; the illustration does not supply real hours or rates.

Table 1: Synthetic worked example
Included cost Baseline (USD) Alternative (USD)
Initial inference 600 300
Retries 120 200
Tools and storage 80 80
Verification 150 180
Allocated human review 250 540
Total included cost 1,200 1,300
Table 2: Synthetic worked example
Outcome at cutoff Baseline tasks Alternative tasks
Verified successes 800 650
Rejected 120 Not separately assigned
Pending 50 Not separately assigned
Cancelled 30 Not separately assigned
Other outcomes combined 200 350
Fixed cohort 1,000 1,000

The combined row summarizes the non-success categories; it is not an additional set of tasks. The alternative fixture specifies only verified successes and the remaining combined count. It does not establish how its 350 other outcomes split into rejected, pending and cancelled tasks. That missing breakdown limits comparison of unresolved work, but it does not prevent reproducing the stated provisional ratios.

Cost per verified success = all included cohort cost ÷ current verified-success count. Baseline: 1,200 ÷ 800 = USD 1.50. Alternative: 1,300 ÷ 650 = USD 2.00. These are provisional cutoff measures, not final costs of completing every task. First-attempt inference falls by 50%, yet the provisional cost per success rises by approximately 33.3%.

Make verification error part of the interpretation

A verifier can approve an incorrect summary or reject a correct one. If erroneous approvals enter the success count, the ratio becomes artificially attractive. If correct outputs are rejected, the denominator shrinks and delivery may appear more expensive than it actually is. The direction of error matters, not merely the fact that verification exists.

Distinguish the operational acceptance rule from an independently reviewed audit sample. The operational verifier determines how the product advances a task. The audit examines whether those decisions agree with a sufficiently careful reference judgment for the evaluation's purpose. Neither should be described as infallible.

Sampling needs a documented selection rule. Reviewing only successful-looking outputs cannot establish the verifier's rejection behavior. Reviewing only obvious failures cannot establish its acceptance reliability. A useful audit can draw from accepted, rejected, and uncertain outputs while retaining the sampling weights needed to interpret the estimates. For consequential domains, the appropriate reference judgment may require qualified expertise outside the application team.

If an audit changes the estimated number of valid successes, show both the operational count and the audited interpretation rather than silently replacing one. The cost worksheet should say which count its headline ratio uses. That makes disagreement informative: it reveals whether a delivery improvement depends on the system being too lenient about what it calls successful.

Report more than one number

Cost per verified outcome is a useful summary, but it does not describe every property a user cares about. Two systems can have the same ratio while differing in completion rate, delay, severe failure frequency, accessibility, or the burden placed on individual reviewers.

For the baseline worksheet, I would report the cohort size, verified-success count, pending count, total included cost, cost categories, and provisional ratio together. I would also report how many attempts were needed, how long tasks remained unresolved, and how much review work was concentrated in difficult cases. Those fields let a reader investigate the ratio instead of merely accepting it.

An especially important boundary is zero verified successes. The ratio is undefined in that case. Reporting zero cost per success would invert the meaning of the failure; inventing a small denominator would disguise it. Report the costs, the zero success count, and the reason no ratio is meaningful.

Latency deserves its own view as well. A workflow that eventually finishes at an attractive cost can still miss the user's deadline. If a task contract includes a completion window, a correct output arriving after that window may not be a successful delivery under that contract. The evaluation should not add a deadline only to one configuration after observing its speed. Define the requirement before the comparison and retain the actual completion distribution.

Separate recurring costs from one-time corrections

During an evaluation, an engineer may repair an adapter, revise a prompt, or correct a malformed dataset. Those interventions affect both the cost and the interpretation of subsequent runs. Treating them as ordinary failed attempts can obscure a system defect; excluding them without explanation can make the trial look cleaner than it was.

Keep an intervention log with the time, changed component, affected tasks, and reason for the change. If the change materially alters the configuration, the later work belongs to a new configuration version. A report can still explain the complete development process while comparing versions on appropriately defined cohorts.

One-time migration work may reasonably sit outside a recurring delivery ratio. The decision then needs a separate estimate of how much recurring improvement would be required to justify that migration cost. The estimate should use declared assumptions about workload and review effort rather than turn the synthetic ratios in this article into an investment recommendation.

The same principle applies to shared tooling. If a verifier serves several products, allocating all its infrastructure cost to one trial may be misleading. Allocating none of it may be misleading too. The important property is a reproducible rule, plus enough detail to examine alternative allocations. A sensitivity analysis can show whether the conclusion survives plausible changes rather than presenting one allocation as the only legitimate answer.

Plan the comparison before optimizing the router

A router that sends tasks to different model configurations can change which tasks each path receives. An apparent cost improvement might reflect its assignment decisions rather than better behavior within any route. If the system also learns from feedback, the decision policy may change during the observation window.

For an initial comparison, keep task identity, configuration version, route decision, and verification result together. A paired or otherwise carefully designed comparison can make the assignment question explicit. A production observation with uncontrolled assignment should be labeled as such; it should not automatically become evidence that one model caused the improvement.

Ethen's learned-routing research protocol is explicitly a planned experiment rather than a completed result. That distinction is relevant here: a promotion criterion explains what evidence would justify a future change. It does not establish that the criterion has been met. I make no claim that Faros achieved a cost reduction, or that I implemented its routing evaluation.

The protocol for a real incident-summary trial should state which comparison would change the deployment decision, how pending work and verifier errors will be treated, and when the analysis stops. It should also preserve disappointing results. A router that saves inference cost while increasing total review effort may teach the team something useful, even when it fails the criterion for promotion.

Use uncertainty to decide what to inspect next

An evaluation can leave uncertainty about several different quantities. The invoice total might be exact while the human-time allocation is approximate. The task count might be exact while the verifier's agreement with an independent reviewer is uncertain. The reported ratio should not imply that all these inputs have the same evidential quality.

Use a short uncertainty register beside the worksheet. It names each uncertain input, the evidence currently supporting it, a plausible range where justified, and the next observation that could improve it. If changing the review-cost assumption reverses the conclusion, validating that assumption is more valuable than calculating the headline ratio to another decimal place.

Similarly, a disagreement over outcome validity should lead to targeted verifier review. More tasks do not repair a systematically mistaken acceptance rule. More precise model-price data does not repair a biased task assignment. The useful next experiment is the one that addresses the uncertainty determining the decision, not necessarily the one that generates the largest volume of new logs.

This is why I would describe the worksheet as a decision aid with a stated boundary. Its purpose is to connect resources to accepted work in a way that another engineer can inspect. Precision follows from better definitions and better evidence, rather than from a more impressive dashboard.

Test the allocation that could change the decision

The alternative's costs excluding human review total USD 760. At 650 verified successes, matching the baseline's USD 1.50 ratio permits USD 975 of total cost. That leaves USD 215 for human review: 975 minus 760. The worksheet currently allocates USD 540, so the difference is substantial. This is a break-even sensitivity calculation, not evidence that the review allocation can actually fall to USD 215.

Both configurations must use a consistent, defensible allocation rule. Changing only the unfavorable column to obtain a preferred conclusion would defeat the comparison. If review time is uncertain, collect the hours, roles and rates that would justify an allocation, and examine the baseline's uncertainty too. This aggregate fixture contains no underlying task traces, so it cannot support an audit of real invoices or individual outcomes.

Account for the whole attempt to deliver

The synthetic worksheet shows how a cheaper first request can coexist with a higher cost of verified delivery. It also shows why a ratio needs a task definition, a cohort, a cost boundary, an acceptance rule, and a clear treatment of pending work.

Keep those definitions beside the result and retain every failed attempt that belongs to the cohort. Then I would audit verification, separate empirical observations from estimates, and test how sensitive the conclusion is to the inputs that remain uncertain. No single number should stand in for that record.

The practical question is whether the system delivers the intended work under a defensible definition of success, and what resources doing so requires. That question connects model choice, product design, verification, and human effort. More original commentary on evaluation methods is collected on the site's research page and in the writing hub.

Sources

Back to blog