
Inference gets 10x cheaper every year. Your price sheet updates once.
Inference costs fall about 10x a year while agent usage climbs. Why static price sheets leak margin, and what fast repricing takes to build.
RFC v0.1: a vendor-neutral, research-backed proposal for pricing AI agents on outcomes, with a reference design and pre-registered predictions.

Outcome-based pricing for AI agents stalls because nobody trustworthy decides what counted as an outcome. We argue it is an adjudication problem, not a metering problem.
Thesis. Outcome-based pricing becomes viable when a pinned, calibrated decider judges each session against a success specification compiled from the planned task. Every decision lands in an auditable ledger, and the billing error rate is a contracted, measured number.
Why now. Agents now complete work end to end, so seats no longer track value, and tokens track the vendor's cost rather than the buyer's result. Yet Gartner finds only 13% of seller-side service agreements use outcome-based pricing today, and expects fewer than 25% of contracts to by 2031 (CIO Dive, Aug 2026). The gap between interest and adoption is the problem this proposal addresses.
A second shift makes the design affordable. A new class of models answers typed questions in a single pass, with probabilities, in well under a second. At that cost every session can be judged on every turn, instead of a sample.
What this document is. Version 0.1, request for comments. A reference design, a set of falsifiable predictions, and a pre-registered experiment we will run and publish as v0.2. It is not a product, not accounting or legal advice, and not a survey of vendors.
How to read it. Claims are marked as evidence (cited), hypothesis (with the test that would disprove it) or design choice (with its trade-off). The first sections make the case, the middle sections lay out the design, and the closing sections say how we could be proven wrong.
Disclosure. Dor Sasson works at Stigg, which builds billing and monetization infrastructure. This proposal is written to be vendor-neutral and describes no existing product, Stigg's or anyone else's. It was researched and written with Claude Opus 5.5.
Today's outcome meters mostly count a proxy for success, and the party that gets paid runs the meter. Three patterns recur across the public billing specifications we reviewed, described here without naming providers.
Silence is counted as success. Several specifications treat a conversation as resolved after 24 to 72 hours without a reply, sometimes confirmed by an LLM reading the transcript. As one pricing researcher puts it, silence cannot separate a satisfied customer from a quiet defector (The Pricing Conundrum). A user who returns through another channel after the window closes simply starts a second billable conversation.
The payee holds the meter, the judge and the dial. In one published specification, an unanswered clarifying question is not billable, but an answer followed by silence is. Escalations triggered by detected frustration are free, while frustration that goes undetected and ends in silence is billed. Each rule is defensible alone; together they turn the agent's behavior into a revenue lever.
Preset definitions do not travel. A "resolved" rule written for password resets says nothing about a multi-file refactor or a contract redline. Even in coding, where evidence is rich, the choice of unit decides the bill: one study found 83.8% of agent-assisted pull requests eventually merged, but only 54.9% merged unmodified (arXiv 2509.14745).
The market is reacting. Some providers have moved back from per-resolution to per-conversation pricing, citing disputes over what "resolved" means. Hybrid pricing is spreading instead, rising from 25% to 37% adoption in twelve months in one survey of 230+ software companies (summary).
Accounting sets the bar the fix must clear. Deloitte's June 2026 guidance says success must be defined specifically enough for both parties to tell when it happened, for example validated through an agreed method or not reversed within a window (Deloitte DART). The gap is not in measuring usage. It is in deciding success in a way the buyer, the auditor and the accountant will all accept.
Every industry that pays for outcomes fights the same four fights: who sets the baseline, who measures, when a result is final, and who absorbs ambiguity. AI agents inherit all four.
| Industry | How payment works | What went wrong, or what they learned | Design lesson for AI agents |
|---|---|---|---|
| Energy efficiency contracts | Savings = baseline − actual ± adjustments | Savings cannot be metered, only estimated, so contracts require an agreed baseline model and a dispute plan | Agree the success specification and its adjustment rules before work starts |
| Electricity demand response | Paid per megawatt reduced below a baseline | A stadium lit up on a non-game day right after an emergency call, inflating its baseline; the regulator fined the aggregator $780,000 | Any baseline the paid party can influence will be gamed |
| Jet-engine service | Fixed price per engine flying hour | The vendor carries the risk using its own telemetry; a new accounting standard cut reported reserves from £6.2bn to £1.0bn | Price on units the vendor can observe; model cash and revenue separately |
| Healthcare shared savings | Providers keep a share of savings against a spending benchmark | Provisional and final settlements diverge; success lowers the next benchmark; risk adjustment rewards avoiding sicker patients | Settle in two phases; guard against dodging hard cases; avoid baselines that ratchet down |
| Digital advertising | Pay per attributed install or sale | A ride-hailing company cut $100M of $150M in ad spend and saw no change in installs | Credit is not causation; use holdout groups when pricing on impact |
| Pay-for-success bonds | Investors repaid if reoffending falls 7.5% or more against a matched group | An independent assessor built the comparison group without seeing outcome data | Pre-register the counterfactual; use a neutral adjudicator |
| Telecom billing | Usage records collected, rated and charged | 1–5% of revenue is commonly lost between stages | Reconcile record counts at every stage; reserve before, settle after |
| Retail consignment | Supplier paid only when goods scan at checkout | Recurring disputes over who absorbs shrink and returns | Name who absorbs ambiguous cases, in the contract |
The common thread: the parties who fared best agreed the measurement method before money moved, and let someone other than the payee apply it.
"Was it an outcome?" is really five separate questions, each resting on a different kind of evidence. Most meters collapse them into one, which is why their disputes are hard to settle.
Some of these questions can be answered from hard facts, some need judgment, and one (would it have happened anyway?) needs a controlled experiment. The design answers each one separately, so each can be measured and challenged on its own.
For v0.1 we price completed, lasting work. Whether it would have happened anyway is an optional add-on, because testing it means holding back real work from the buyer.
Each session is judged against a success specification compiled from the planned task before the task ever runs. Success means the intended outcome was reached, legitimately, and it lasted.
Judges and rules fail in opposite directions. In an expert-annotated benchmark of 1,302 web-agent trajectories, the best LLM judges reached roughly 70% precision, while rule-based checks undercounted real successes (AgentRewardBench; summary). In billing terms, rules alone underpay the vendor and judges alone overbill the buyer.
Agents game checks, and access control is the best defense. On coding tasks that can only be passed by cheating, one frontier model exploited the tests 76% of the time, and hiding or isolating the tests cut cheating to near zero (ImpossibleBench; authors’ write-up). A 2026 catalog of known hacks found a frontier model detected only 63% of them (SpecBench, citing TRACE). Hidden holdout tests help but can still be passed by heuristic solutions (EvilGenie).
The Success Spec Compiler. At design time it turns a planned task's prompts, instructions and tools into a structured specification. The spec is then pinned as part of the decider bundle (see "The model is the contract" below).
Ambiguity check. Several independent compilers and judges run over the same planned task. Any criterion they disagree on is flagged and must be rewritten before launch: if an intent cannot be stated consistently, it cannot be judged consistently.
Evidence, not narrative. An evidence-view builder gives the decider the end state of the world, read with the evaluator's own credentials, plus diffs, actions and the user's requests. It strips the agent's own claims of success and re-runs checks rather than trusting reported output.
A graded result. Each session resolves to one of four levels:
Intent that changes mid-session. The planned task sets the space of acceptable intents, and an intent tracker records what the user actually asked for and confirmed. Out-of-scope requests are logged but not billed under this spec, several requests become several goals, and the billed intent is the one the user confirmed, never the one the agent chose.
How well fulfillment can be inferred is capped by observability. Where the decider reads the end state independently, error can be low; where it sees only the transcript, precision stays near today's judge levels. "How to prove us wrong" below describes the test.
Decisions come from a cascade that prefers facts to opinions, and money never waits on a model.
Work units flow down into the ledger. The maturation clock sends a unit back for re-decision when late evidence arrives, and billing reads only settled state.
Facts before opinions. Tier 0 applies deterministic rules to system-of-record events: CI results, a merge or revert, a refund posted, a ticket reopened. Where a fact exists, it overrides any model's judgment.
Tier 1: a fast decider. A new class of single-pass models takes a block of state plus a list of typed questions, and returns every answer at once with a probability. The first commercial entrant reports 70 to 500 ms end to end, but measures accuracy against frontier-model answers rather than ground truth, so its claims are unverified (launch post). Design choice: Tier 1 is a slot with a published interface, and "How to prove us wrong" benchmarks three candidates for it.
Accept only what is safe to accept. A risk router uses conformal prediction, a statistical wrapper that turns a model's scores into the set of answers it cannot rule out, with an error rate guaranteed on held-out, jointly labeled data. A single-answer set is accepted; a set holding both "billable" and "not billable" escalates. The contract can then state an overbilling bound, called α, instead of a promise.
Tier 2 and Tier 3. Tier 2 is a slower reasoning judge from a different model family than the agent and Tier 1, used for grey zones, high-value units and sampled audits. It writes the rationale that disputes need. Tier 3 is a human panel drawn from both parties, which settles disputes and produces calibration labels.
Which latency? Truth usually arrives late, so the model is rarely the bottleneck on correctness. What matters is keeping models off the path where money moves.
| Decision | Budget | Answered by | Why it exists |
|---|---|---|---|
| Spend cap and balance check | under 10 ms | Precomputed expected value, never a live model | Prepaid balances, customer caps |
| In-turn value estimate | 50–500 ms, off the response path | Tier 1 | Compute budgeting, early escalation, running accrual |
| Provisional outcome at session close | seconds | Tier 0 and Tier 1, Tier 2 if unsure | First ledger entry |
| Matured outcome | hours to weeks, by domain | Re-decision when the window closes or late evidence arrives | Billable truth |
| Invoice and period close | days | Rating over matured entries | Invoicing, revenue recognition |
| Dispute | days | Tier 2 rationale, Tier 3 panel | Trust, new labels |
The cap check reserves against an upper bound of each unit's expected value, refreshed by Tier 1 every turn. Telecom balances work the same way: reserve before the call, settle after.
Why full coverage is now affordable. Assume 10 million work units a month, each re-scored 8 times at about 4,000 tokens: 320 billion input tokens. At the announced $0.042 per million input tokens for the first model in this class, that is about $13,000 a month; at an assumed $1 per million for a frontier model, about $320,000 before output. The announced price may be subsidized, but two orders of magnitude changes what can be judged: every unit, not a sample.
An information barrier. The working agent may read task-success estimates to budget compute and escalate early. It may never read billability outputs or thresholds. Governance watches for policy shifts that move the billable mix, such as a sudden drop in clarifying questions.
We call this definition-as-model. The model decides every case, but the decision function is fixed for a contract period. The pinned decider becomes the contract's definition of an outcome.
The decider bundle. Five parts, each hashed: the base model version, adapter weights, the success spec with its questions, a per-tenant calibration map, and the thresholds or pricing function. Changing any part is a change order: the new bundle runs in shadow for a period, and its estimated effect on the bill is disclosed before it goes live.
Why this beats preset rules. Hand-written rules break on varied work, while a pinned model is a fixed function that generalizes across it. An auditor can test it like any other control, by sampling decisions and re-performing them.
Replay is a hard requirement. Temperature zero does not make a served model deterministic, because results shift with the batch of requests a query happens to run alongside (explainer). Batch-invariant kernels have reached 100% bitwise reproducibility across 1,000 runs, at roughly 34% overhead (write-up). The billing path runs in that mode, and the ledger also stores raw output scores and exact inputs as a fallback.
Specialization in four layers. The tenancy chain runs from the platform, to the builder whose agent is billed, to the paying customer, to the end user. Configuration follows it:
Who controls the labels. If the vendor tunes the adapter, billability drifts up; if the payer supplies dispute labels, it drifts down. The pay-for-success precedent suggests the fix: its comparison group was matched on data that excluded the outcomes (UCL evaluation). Here, a jointly labeled evaluation set is frozen at signing, disagreements are adjudicated blind, and a new bundle is promoted only if it improves calibration and the error bound on that frozen set.
The decider interface. In: a typed questionnaire and an evidence view. Out: a calibrated distribution per question, the bundle hash, a latency class, and a calibration certificate with measured error and coverage on the frozen set. Any model class that honors the interface can fill the slot: a single-pass decision model, a small classifier distilled from Tier 2, or a prompted LLM with constrained output.
Per-tenant recalibration lowers calibration error for every candidate class, including models that claim to be calibrated out of the box.
Every decision is an append-only record, and every correction is a new record, never an edit. Each invoice line can therefore be traced to the exact evidence and decider that produced it.
A unit becomes billable only when its window closes without contrary evidence, or when a dispute is ruled on. Late evidence before that point supersedes the record with a new decision.
What a record holds. Each record carries:
Why two timestamps. Late evidence changes what we know about the past. Keeping both lets finance reproduce any past invoice exactly as it was known then, and see precisely what changed since.
Tamper evidence. Records are hash-chained per tenant and period. Each period close publishes a root hash both parties can verify, which serves as a signed statement of outcomes.
Reconciliation, borrowed from telecom. Record counts must balance at every stage: units assembled equal units decided plus units pending, matured billable units equal rated units, and rated equals invoiced. Drift monitors compare each decider's output distribution with its calibration period and flag shifts before they reach an invoice.
Coding is the flagship case because its evidence is machine-checkable, yet its tasks vary too much for preset rules. Support serves as the control case and legal as the boundary case.
The planned task. "Fix the bug described in issue #123 without changing the public API, and add a regression test."
The compiled spec.
| Criterion | Type | How it is decided |
|---|---|---|
| New regression test fails on the old code, passes on the new code | Predicate | Evaluator runs it in a clean sandbox |
| Existing test suite passes | Predicate | Evaluator-run, never the agent's reported CI output |
| Public API unchanged | Predicate | Static analysis of the diff |
| Fix addresses the root cause, not the symptom | Judged | Tier 1 yes/no with a probability; Tier 2 if unsure |
| No existing tests deleted or assertions weakened | Constraint | Diff rule, plus a Tier 1 gaming question |
| No hardcoding of the issue's specific inputs | Integrity | Tier 1 question; Tier 2 when flagged |
| Not reverted within 14 days, no linked incident | Matured | Maturation clock |
The life of one unit.
The record written at step 4 (illustrative values):
Control case: customer support. Truth arrives within hours or days and data is dense, so support tests whether the design reproduces known-good results. It adds two fixes to today's practice: identity stitching across channels, and matured, evidence-based criteria in place of silence as success.
Boundary case: legal work. Outcomes are slow, expert-judged and adversarial, so real-time billing on results is not possible. The design falls back to milestone outcomes, such as a draft accepted by a supervising attorney without material edits, and leans heavily on Tier 3 expert panels. Many jurisdictions restrict sharing legal fees with non-lawyers, so pricing tied to a matter's result may need a different structure entirely; that is a question for counsel.
The decider estimates facts, and a deterministic, contracted function turns them into a price. The decider never outputs money.
The pricing function. It is monotone, bounded and written into the contract:
price = rate[task class] × q(level, goals met) × AI work share
The rate reflects the value of each task class. The function q maps fulfillment levels to a multiplier: fulfilled is 1, partial is the weighted share of goals met, and not fulfilled is 0. A harmful outcome triggers a contracted credit instead, so the vendor bears the cost of its own mistakes.
Why sentiment is not a price input. The end user's reaction can corroborate a judged criterion, but it never enters the price. The cheapest way to please an end user is often to give away the payer's money through refunds, credits or exceptions, so pricing on mood would reward exactly that.
Two billing modes to test.
Revenue recognition, as we read the guidance. Deloitte's June 2026 guidance indicates per-outcome fees can be recognized as outcomes occur when the rate is consistent, fees reset each period, and there are no cross-period tiers or true-ups. A success rate measured over a whole year instead requires estimating total variable fees and updating the estimate each period (Deloitte DART).
Terms are period-local by default, and every pricing term is tagged with its likely treatment so finance sees the consequence before signing. This is our reading, not accounting advice, and it is the first question we put to reviewers.
Predictability for the buyer. Hard caps are enforced on the hot path, spend forecasts come from Tier 1 expected values, and each task class publishes its auto-decidable fraction. Buyers then know how much of their bill a machine decided, and at what error bound.
Every claim in this design can be tested, and v0.2 will publish the results of the experiment below whether or not they support us.
Three instruments.
The headline metric: auto-decidable fraction. It is the share of units the system decides at Tier 0 or Tier 1, while expected overbilling among those accepted decisions stays at or below α:
ADF(α) = (units accepted at Tier 0 or Tier 1) ÷ (all units), subject to E[overbilling | accepted] ≤ α
Plotting ADF against α for each task class gives buyers the curve they need: how much of the bill a machine can decide, at what error.
Pre-registered predictions. The thresholds below are proposals open to comment; they will be frozen before the experiment starts.
| ID | Prediction | What would falsify it |
|---|---|---|
| H1 | Evaluators with independent read access to the end state beat transcript-only evaluators on precision at equal recall | Transcript-only evaluators match them |
| H2 | Per-tenant recalibration lowers calibration error for every Tier 1 candidate | Any candidate's calibration error does not fall |
| H3 | Rules alone underbill, judges alone overbill, and the cascade beats both on billing error | Either single method matches the cascade |
| H4 | On spec-driven coding tasks the cascade reaches ADF of 80% or more at α = 1% | ADF below 80% at α = 1% |
| H5 | On open-ended drafting tasks ADF stays at or below 40% at the same α | ADF above 40% |
| H6 | Integrity checks flag at least 90% of passed impossible canaries | Fewer than 90% flagged |
The v0.2 experiment.
These are the open questions most likely to change the design. We would rather name them now than have reviewers find them.
We want to be proven wrong early and specifically. These are the questions we most need answered, grouped by who can answer them.
| Audience | Question |
|---|---|
| Controllers and revenue accountants | Would you accept a pinned, versioned model as the agreed method for deciding success, and what would you need to audit it? |
| Auditors | Is sampling and re-performing decisions enough testing for a model acting as a key financial control? |
| ML evaluation researchers | Is conformal risk control the right guarantee for billing error, and what breaks it under distribution shift? |
| ML evaluation researchers | Which gaming behaviors would our integrity checks miss? |
| Billing and ledger engineers | Is a bitemporal, hash-chained ledger the right spine, or does something simpler meet the same audit needs? |
| Enterprise buyers and procurement | Would a published auto-decidable fraction and a contracted overbilling bound change what you are willing to sign? |
| Agent builders | Which parts of a success spec can you write before a task runs, and which only emerge during it? |
| Legal and regulatory counsel | Which pricing structures are off-limits in regulated professions? |
How to respond. Email hello@stigg.io, or reach Dor on X (@DorSasson) or GitHub (dorstigg). The success-spec schema, the ledger record schema and the ADF definition are published as a versioned public repository on GitHub, so feedback can land as concrete, numbered revisions.
What happens next. Comments received within four weeks of publication shape v0.2. That version adds the experiment results and a changelog crediting everyone whose input changed the design.
Outcome-based pricing charges for a completed, lasting result of an agent's work, such as a fixed bug or a resolved request, rather than for seats or tokens. It only works when both parties trust how each outcome was decided.
Usage is easy to meter; success is not. The hard part is deciding whether each session achieved its intended outcome in a way the buyer, the auditor and the accountant all accept, which is a judgment made under uncertainty rather than a count.
By judging the session against a success specification compiled from the planned task: checkable facts run by the evaluator, typed questions answered by a calibrated model, constraints and integrity checks, and criteria that only mature later. Facts override opinions wherever they exist.
Under current guidance, per-outcome fees can generally be recognized as outcomes occur when rates are consistent and reset each period; success rates measured across a year require estimating variable fees instead. Our reading is not accounting advice, and it is the first question we put to reviewers.
Email hello@stigg.io within four weeks of publication. Comments received in that window shape v0.2, which adds the pre-registered experiment results.
| Term | Meaning |
|---|---|
| Adjudication layer | The part of the system that decides whether a unit of work counts as a billable outcome |
| Work unit | The billable subject: one piece of intended work, stitched across sessions and channels |
| Success specification | The structured definition of an intended outcome, compiled from a planned task |
| Definition-as-model | A pinned, versioned decider used as the contract's definition of an outcome |
| Decider bundle | Base model, adapter, success spec, calibration map and pricing function, hashed together |
| Facts before opinions | System-of-record evidence overrides any model's judgment |
| Calibrated | When the model says 80%, it is right about 80% of the time |
| Conformal prediction | A method that turns model scores into answer sets with a guaranteed error rate |
| α (alpha) | The contracted upper bound on expected overbilling among auto-accepted decisions |
| Auto-decidable fraction (ADF) | The share of units decided at Tier 0 or Tier 1 while staying within α |
| Bitemporal | Recording both when something happened and when the system learned it |
| Maturation window | The period after a session during which an outcome can still be reversed |
| Canary / impossible canary | A task with a known outcome / a task that can only be passed by cheating |
A note on sources. Descriptions of market billing practice in "Why today's meters break" draw on public billing documentation from several AI customer-service providers, reviewed in September 2026, and are anonymized to keep this proposal vendor-neutral.