%20(1).png)
Software Billing Models: 8 Types and How to Choose
Software billing models decide how you charge and what you must build. Compare subscription, usage, credits, and hybrid models, and how to choose.
Follow the metering and billing path from usage event to invoice, including ingestion, aggregation, rating, reconciliation, credits, and enforcement.
%20(1).png)
A customer opens a billing ticket with a simple question. “Why did we get charged for this twice?” Engineering traces event IDs, finance checks the invoice, and nobody knows where the bad number entered the system.
Metering and billing connect product usage to what the customer pays. The interesting part lives between those endpoints, across ingestion, aggregation, rating, reconciliation, and the runtime controls that act before more AI compute runs.
Metering measures product consumption. Billing applies commercial rules to that measured usage and creates the financial record.
Metering starts inside the product. Each billable action emits an event tied to a customer, feature, quantity, timestamp, and other dimensions needed downstream.
Those events might represent:
AWS’s SaaS architecture guidance separates metering, operational metrics, and billing into distinct concerns, which is a useful separation when assigning ownership across services.
Billing picks up from the measured quantity. Rating applies the relevant price or rate card, then invoicing turns rated charges into the amount the customer owes.
The metered billing layer covers charging based on measured consumption. The broader metering and billing system also includes the event path that produces that measurement.
Metering and billing work as a pipeline that turns raw product activity into a rated charge, invoice, and reconciled financial record.
The important detail is state ownership between stages. Each layer should consume a defined output from the previous one without rebuilding its logic.
A duplicate event can pass cleanly through aggregation, rating, and invoicing. Every downstream calculation may be correct while the final bill is still wrong.
Event ingestion captures each billable action as a durable, uniquely identifiable usage event before aggregation begins.
The event envelope needs enough context to support attribution, deduplication, replay, and downstream grouping. A typical event might carry:
The distinction between event_id and business identity is important. Two valid requests may look identical while representing separate consumption. A retry of one request should still collapse to one event.
If your transport uses at-least-once delivery, duplicate delivery is part of the design space. Deduplication needs to happen before a retry can inflate the billable quantity, and replay deserves the same treatment.
If ingestion pauses, durable source events should let you rebuild the missing window without creating a second copy of usage already accepted.
Aggregation converts raw events into the exact quantity rating will price.
The aggregation rule depends on what the meter represents:
The grouping key matters as much as the function. A token meter might aggregate by:
tenant_id + model + billing_period
Add workspace, region, feature, or agent and the cardinality grows quickly. Every extra dimension gives you more attribution detail and more state to query, store, and reconcile.
Window semantics also need one owner. UTC boundaries, customer-local billing periods, late events, and backfills can all move otherwise valid usage into a different period.
Late-event policy should be explicit. You need to know whether a late event updates the closed period, creates an adjustment, or moves into the next billing cycle.
Rating turns an aggregated quantity into priced usage under the commercial terms that applied when the usage occurred. A typical path looks like this:
measured quantity → included allowance → applicable tier → contract override → overage rule → charge
The rating layer may need to resolve:
The architecture of a billing system gets harder to reason about once multiple versions of those rules can apply to the same customer over time.
Pricing configuration therefore needs versioning.
Usage generated on August 31 should resolve against the terms active on August 31, even if the customer upgrades on September 1.
A useful rating record should preserve the quantity, pricing version, rate, and resulting charge. That gives reconciliation something concrete to compare later.
Invoicing converts rated usage into the financial record the customer receives and accounting systems recognize.
By this point, the meter has already answered how much was consumed, and rating has answered what that usage costs.
The invoice layer still has several pieces of state to preserve:
The useful debugging direction runs backward.
An engineer should be able to start from an invoice line and trace it to the rated charge, aggregated quantity, and source usage events.
Without that chain, a customer dispute turns into a search across logs, billing exports, and ad hoc SQL.
Reconciliation verifies whether the usage that entered the pipeline matches the usage that became a charge. This is where the pipeline proves its own work. A useful reconciliation process should answer:
Reconciliation becomes much more useful when mismatches identify a stage.
“Usage and billing differ” gives you an alert. But “Seventeen late events entered after the August window closed” gives an engineer a starting point. That traceability is what turns metering and billing from a chain of counters into an operable production system.
Reliable metering and billing depends on stable event identity, durable ingestion, consistent aggregation rules, versioned pricing configuration, and reconciliation. The core pieces include:
The usage metering layer owns the measurement side of this path. Billing should be able to consume that state without reconstructing product activity from application logs or analytics tables.
Metering and billing for AI products gets harder because one customer action can create several different kinds of usage across models, tools, and execution paths.
An agent request might call one model, query a vector store, invoke an external tool, then call another model before returning a result. The meter has to preserve enough context to connect those events to the same customer and workload.
The next choice is the commercial unit. Different units expose different parts of the workload:
Tokens are useful when model consumption itself drives the commercial rule, but their economics still vary by model, context length, caching behavior, and workload.
The AI token cost model becomes useful when those differences need to feed usage attribution, internal costing, or customer-facing pricing.
Requests need more care once workload size varies. A short extraction and a multi-step agent run may both count as one request while consuming very different resources.
Credits solve a different part of the problem, by letting several workload types draw from one balance while the metering layer keeps the lower-level events available underneath.
A production credit system therefore needs richer state than a single balance field:
The usage store and credit ledger should have separate jobs. The usage store records what was consumed, while the credit ledger records how grants, debits, expirations, refunds, and adjustments changed the available balance.
Keeping those states distinct makes attribution, reconciliation, and request-time enforcement easier to reason about as AI workloads become more complex.
Pricing models change the metering path by adding different state and decision rules to the same usage events.
Pay-as-you-go needs a measured quantity, an applicable rate, and a reproducible path from event to charge. You can rate usage continuously or aggregate first and rate later, but both approaches need stable events, deterministic aggregation, and versioned rates.
Included usage adds allowance state and an overage boundary. The system needs to track:
The usage-based billing layer then prices consumption beyond the included amount. Late events matter here because they can push an account across the allowance boundary after a period has been calculated.
Subscription plus credits adds a consumable balance with its own lifecycle.
Each usage event can affect the credit deduction, remaining balance, burn order, depletion rule, and billing state.
Shared balances also introduce concurrency. Several requests can draw from the same balance at once, which calls for atomic debits, idempotency, and ledger-backed credit state.
The pricing model changes how much state the metering path has to carry forward before billing or runtime enforcement can act.
Metering and billing break when different stages disagree about event identity, time, quantity, pricing configuration, or financial state.
Duplicate usage is easy to miss because every downstream step can still behave correctly. Aggregation, rating, and invoicing simply preserve the wrong input.
When it comes to late events, you need a clear policy for whether they reopen the period, create an adjustment, or move into the next billing cycle.
Observability should surface both problems early. Useful signals include:
Those metrics give you a chance to catch a bad usage trail while it is still an engineering problem, before it turns into a billing dispute.
Metering and billing stop at measurement and financial processing. Request-time enforcement handles the separate decision of whether the next protected workload can run.
For AI products, timing matters. Usage can keep accumulating while earlier events are still moving through aggregation, rating, and billing.
Concurrency adds another failure mode. Several requests can read the same allowance before any debit commits, which makes stale state look valid to every request.
A production request path needs:
Hard and soft limits define what happens when the allowance is exhausted. A hard limit fails closed and denies further usage. A soft limit allows the request, flags it as overage, and lets you track and bill it instead of blocking it.
The AI billing infrastructure layer connects those runtime decisions with the metering and financial systems downstream.
Once usage needs to affect the next request, the problem has moved beyond counting and charging into live state and enforcement.
Knowing that an account has 120 credits left is useful. However, deciding which of several concurrent requests gets to spend them is the harder infrastructure problem.
That decision needs current metering state, entitlement rules, and credit state in one request-time path. Stigg handles that runtime layer, enforcing entitlements, credits, usage limits, and spend governance synchronously in the request path.
That runtime holds p99 entitlement checks under 10ms while ingesting 1M+ events per second on BYOC across 1M+ entities per hierarchy root, with multi-region active-passive failover and up to 99.99% uptime. This is the scale budget most in-house builds run out of before the second product line ships.
For a metering and billing stack, the relevant pieces include:
A useful architecture check is to follow one live request. Find where usage is measured, where the allowance lives, where credits are debited, and where execution is finally allowed or blocked.
The Stigg docs map metering, Sidecar checks, credits, entitlements, and BYOC onto that production request path.
The main difference between metering and billing is measurement versus charging. Metering captures and aggregates usage into a billable quantity, while billing applies pricing to that quantity and creates the customer’s financial charge.
No. Metering and billing describe the full path from usage event to financial record, while metered billing charges customers based on the usage measured by that path.
You meter usage for an AI product by tracking a defined unit such as tokens, requests, agent actions, compute time, or credits. Stable event IDs, attribution, and deduplication keep retries and concurrent activity from distorting the total.
Incorrect usage-based bills often come from duplicate or missing events, inconsistent aggregation windows, stale pricing rules, late usage, or reconciliation errors. Durable ingestion, stable event identity, consistent windows, and reconciliation help catch those failures.
No. Metering and billing record consumption and turn it into a charge after usage occurs. Preventing further usage requires request-time enforcement against the current entitlement, credit balance, or usage limit before protected compute begins.