%20(1).png)
Software Billing Models: 8 Types and How to Choose
Software billing models decide how you charge and what you must build. Compare subscription, usage, credits, and hybrid models, and how to choose.
Metered utilization is the measurement layer beneath every AI bill. Learn what gets metered, how the pipeline breaks, and why measuring can't enforce limits.
%20(1).png)
A duplicate usage event can feed the same credit balance, usage limit, and invoice twice. AI products add extra challenges because one action can fan out across several models, tools, and concurrent requests.
Metered utilization is the measurement layer that turns this activity into a trustworthy usage record. This guide covers what AI products meter, how events become billable usage, where metering pipelines fail, and where request-time enforcement begins.
Metered utilization is the measurement layer that records how much product usage belongs to a customer, feature, and time window.
That usage record can feed several downstream systems:
The measurement layer is upstream of rating and invoicing. Our broader guide to usage metering covers instrumentation, counters, and aggregation.
AI products meter the usage unit that best represents what customers consume. Common units include:
One user action can fan out across several model calls and tools. Event identity and attribution become part of the metering design when those calls contribute to the same customer balance.
Credits give mixed workloads one customer-facing unit while the metering layer keeps lower-level usage underneath.
In Stigg, credits are issued as blocks with their own expiry, cost basis, paid-versus-promotional category, and burn order, and the depletion behavior is configurable per feature (hard limit denies the request, soft limit allows the balance to go negative and reconciles later).
The metered unit also shapes event volume, cardinality, aggregation windows, and attribution logic. Raw tokens create high event volume and per-model metadata. Agent actions introduce fan-out and attribution work, and the unit you choose affects the architecture around it.
Metered utilization moves from event to billable signal through instrumentation, capture, normalization, aggregation, and handoff.
Each stage has a different failure mode. A missing event lowers the count, a duplicate inflates it, and a bad window boundary can assign valid usage to the wrong billing period.
The output is a usage record that billing, credits, reporting, and runtime controls can read.
A metering system relies on durable event capture, stable identity, aggregation, storage, and reconciliation.
Two components deserve extra attention:
If your ingestion layer uses at-least-once delivery, the same event can arrive more than once.
Stable IDs and deduplication keep retries from inflating usage totals. You can test this by replaying known duplicates and confirming that the aggregate stays unchanged.
Credit state needs a separate append-only ledger for grants, debits, expirations, refunds, and adjustments.
The usage store records consumption events, while the credit ledger records balance changes.
Useful operational signals include:
Those signals can expose metering problems before they reach an invoice or customer balance.
Metered utilization differs from adjacent billing concepts because it owns the measurement of consumption.
Metered utilization produces the usage record, while metered billing reads that record during rating and invoicing.
Usage-based pricing and consumption-based pricing define how measured usage maps to commercial terms.
Clear ownership between these layers makes changes easier to reason about. A pricing update can leave the event pipeline untouched when the underlying metered unit stays the same.
Metering ends after it records and updates usage state, and runtime enforcement uses that state to decide whether a protected compute can run.
Agent workflows make that decision harder because concurrent calls can contend for the same allowance. Several requests may read the same remaining balance before any debit commits, and a production request path needs controls for that shared state.
Key requirements include:
Hard and soft limits define what happens when usage reaches the configured boundary.
A hard limit blocks further usage when the allowance is exhausted, and a soft limit follows a configured grace or overage policy.
If the product needs to stop further compute, the threshold has to be evaluated in the request path.
Metered utilization tells you how much a customer has consumed. A usage runtime uses that state to decide what can happen next while the request is still active. Stigg combines metered usage with credits, entitlements, and limits to make a synchronous decision before more AI work runs.
Your billing system can keep handling invoices, payments, tax, and financial records. Stigg handles the product-facing state that determines what the customer can consume next.
A useful way to map the architecture is to trace one request from measurement to decision:
Metered utilization creates the usage record. The runtime uses it to make the next allow, limit, or credit decision. If you’re mapping that flow through your own stack, the Stigg docs show how metering, Sidecar checks, credits, entitlements, and BYOC fit into the production request path.
Metered utilization is the measurement of how much product usage a customer, feature, or workload consumes within a defined period. For AI products, that can include tokens, credits, inference units, API calls, agent actions, or compute time.
No. The main difference between metered utilization and metered billing is measurement versus charging. Metered utilization produces the usage record, while metered billing uses that record to calculate a charge.
AI products meter tokens, credits, inference units, API calls, agent actions, and compute time, depending on the workload. The best unit is one the product can measure consistently and attribute to the correct customer, feature, or usage window.
Accurate metering gets harder under high event volume because retries, duplicate events, concurrency, and high-cardinality dimensions can push usage totals out of sync. Stable event IDs, durable ingestion, deduplication, and reconciliation help keep the usage record consistent.
No. Metered utilization records consumption and supplies the usage state that runtime controls can evaluate. Preventing further consumption requires a request-time decision against the current entitlement, credit balance, or usage limit before protected compute begins.