%20(1).png)
Software Billing Models: 8 Types and How to Choose
Software billing models decide how you charge and what you must build. Compare subscription, usage, credits, and hybrid models, and how to choose.
Real-time billing calculates charges the moment usage happens. See how it works for AI products, what breaks under load, and where enforcement fits in.
%20(1).png)
An enterprise customer turns on a new agent workflow and requests start fanning out across your API. The meter is working, but the balance still updates in a scheduled job.
A few hours later, billing catches up. The customer crossed their limit long ago, and the compute is already spent.
Real-time billing keeps usage and charges current as events arrive. That gives engineering fresher commercial state to work with, but it also introduces new problems around concurrency, retries, pricing versions, and latency.
Real-time billing calculates and records usage charges as billable events arrive, which keeps spend and account state close to the latest customer activity.
Traditional billing can wait for a nightly job or the end of a billing period. That works well for predictable charges where a few hours of stale usage data carries little operational risk.
AI workloads change the timing, though. One user action can trigger several model calls, tool invocations, or agent steps, each carrying its own cost.
You now have two clocks to think about. One tracks when the financial state updates, while the other tracks when the product needs to decide whether another request can run.
Here's the practical difference:
Real-time billing solves the timing problem on the financial side. A live product decision still needs access to current usage, balances, and commercial rules before expensive work begins.
That live decision (allowed or not, before the compute runs) is where a usage runtime sits, separate from the billing system that records the charge.
Real-time billing moves each usage event through metering, normalization, rating, state updates, and reconciliation.
You can think of it as a short financial pipeline running alongside the product.
The process starts when the product emits a billable event. An API call might include the customer ID, feature, model, timestamp, quantity, and a stable event ID. An agent workflow may emit several events as it moves through tools and model calls.
Good usage metering tends to be fairly boring, which is a compliment.
Every valid event needs the right customer, quantity, timestamp, and identity. If a retry becomes a second charge here, every later calculation can be perfectly correct, and the invoice will still be wrong.
Raw infrastructure rarely speaks the same language as pricing.
One model might report input and output tokens separately, while another workload tracks inference seconds. Your customer may buy credits that hide both.
Normalization converts those events into a consistent schema while preserving enough detail for debugging later. Aggregation then rolls them into whatever unit the pricing model expects, such as tokens per request, GPU seconds per job, or API calls per account.
Rating applies the commercial rule to the measured usage.
A customer may have included usage, a negotiated overage rate, a volume discount, or a credit conversion that differs by model. Pricing versions matter here. If an account changes plans at noon, usage from 11:59 a.m. should still carry the earlier commercial state.
A good billing software architecture leaves a trace from the original event to the rate that produced the charge. You want to answer "why is this $37.42?" without reverse-engineering the invoice.
Once rated, the event can update current spend, accrued charges, or another customer-facing balance. This is where real-time starts to feel useful outside finance.
A usage dashboard can move during the session. An alert can fire before the billing cycle closes. Support can see what the customer has consumed without waiting for yesterday's batch.
Current state can also feed product controls, though the live allow-or-block decision needs stricter timing.
Real-time processing does not make reconciliation disappear. Raw events, aggregates, rated charges, credits, adjustments, and invoices still need to tell the same story.
The advantage is fresher data. Engineering can investigate a mismatch while the underlying events and pricing context are still easy to trace.
Replayable history gives the team a path back through every decision. When an account moves from one balance to another, engineers should be able to follow the events, pricing rules, and adjustments that led there. A missing trail turns a billing issue into detective work.
Real-time billing, metered billing, and usage-based billing describe different parts of the same system.
The terms often get blended together because a single product can use all three:
Metered billing can run in a batch or close to event time. Usage-based pricing can work with either approach.
Real-time billing describes how quickly financial state changes after usage occurs.
Request-time enforcement has a different timing requirement. It needs the answer before protected compute begins.
Keeping those concepts separate makes architecture discussions much cleaner, especially once credits and limits enter the picture.
Real-time billing matters for AI workloads because usage and cost can change with every request. A short chat response may cost very little, but a coding agent may run several models, call external tools, and retry failed steps before returning an answer.
One button click can trigger:
A single session can leave a long cost trail. The commercial unit may be tokens, credits, inference seconds, agent actions, or a weighted mix. The underlying AI token cost keeps growing as the workflow runs.
A delayed billing pipeline may still calculate the correct charge later. During the delay, another expensive request can start after the account has crossed its allowance or budget.
Real-time processing gives teams current spend visibility during the session. Customers can track consumption, while engineering gets fresher data for alerts, budgets, and product controls.
A soft overage can continue and reach the next invoice. A hard limit needs a live balance check before another workload begins. The second case moves billing logic into the request path.
Real-time billing usually gets difficult around concurrency, retries, pricing state, cache freshness, and reconciliation.
Production traffic is where the assumptions get tested.
Concurrency becomes a problem when several requests try to spend the same visible balance.
A wallet has enough credit for 40 operations, and 50 requests arrive within a few milliseconds. Several workers may read the balance before the first debit commits, and each one sees available value and approves the work.
The code often follows a familiar sequence:
Running those steps across many workers can create an overdrawn wallet. The fix is an atomic check and debit that treats approval and consumption as one operation.
Longer jobs may need a reservation. If the reservation can't be covered, the request fails closed rather than running on hope.
Replicate documents a similar edge case in its prepaid credit system. New work stops when the balance reaches zero, though an in-flight prediction can still run past the available credit.
The test is simple. Fund a wallet for 40 operations, send 50 requests in parallel, and count the approvals. Exactly 40 requests should clear across repeated runs.
Retries can turn one customer action into two billing events. A network drops the response, a queue redelivers the message, or a client resends the request after a timeout. The second copy reaches the pipeline looking like fresh usage.
Mistral's workflow documentation describes activities that may retry after partial side effects. Replicate makes the same warning in its webhook documentation, where identical events can arrive more than once.
A stable event ID lets the pipeline recognize each copy as the same logical action. The first event records the charge or debit, and later copies receive the existing result.
The idempotency check needs to cover the financial side effect itself, including the ledger entry, credit deduction, or rated charge. Checking the API response alone can still leave the customer with two charges.
Pricing changes create historical state because each usage event belongs to the commercial rules active when it occurred.
Let's say a customer upgrades at 2:03 p.m. An agent run begins at 2:02, finishes at 2:05, and reports usage at 2:06. The event still belongs to the rules active when the run started.
Enterprise contracts add more cases. A discount may begin at midnight in the customer's timezone, one workload may use a custom credit conversion, or a mid-cycle amendment may change the included allowance.
Each event should preserve:
Replaying an old event against the current catalog can change a settled charge. Versioned pricing state keeps each event tied to its original commercial rules.
Late events make the risk easy to miss. Tuesday's usage may arrive on Friday after the customer has changed plans twice. In this case, Friday's configuration cannot explain Tuesday's charge.
Caches create a consistency problem when fast checks use outdated commercial state.
Local caching removes a central API call from the request path, and also creates another copy of the customer's plan, balance, or limit that needs to stay current.
If a customer upgrades while requests are running across several regions, one node sees the new plan immediately. Another, meanwhile, continues using the previous allowance. The customer gets different access depending on which node handles the request.
A production design needs clear rules for:
The right failure policy depends on the commercial risk. A low-cost feature may continue briefly with older state, but an expensive agent run tied to a hard balance may require a fresh check before execution.
Cache freshness becomes part of the product experience. When two nodes disagree about a customer's allowance, the inconsistency reaches the customer before it reaches the invoice.
Real-time billing updates financial state as usage happens, while real-time enforcement checks commercial state before protected usage begins.
Credits make the difference easy to see. A customer has three credits left when two expensive agent jobs arrive together. Both debits may post immediately and leave the wallet below zero. The balance is current, and the extra compute has already run.
A request-time control reads the current balance, applies the relevant entitlement or limit, and reserves or debits the required amount before the workload continues.
For usage with real marginal cost, timing becomes part of the economics. Billing can own invoices, tax, collections, and financial records. The runtime layer handles the live commercial decision close to the request.
The broader runtime and billing stack connects those layers while giving each system a clear job.
Runtime controls become useful when current usage needs to affect the next request before more cost is created.
Stigg is the usage runtime for AI products. It can enforce entitlements, credits, usage limits, and spend governance synchronously in the request path while your billing provider keeps handling the financial record.
For a real-time billing architecture, the pieces look like this:
This gives engineering a cleaner boundary. Billing can keep answering what the customer consumed and owes, while runtime controls handle the live decision before another protected workload starts.
The Stigg docs walk through how metering, credits, entitlements, Sidecar checks, and BYOC fit into that production request path.
Real-time billing calculates and records charges as usage events occur, keeping current spend and billing state close to the latest customer activity.
Real-time billing processes usage events as they arrive.
Each event is validated, deduplicated, assigned to the correct customer, and rated against the active pricing rules. The system then updates balances, limits, or accrued charges. Reconciliation checks those results against invoices and financial records.
The main difference between real-time billing and metered billing is timing versus measurement. Metered billing measures and rates consumption. Real-time billing describes how quickly those rated charges update.
No. Real-time billing can keep usage and charges current, while overage prevention requires a control before protected usage runs. Hard limits, atomic credit debits, and entitlement checks can make that decision.
The main difference between real-time billing and real-time enforcement is when each system acts. Billing updates the financial effect of usage. Enforcement approves, limits, or blocks protected usage before the workload begins.