%20(1).png)
Software Billing Models: 8 Types and How to Choose
Software billing models decide how you charge and what you must build. Compare subscription, usage, credits, and hybrid models, and how to choose.
See how a real time billing system keeps AI usage, credits, pricing, latency, and reconciliation aligned before small billing errors get costly under load.
.png)
A customer upgrades mid-session, launches another agent run, and the product keeps using the previous allowance because billing state hasn't caught up.
Billing tells you what happened, but a real-time billing system decides what's allowed while the request is still open.
A real-time billing system processes usage events continuously and updates rated charges or billing state close to the time consumption occurs.
The path typically covers:
The point is preserving enough context to explain how each number got there. “Real-time” doesn’t mean every billing step has to sit inside the application request.
You can process usage within seconds and still keep invoicing downstream. If the product needs to allow or block the next request before compute starts, that becomes a separate request-time decision.
The metered billing layer turns measured consumption into a charge. Real-time billing keeps that path current as new usage arrives.
Real-time billing processes usage continuously, while batch billing collects events and processes them on a defined schedule.
Batch processing can work well when usage is predictable, and the product doesn’t depend on current billing state during execution.
The trade-offs become clearer when event volume, pricing rules, or usage costs increase.
A nightly job gives errors more room to pile up, as one bad window boundary can misclassify a whole batch of usage before anyone spots the mismatch.
Real-time processing keeps that window smaller, but asks more from ingestion, metering, pricing state, and storage.
Either way, the billing system architecture needs clear ownership for each piece of state as it moves through the pipeline.
A real-time billing system works by moving each usage event through ingestion, metering, rating, financial state, and reconciliation without waiting for a period-end batch.
Each stage inherits the output of the previous one.
A duplicate at ingestion can move through every later stage cleanly, while rating and invoicing may behave exactly as designed while the final quantity is still wrong.
Event ingestion creates a durable record of what happened, who consumed it, and when.
A usage event commonly carries:
Retries are part of normal distributed-system behavior.
AWS Standard SQS queues, for example, use at-least-once delivery and may deliver more than one copy of a message. Applications using that pattern need idempotent processing.
Stable IDs let the ingestion layer recognize a repeated event before it changes the measured quantity.
Replay needs equal attention. If processing pauses, you should be able to replay durable source events without billing the same consumption twice.
Aggregation converts raw usage into the exact quantity the rating layer will price. Different meters need different functions:
A token meter might begin with:
tenant_id + model + billing_period
Then you add workspace, feature, region, or agent. Each dimension gives you better attribution, but it also increases the state the system has to store, query, and reconcile.
Time boundaries need the same level of care. A late event can arrive after its usage window closes, which forces a policy decision.
You need to decide whether that event reopens the period, creates an adjustment, or lands in the next billing cycle. Leaving that undefined is how two systems can process the same event correctly and still disagree on the bill.
Rating turns measured usage into the charge that should apply under the customer’s commercial terms at that moment. A typical path might look like:
quantity → allowance → pricing tier → contract rule → overage → charge
That path can pull from several pieces of pricing state:
The tricky part is making every charge reproducible later.
Say an enterprise customer negotiates a custom rate halfway through a billing cycle. Usage before the contract takes effect still needs the previous rate, while new usage picks up the override. If pricing state gets overwritten, recreating that invoice later becomes guesswork.
The billing software architecture needs one authoritative source for pricing versions, effective dates, and contract overrides before those exceptions start piling up.
Billing state turns rated usage into charges, balances, adjustments, and the financial record that eventually reaches the customer. The important part is keeping enough context to explain how each number got there.
A useful billing record should preserve:
You don’t want an invoice line to be a dead end. If a customer questions a charge, engineering should be able to trace it back through the rated usage, the aggregated quantity, and the original events without stitching the story together from three different systems.
Reconciliation checks whether the usage you recorded, the quantity you rated, and the charge you produced still agree.
That means comparing the trail end to end:
A mismatch by itself doesn’t tell you much. “Usage and billing differ” gives you a symptom. Something more specific, like “Late events landed after the window closed,” gives you a failure mode to investigate.
That level of traceability is what makes reconciliation useful in production, especially once usage volume and pricing rules get harder to reason about.
Latency matters once usage or balance state can change what the product does next. At that point, a few seconds can be perfectly fine for one part of the billing path and far too slow for another.
A usage dashboard updating a few seconds late probably won’t hurt anyone, but a credit check returning after the model call has started is a different story. The compute is already running, and the cost is already yours.
Caching helps keep request-time checks fast, but now the cache becomes part of billing correctness.
You need clear rules for:
The fastest response isn’t useful if it carries stale balance or entitlement data. For credits and access, latency and correctness have to be designed together.
AI workloads change real-time billing because one customer action can create several usage events across models, tools, and execution paths.
An agent request might invoke a model, query retrieval infrastructure, call an external tool, then invoke another model before returning one customer-facing result.
The meter needs enough context to connect those events to the right customer and workload.
Common commercial units include:
Tokens give you precise model consumption, but their economics vary across models and workloads.
The AI token cost model becomes relevant when those differences feed usage attribution, internal cost analysis, or customer-facing pricing.
Credits sit one level above those raw units. One balance can represent several workloads with different burn amounts underneath, while metering still preserves the lower-level consumption record.
A production credit model also needs richer state:
The usage store and credit ledger have separate responsibilities. The usage store records consumption, and the credit ledger records grants, debits, expirations, refunds, and adjustments.
Pricing models change the real-time billing path by changing what state has to be read, updated, and preserved as each usage event moves through the system.
The usage event may look identical at ingestion. What happens after that depends on whether you charge every unit, include an allowance, or draw from a shared credit balance.
Pay-as-you-go has the lightest state model. Each measured unit needs a rate and a reproducible path from event to charge. The basic flow is:
usage event → measured quantity → applicable rate → charge
You can rate events continuously or aggregate them first. Either way, the system still needs stable event IDs, deterministic aggregation, and versioned rates.
The tricky part shows up when rates vary by model, region, customer, or effective date. The rating layer has to resolve the correct rate for that event without rewriting historical charges later.
Included usage adds a running allowance that has to stay aligned with the billing period.
Now the system needs to track:
Each event updates consumption against that allowance. Once usage crosses the boundary, the metered billing layer starts pricing the excess.
Late events make this harder. An event that arrives after the period closes can push the account over its allowance retroactively.
Your billing logic needs a clear policy for whether that event reopens the period, creates an adjustment, or enters the next cycle.
Subscription plus credits adds a shared balance with its own ledger, depletion rules, and concurrency problems. A single event can affect several pieces of state at once:
usage → credit conversion → debit → remaining balance → downstream billing state
The system may also need to decide which credit block gets spent first, especially when paid credits, promotional grants, and expiring balances coexist.
You also need to consider concurrency, where requests try to spend from the same balance at the same time.
That calls for atomic debits, idempotent writes, and ledger-backed state. Otherwise, multiple requests can read the same available balance and all pass before any debit commits.
As the pricing model gets richer, real-time billing becomes much more about keeping several pieces of commercial state consistent while usage is still arriving.
A real-time billing system tends to fail at the handoffs between events, state, pricing, and reconciliation. These failures often stay invisible for a while. The pipeline keeps moving, but the wrong state moves with it.
Duplicate usage is a good example. If the same event enters twice, aggregation can sum it correctly, rating can price it correctly, and billing can invoice it correctly. The pipeline works as designed around bad input.
Late usage creates a different failure mode. An event may arrive after the billing window has closed, which forces a decision about whether to reopen the period, issue an adjustment, or push that usage into the next cycle.
Wrong pricing often comes from stale commercial state. A usage event can be valid, but the rating layer may read the wrong contract version, tier, or override.
Concurrency causes trouble when several requests touch the same balance at once. Without atomic updates, each request can read valid state and still leave the account overdrawn after the writes land.
That is why observability has to follow the billing state itself, not only service health. Useful signals include:
Those signals help narrow the failure domain quickly. A billing mismatch is much easier to debug when you already know whether it started at ingestion, aggregation, rating, or shared-state updates.
Real-time billing gets you fresh usage and pricing state. The next challenge is making that state useful before another request consumes compute.
Take a shared credit balance, where one request reads 50 credits remaining. Before its debit commits, three more requests read the same 50. All four can pass unless the balance update is atomic.
That is why request-time enforcement needs a tighter set of controls:
The billing system architecture can process charges and financial state downstream. The enforcement layer has to make its decision while the request is still waiting.
Build a real-time billing system when the billing logic itself gives your product an advantage and your team is prepared to own the failure modes that come with it.
The first version can look tiny:
event → counter → rate → invoice
The decision gets harder once production adds retries, late events, pricing versions, credit balances, concurrent debits, cache invalidation, replay, reconciliation, and audit history.
A useful way to decide is to look at what you actually want to own:
The biggest mistake is treating this as a one-time build.
A home-grown system becomes something you operate every day. Someone owns duplicate events, broken backfills, stale pricing state, negative balances, failed reconciliation jobs, and migrations when the commercial model changes.
That can be the right trade when those capabilities are strategically important.
If most of that work exists only to keep billing correct, a commercial infrastructure layer can remove a large amount of engineering ownership while leaving your product logic where it belongs.
Our build-vs-buy analysis for billing infrastructure goes deeper on which parts tend to become long-term operational work and which ones are worth keeping in-house.
Real-time billing software should be evaluated on event correctness, pricing state, latency, reconciliation, deployment, and how cleanly it fits your existing financial stack.
A useful technical review covers:
The billing system requirements for AI SaaS provide a broader checklist for evaluating those trade-offs.
You should also draw one real request through the architecture before choosing anything. Mark the event source, meter, pricing version, balance state, invoice destination, and any synchronous decision in the path.
That exercise tends to expose fuzzy ownership faster than a feature checklist.
A real-time billing system may know that 20 credits remain and still let several requests spend the same 20 credits. That happens when billing state is current, but the request path lacks atomic control over it.
That decision layer is what Stigg provides. Stigg is the usage runtime for AI products. It enforces entitlements, meters usage, and governs AI spend in the request path, not after the bill. It is the runtime layer that answers the enforcement question your billing stack alone leaves open.
With Stigg, teams can:
When Miro launched its AI collaboration features, it shipped a credit-based pricing model in under 6 weeks, with credits allocated by plan, tracked in real time, and reset each cycle, all without rebuilding its stack.
Keep the billing you have, and add the runtime decision it was never built to make.
If enforcing credits and entitlements in the request path is eating sprint capacity, the runtime layer is missing from your stack. See how Stigg fits into that architecture.
A real-time billing system processes usage events continuously and updates rated charges or billing state close to the time consumption occurs. It typically combines event ingestion, metering, rating, billing state, and reconciliation.
The main difference between real-time and batch billing is processing cadence. Real-time billing processes usage continuously as events arrive, while batch billing processes accumulated usage on a scheduled cycle.
No. Real-time billing describes how quickly the billing architecture processes usage, while usage-based billing describes a pricing model where charges depend on consumption. A usage-based model can run through either real-time or batch processing.
A real-time billing system handles high event volume through durable ingestion, stable event identity, partitioned processing, deterministic aggregation, and replay support. The exact architecture depends on event volume, cardinality, latency requirements, and recovery policy.
A real-time billing system can keep usage and rated state current, while preventing the next AI workload from exceeding a limit requires synchronous request-time enforcement. That decision needs current entitlement, credit, or usage-limit state before protected compute begins.