%20(1).png)
What is Metered Utilization? Measuring AI Token Usage
Metered utilization is the measurement layer beneath every AI bill. Learn what gets metered, how the pipeline breaks, and why measuring can't enforce limits.
Software billing models decide how you charge and what you must build. Compare subscription, usage, credits, and hybrid models, and how to choose.
%20(1).png)
I always find billing most interesting when the invoice and product dashboard disagree. Finance sees one total, the customer sees another, and engineering gets pulled into reconstructing weeks of usage events.
The root cause often stems from the software billing model and how the product implements it. Subscriptions, seats, usage, credits, and hybrid plans each create different state to manage.
Here’s what your team should expect from all eight.
Software billing models define how a company calculates what a customer owes for access to or use of a software product. A model may charge by time, seats, consumption, credits, outcomes, or a mix of those units.
Pricing meetings tend to blend billing, packaging, and licensing into one conversation. Keeping them separate makes the engineering work much easier to reason about:
A yearly Pro plan makes the distinction concrete. The license grants access, the package includes five seats and 20,000 credits, the price is $12,000, and the billing model collects the fee annually.
All four choices meet inside the product. A mid-cycle plan change can update access, issue a new credit grant, and activate a different overage rate while customer jobs are still running.
Our guide to software licensing models covers the licensing side in more detail.
The main software billing models are subscription, per-seat, usage-based, tiered, freemium, prepaid credits, hybrid, and one-time licensing. Each model creates a different balance between predictable revenue, customer flexibility, and engineering work.
Here is a quick map before we get into what each model asks your team to build:
The tradeoff tends to appear after launch. A flat subscription is easy to explain until one account consumes ten times more compute. Usage billing follows consumption closely, though customers may worry about an open-ended bill. Credits cap spend, but the wallet behind them needs careful engineering.
Subscription billing charges a recurring fee on a fixed schedule, often monthly or annually. It works well when customers receive steady value from ongoing access, recurring allowances, or a defined set of features.
Midjourney uses this model across its Basic, Standard, Pro, and Mega plans. Each plan renews monthly or yearly and includes a different amount of Fast GPU time and product access.
The customer experience feels simple. You pick a plan, pay on a schedule, and keep creating. But engineering still has to handle everything between those renewal dates.
A customer can upgrade halfway through the month while image jobs remain active under the old allowance. Proration, access, and plan state now need the same effective time.
Your billing provider can calculate the new charge and retry a failed card. The application still needs a current answer about which plan applies, how much usage remains, and whether the customer stays active during a grace period.
Subscriptions give customers predictable bills, so they fit products with steady delivery costs or plans that include clear usage limits.
Per-seat billing charges according to how many people have access to the product. The unit is familiar, contracts are easy to quote, and buyers can budget against headcount.
Microsoft 365 Copilot follows this model. Businesses buy Copilot licenses per user, then assign those licenses to people inside the organization.
The pricing math is simple, but defining what counts as a seat takes more thought. An invited user, an activated account, and an assigned license can represent three different billing events.
Those rules need to stay aligned across identity, provisioning, entitlements, and billing. If the product counts 30 active seats while the billing system still sees 35 licenses, the discrepancy eventually lands with finance or support.
Per-seat billing works cleanly when access maps closely to individual users. It gets more complicated once licenses move frequently, admins need special treatment, or usage varies widely between people.
Usage-based billing charges for measured consumption, such as tokens, API calls, compute seconds, generated images, processed documents, or completed runs.
The OpenAI API uses this model. API charges depend on the model and the number of input and output tokens processed, which connects the bill directly to customer activity.
A production meter validates events, assigns them to the correct account, removes duplicates, places late usage in the correct period, and applies the pricing version active when the request ran.
The messy version starts with a timeout. The client retries, the original request finishes in the background, and two usage events reach the pipeline for one customer action. Stable event IDs keep the retry from becoming a second charge.
Our metered billing guide explains the event pipeline in detail. Consumption-based billing covers the commercial model when each action carries a different cost.
Usage pricing feels intuitive when consumption tracks value, but the first unexpectedly large invoice can change that reaction quickly. Budgets, alerts, commitments, and prepaid balances help customers keep spending visible.
Tiered billing applies different rates at defined usage thresholds. The first 10,000 units might carry one rate, the next 40,000 another, and everything beyond 50,000 a third.
You see this in API products such as AssemblyAI. Its public pay-as-you-go pricing gives customers a clear starting point, while higher-volume customers can move onto custom agreements with volume discounts.
The detail that matters is how the threshold changes the rate. Plan tiers such as Starter, Pro, and Enterprise define product packages. Billing tiers determine how measured usage gets priced inside those plans.
Two common approaches behave differently:
Crossing 50,000 units can therefore produce two very different invoices, even when the published tiers look almost identical.
Live traffic makes the boundary harder to manage. Several requests may cross the threshold together while a late usage event lands at the same time. Metering, rating, and billing need one consistent view of which usage belongs in which band.
Our tiered pricing guide goes deeper into graduated, volume, and package tiers, plus the infrastructure needed to keep those boundaries consistent.
Freemium gives customers ongoing free access with paid limits or features above it. It works best when the free plan delivers enough value to build a habit, while paid plans raise capacity limits or open additional capability.
Perplexity is a good example. Its Standard plan gives users free access to core search, while Pro and Max increase usage limits and open access to more advanced features.
The cost profile gets interesting once people start using the free tier heavily. A few users can automate repeated searches, create several accounts, or lean on expensive models long before subscription revenue catches up.
For engineering, “free” is really a bundle of entitlements. A plan might define:
Each allowance needs a clear enforcement point. The product has to know when the user reaches a search cap, which models they can call, and whether another upload should go through.
The harder decision is where to put the boundary. A tight free tier can end the experience before users see enough value, while a generous one can leave you paying for meaningful production usage with little reason for heavy users to upgrade.
For products with variable compute costs, freemium design eventually becomes an infrastructure question too. You need current usage state, plan entitlements, and limits that behave consistently across concurrent requests.
Prepaid credit billing lets customers purchase a balance before they consume the product. Credits give buyers a firm spending limit while the product translates several infrastructure costs into one visible unit.
Replicate requires prepaid customers to purchase credit before running models. Usage draws down the balance, and new work stops once the balance reaches zero.
The request path may look simple from the outside. A customer runs an image model, and the balance decreases. Behind that action, compute time can vary by model, hardware, input, and runtime.
The first implementation often looks like balance = balance - 5. Production soon adds concurrent requests, auto-reloads, promotional grants, refunds, expiry dates, and shared account balances.
A reliable wallet needs:
The ledger earns its keep during support calls. When a customer asks why 2,000 credits disappeared, your team can trace each debit back to the model run that created it.
Hybrid billing combines two or more billing models inside one offer. A common setup pairs a recurring subscription with included usage, then charges for extra consumption once the allowance is gone.
Cursor is a useful example here. Its plans include monthly model usage, and customers can enable on-demand billing when they run through the included amount.
The offer feels simple from the customer side. You pay a predictable subscription, use the included allowance, then keep working during a heavier month.
Underneath, though, several systems have to agree about the customer at the same time.
The awkward moments happen around state changes. A renewal resets included usage just as automated coding agents start new sessions, or a plan upgrade takes effect while requests are already in flight.
Every request needs the same commercial state. If the editor sees the new allowance while billing still uses the previous contract, you can end up with usage that is allowed under one system and charged under another.
Hybrid plans therefore need clear ownership across subscription state, metering, allowances, entitlements, and rating. The hybrid pricing model guide goes deeper into how those pieces work together.
One-time billing collects a single payment for a permanent or version-bound software license. Support, upgrades, hosting, and external services can carry separate charges.
TypingMind sells its personal plans as one-time purchases with licenses that remain valid permanently. Customers provide their own model API keys and pay the model provider separately for usage.
The arrangement separates the interface license from the ongoing cost of model calls. TypingMind receives the one-time software payment, while the API provider handles consumption charges.
The engineering lifecycle continues after the sale. A customer replaces a laptop, adds a device, restores an offline installation, or requests access to a newer version. License issuance and device state still need clear rules.
Revenue arrives near the start, while product updates and support can continue for years. Paid upgrades, maintenance agreements, and separate usage charges give vendors a way to fund the work that follows the original sale.
Every software billing model creates commercial state that has to survive production. Retries, concurrent requests, upgrades, and late events all test assumptions that looked perfectly reasonable in the first implementation.
I’ve found the same five engineering concerns keep coming back, whether the product charges by seats, usage, credits, or some hybrid of them.
Billing events rarely take a clean path. An API times out, the client retries, and a queue redelivers the original message. One customer action can suddenly look like three billable events.
Give each event a stable ID and make the financial side effect idempotent. Deduplicating the API request alone is not enough if two workers can still write separate charges or ledger entries.
Late events need a rule too. You might rate them into the original period, create an adjustment later, or reject them after a cutoff. Pick the behavior up front, because support will eventually need to explain it to a customer.
A customer upgrades at noon after using the product all morning. Those earlier events still belong to the old commercial terms.
Store effective dates and immutable pricing versions with enough context to reconstruct the charge later. Re-rating historical usage against today’s plan is a quick way to make last quarter’s invoice change underneath you.
A useful test is simple. Open an old invoice and ask whether engineering can reproduce every charge using the exact pricing rules that applied at the time.
Credits get interesting the moment several workers touch the same wallet. Say a balance can fund 40 jobs and 50 arrive together. A basic read, check, write flow lets several requests see the same available credits before another debit commits.
The safer options are:
Reservations are especially useful when the final charge is unknown up front. An agent run can reserve enough credits to start, then reconcile the actual usage after completion.
An accurate invoice can tell you an account went over its limit. It cannot recover compute that already ran. Hard limits have to act before the expensive operation begins.
This puts commercial state close to the request path. Before an agent run, API call, or model request proceeds, the application may need to check the customer’s current entitlement, allowance, or credit balance.
Caching keeps those checks close to the application, but your team still needs clear rules for:
The last one is easy to overlook until there’s an outage. A low-cost feature may continue safely, while an expensive agent run may need to wait. This means failure behavior should be designed ahead of the incident.
Eventually, someone will ask why the product says one number and the invoice says another.
Raw usage events, aggregates, ledger entries, rated charges, credits, and invoices should give engineering one traceable history of the account.
Fresh processing helps you catch mismatches early, and replayable history is what helps you fix them.
When a total looks wrong, you want to rebuild it from the original events and pricing versions, find the first divergence, and explain exactly how the account arrived there.
AI workloads change software billing models by making cost variable at the request level and making the billable unit harder to define.
A customer clicks “run” once. Behind that action, an agent may plan, retrieve data, call tools, retry failed steps, and invoke several models before returning a result. The UI shows one action, but the infrastructure may record a long chain of costs.
That creates a few common billing patterns:
Outcome pricing needs especially careful event design. Engineering has to define what counts as completion, how partial success is treated, and whether a retried workflow represents the same billable outcome.
I’d pay close attention to retries here. A workflow can fail after most of the compute has already run, then restart and eventually succeed. Billing logic needs a stable identity for the outcome, or else one customer action can become two charges.
Real-time visibility also becomes more useful as cost accumulates during a session. Some products only need a current dashboard and spend alerts.
Others need request-time enforcement. If a wallet reaches zero or a budget cap is hit, the next model call or agent job has to be checked before more compute starts.
Choose a software billing model by matching the billing unit to customer value, delivery cost, buying preference, and what your systems can enforce. I would give the last factor more weight than it often gets, as production has to measure and apply every rule the pricing team creates.
Start with these questions:
Here is a practical starting point:
The model can evolve as the product matures. Keep the underlying events, catalog, entitlements, and ledger modular enough to support a new package without rewriting the entire commercial path.
Software billing models break most often where pricing rules meet live traffic. The plan looks calm in a spreadsheet, while production brings retries, concurrent requests, stale caches, and contract changes into the same boundary.
Fund a wallet for 40 operations, send 50 in parallel, retry several requests, change the plan halfway through, and confirm which jobs ran and which charges appeared.
Run the same test against tier boundaries and billing-period rollover. Your future support team will thank you.
Building the runtime behind software billing models means connecting plans and prices to live product behavior. The invoice sits at the end of that path. Before it exists, the product still has to decide who can run a job, which balance pays for it, and what usage gets recorded.
Stigg is the usage runtime for AI products, which means the enforcement layer that resolves entitlements, meters usage, and governs credit spend in the request path, while your existing billing stack keeps running underneath it.
Before an expensive job starts, your application can resolve the active plan, check an entitlement, reserve credits, and record usage while the billing provider continues to own invoices, payments, tax, and collections.
Here is what it adds:
Miro, for example, shipped a credit-based AI pricing model in under 6 weeks and saved roughly 5,000 engineering hours by moving credit and entitlement handling into a dedicated runtime.
You can adopt each component as the need appears. Metering may come first, followed by credits or entitlement checks as the pricing model grows. The result is one consistent commercial decision across the live product and the final invoice.
Follow the path from a usage event to a request-time decision in the Stigg docs, then map the pieces to your current architecture.
A software billing model defines how customer activity turns into a charge. A company might bill by subscription period, seat, usage, credits, outcomes, or a mix of several units.
The most common software billing models are subscription, per-seat, usage-based, tiered, prepaid credits, freemium, hybrid, and one-time licensing.
The useful distinction is what each model tracks, since that determines the metering, billing, and enforcement infrastructure you need behind it.
The main difference between a billing model and a pricing model is what each one defines. The billing model determines how charges are calculated and collected, while the pricing model sets the rate, tiers, discounts, and other price rules applied to that model.
Usage-based, prepaid credit, and hybrid billing often fit AI products because compute cost changes with customer activity. Usage billing exposes the meter directly, credits give customers a defined balance, and hybrid models pair a predictable base fee with variable consumption.
Yes, a company can change its software billing model, but existing contracts, usage history, balances, and pricing versions need a clear migration path. Versioned pricing and durable usage records make it much easier to move customers without rewriting historical charges.