%20(1).png)
Best Metered Billing Software: 9 Tools Ranked (2026)
I tested 9 metered billing software platforms for metering accuracy, pricing, and real-time enforcement, with honest pros, cons, and prices for 2026.
The 8 API pricing models that show up in production, from per-call to token and credit pricing, plus the failure mode that breaks each one.
%20(1).png)
Picking an API pricing model takes an afternoon, and making it hold on every request takes the next two quarters.
Per-call pricing suits spiky, transaction-shaped usage, tiered pricing rewards volume, and flat plans keep delivery costs predictable. Still, credits fit multi-feature AI products best because one balance covers every feature.
If you're the engineer who has to make one of these survive real traffic, each model below comes with the failure mode to plan for.
API pricing models combine access fees, usage charges, and allowances in different ways. This table shows what each approach charges for and where it fits.
These approaches can work together. A subscription can include credits, a usage rate can include volume discounts, and a free tier can lead into any of the paid models below.
Charging for access and API usage gives you three starting points: a fixed subscription, a charge for each unit, or a combination of both. Volume discounts add another choice about how you calculate usage charges.
Flat-rate API pricing charges a fixed monthly or annual fee. A plan can include a request allowance, feature access, support, or a combination of these benefits.
For example, SerpApi offers a Starter plan at $25/month for 1,000 searches and a Developer plan at $75/month for 5,000 searches. A customer on Starter pays the same subscription fee whether they use 200 searches or the full allowance.
Customers get a clear budget, and you get recurring revenue and a straightforward offer to explain, which can work well when you understand the cost of serving the included usage.
The next decision is what happens at the allowance boundary. You can pause access, offer an upgrade, sell another usage package, or allow paid overages. Make that choice visible before a customer builds a production workflow around your API.
A monthly quota and a rate limit serve different purposes. The quota controls total included consumption, while a rate limit controls how quickly requests can arrive. A plan may need both.
Per-call pricing charges for each billable API request. More broadly, per-unit usage pricing can charge for a message segment, a processed document, a data transfer, or another measurable operation.
Suppose your API charges $0.01 per successful lookup. A customer who makes 10,000 billable lookups pays $100, and they can start with a small workload and pay more as usage grows.
Choose the unit carefully, since one request can produce several billable units. Twilio's SMS pricing, for example, charges per message segment. A longer text can span multiple segments, and additional carrier fees can apply.
Per-unit pricing fits best when the unit is easy to understand and has reasonably consistent delivery costs. If one endpoint runs a quick lookup and another launches an expensive AI workflow, separate rates can make more sense.
Before launch, define how failures and retries affect charges. A customer retrying a failed operation needs a clear policy. Your event pipeline also needs to recognize duplicate deliveries and avoid counting the same billable event twice.
Tiered usage pricing changes the unit rate as a customer's volume crosses defined thresholds. It gives larger customers a reason to commit more usage to your API.
There are two ways to apply the discount, and they produce different bills:
Take an illustrative rate card with the first 1,000 requests at $0.10 each and additional requests at $0.05 each. At 2,000 requests, graduated pricing produces a $150 bill, and volume pricing produces a $100 bill.
Real providers also use volume bands. Google Maps Platform's Geocoding pricing includes a monthly free usage cap and lower unit rates in higher usage bands.
Customers like volume discounts because higher usage earns a better rate. Your rate card needs to account for the threshold effect. A customer who crosses into the next tier may pay less overall after the new rate applies.
Show a worked example on your pricing page, as this helps customers estimate spend and gives your team a reference for checking the invoice calculation.
A subscription with usage and overage combines a recurring fee, an included allowance, and a rate for additional consumption. Customers get a baseline budget while their application can keep running beyond the allowance.
For an illustrative plan, charge $100/month for 10,000 requests, then $0.008 for each extra request. A customer using 15,000 requests pays $140, which is the $100 subscription plus $40 in overages.
This model fits products with recurring value and variable delivery costs. The base fee can cover access and a useful amount of usage, while extra charges account for larger workloads.
Make the transition into paid usage clear. Show the remaining allowance, the overage rate, and accrued charges, and let customers choose a spending cap or receive alerts before they exhaust the included amount.
Your application needs an answer when the next request arrives. Is it included, billable as an overage, or blocked by the customer's cap? The overage policy should govern that decision while usage occurs.
Pricing AI work by consumption or results means choosing between tokens, credits, and a defined outcome. Each gives customers a different way to understand the work they're buying.
Token-based pricing charges for the units of content a model processes or generates. Providers can set separate rates for input, output, cache writes, and cache reads.
Anthropic's pricing documentation, for example, lists these categories separately. Model choice and caching can change the cost of serving a request.
Let’s look at an illustrative rate of $3 per million input tokens and $15 per million output tokens. A request using 2,000 input tokens and producing 500 output tokens costs $0.0135. This example excludes caching and other fees.
Token pricing makes consumption visible, particularly for developers choosing models and tuning prompts, and it also gives customers several variables to estimate. Conversation history, response length, and agent retries can all affect usage.
If you're building on another provider's models, track your underlying cost alongside the amount you charge. Your product can set its own markup and packaging. Keep the token categories accurate and associate each operation with the customer who used it.
For spend controls, you may need to reserve a budget before generation starts, then settle the charge against actual token usage when it finishes. Set explicit limits for jobs that can continue making model calls.
Credit-based pricing gives customers one balance to spend across different operations. A short summary might cost one credit, a document analysis five, and an image generation ten. These are illustrative rates.
Customers can buy credits upfront, receive them with a subscription, or add more when their balance runs low. The credit rate lets you account for different workloads while presenting a shared pricing unit.
Stability AI illustrates this approach. Its credit costs vary by endpoint and request parameters, including the number of images a request generates.
Credits help when a product combines several models or features. Customers still need a clear rate card. Show what common actions cost and roughly how much work a credit package covers.
The balance also needs rules. Do unused credits carry over? Can a team share them? Which credits get spent first when a customer has both a paid grant and a promotion?
For example, you might consume trial credits expiring tomorrow before paid credits valid for another month. Tracking each grant's expiry and burn priority makes that policy possible.
Shared balances need safe concurrent deductions. If two agents reach for the last ten credits at the same time, checking the balance alone can let both proceed.
Atomic reservations or debits help prevent them from spending the same available credits. A ledger records the grants, charges, and adjustments needed to explain the balance later.
Stigg's credit pricing guide covers these grant, balance, and enforcement decisions in more detail.
Outcome-based pricing charges for a defined result, such as a resolved support issue or a completed workflow. It works when customers can recognize the result, and both sides agree on how to measure it.
Intercom's Fin AI agent lists pricing from $0.99 per outcome. At that rate, 500 billable outcomes produce a $495 usage charge. Other fees depend on the setup.
The customer can connect spending to a business result. For the provider, delivery costs remain variable. One outcome might take a single response, while another might require several model calls and tool interactions.
Define the billable result precisely. Document what counts, when you confirm it, and how you handle reopened issues or reversals. Keep enough evidence to explain each charge to the customer.
You also need a budget for attempts that fail to produce a billable outcome. Meter the work behind the result and set execution limits that protect your margins. Outcome pricing changes what you charge for, but the underlying resource costs still need management.
Offering free access before the paid plan gives developers room to test your API and understand its value. Freemium and trials define that entry experience and can lead into any of the pricing models above.
Freemium provides an ongoing free allowance. A trial limits access by time, usage, or both. Choose the approach that gives developers enough room to complete a useful test.
SerpApi's free plan, for example, includes 250 searches per month, and developers can try a small workload before choosing a paid allowance.
Size the free offer around a real evaluation. A document API should let someone process representative files, and an AI API should provide enough usage to test response quality across several prompts. The allowance needs to support that work while keeping your costs manageable.
Show remaining usage and explain the upgrade path early. If free access pauses at a limit, developers should know when it resets. If a trial expires, make the date visible.
Free plans need the same care around usage attribution and allowances as paid plans. Where a hard cap protects an expensive operation, check it before the work begins. Delayed usage reporting can leave room for consumption beyond the intended allowance.
Choose an API pricing model by matching the customer-facing unit to the value delivered, then checking whether the price covers the work behind it.
Start with representative customer workloads where you compare a small account, a growing account, and a heavy user. For each, calculate what they'd pay, what you'd spend delivering the service, and how easily they could predict the bill.
Then work through the decisions that affect daily use:
A subscription can include credits, a credit plan can allow top-ups, and an enterprise contract can set a custom allowance and overage rate. Choose combinations your customers can understand, and your application can apply consistently.
Enforcement is the part that breaks because every model above shares one weak point, which is the runtime underneath it, and billing only records what already happened.
Enforcement decides what is allowed before compute runs. A metering-to-invoice pipeline can be perfectly accurate and still let a request through when the balance is gone, since it reads state after the fact.
Picture the check as a question your gateway can't answer: May this customer, on this plan, with this balance, make this call right now?
Answering it takes entitlements, the commercial allowances that define which features a customer can use and how much of each, based on their plan.
An entitlement carries a limit and a running count, which separates it from RBAC (role-based access control) and its question of who may use a feature. Limits resolve across the base plan, a parent plan, add-ons, trials, and promotional grants, and the most generous valid value wins.
Four problems cluster together the moment you enforce a model in real time:
Concurrency is the thread tying idempotency, consistency, and enforcement timing together, which is why treating them as separate edge cases scatters enforcement across middleware and gateways.
Rate limits at the gateway only see volume. They can't see a customer's plan, credit balance, or add-ons, and that's the decision you need to make in the request path.
Your API pricing model leaves 5 questions to the request path, and each one needs an answer before compute runs. Stigg is the usage runtime for AI products, answering them synchronously across entitlements, credits, usage limits, and spend governance.
It sits beside Stripe, Zuora, or your custom billing and keeps performance steady at scale, including large, complex setups with heavy agent traffic.
Got a question your pricing page can't answer? Stigg's docs cover the Sidecar, entitlement layer, and credits ledger behind each one.
The main API pricing models are per-call (pay-as-you-go), tiered, flat/subscription, freemium, usage-plus-overage (hybrid), token-based, credit-based, and outcome-based. Most API businesses combine 2 or 3, such as a subscription base with metered overage.
The best API pricing model depends on the unit your buyer can forecast. Hybrid pricing, a subscription base plus metered overage, suits most scaled APIs, and AI-native APIs often add token-based or credit-based pricing on top.
An API pricing model is the structure that decides what a customer pays for access to an API, including the charged unit (calls, tokens, credits, outcomes), how units get packaged into plans, and how usage gets metered and billed.
The main difference between graduated and volume pricing is how units are priced across tiers. Graduated pricing bills each unit at the rate of the tier it falls into, and volume pricing bills every unit at the rate of the highest tier reached.
Per-token API pricing is calculated by charging separately for input and output tokens, typically per million, often at a lower rate for cached input. The total is tokens consumed times the per-token rate, reconciled against the provider's own token count.