%20(1).png)
Software Billing Models: 8 Types and How to Choose
Software billing models decide how you charge and what you must build. Compare subscription, usage, credits, and hybrid models, and how to choose.
A tiered discount structure rewards bigger commitments with deeper discounts. Here's how to design one, and how to enforce the tiers before compute runs.
%20(1).png)
Your metering service shows a customer crossing 10,000 units at 2:17 p.m. Five requests are still in flight, each reading usage state near the same tier boundary. Finance expects the lower rate to apply cleanly. Engineering has to make it true.
A tiered discount structure can reward higher commitment and protect margin, but its thresholds, pricing math, and runtime behavior must agree. Here’s how to build one that holds up in production.
A tiered discount structure increases the discount at defined thresholds based on volume, spend, contract length, or another commitment metric.
A simple schedule might look like this:
The engineering questions start as soon as someone asks what happens to unit 5,001.
Does the 15% discount apply to that unit, every unit in the period, or the next invoice? If three requests cross the threshold together, which rate does each one receive?
The first design decision is what the tier measures. Units purchased, API calls, annual contract value, and prepaid credits behave differently near a threshold.
Your tiered pricing defines how units are priced. The discount structure defines how customer commitment changes the price.
A discount based on annual contract value can live in a sales workflow, but a discount tied to live usage needs current metering state and a precise boundary rule.
Tiered discount models differ in which units receive the lower rate after a customer crosses a threshold.
Two teams can both say “tiered pricing” and still calculate different invoices. The model name matters less than the boundary behavior specified.
Use 6,000 units across three bands:
Graduated pricing produces a $5,400 total because each band keeps its own rate.
An all-units volume model reprices all 6,000 units at $0.80, which results in a $4,800 total. The same usage and published tiers create a $600 difference.
Graduated tiers make the effective rate move smoothly, while volume tiers create a stronger incentive to cross the threshold, and the price cliff needs careful modeling before launch.
AWS S3 Standard storage is a familiar graduated example. The first 50 TB, the next 450 TB, and everything above 500 TB each price on their own band.
Tiered discounts affect margin by trading unit revenue for higher commitment or consumption.
The economics work when the added volume earns back the lower unit price. Trouble starts when the deepest discount goes to customers who also generate the highest delivery cost.
An API customer can double usage after reaching the top tier. Revenue rises, but every request may also trigger premium model calls, retrieval, storage, and third-party API fees. The discount decision now reaches the infrastructure bill.
Three checks help before the rate card goes live:
Pay extra attention to the top tier. A two-point discount looks small in a pricing review. Across millions of requests, those two points can become a large monthly number.
Engineering should also know the margin floor. If a custom rate can fall below it, the system needs an approval path or a rule that blocks the configuration.
A tiered discount structure works best when the metric, thresholds, discount depth, and runtime rules reflect real customer behavior. Work through the design in this order:
Start with the commitment metric. The metric tells customers which behavior earns a better price, and it tells engineering which state must stay current.
If customer value grows with API throughput, discount throughput. If enterprise buyers commit annual spend, contract value may be the cleaner choice.
A seat threshold on a consumption-heavy product can reward the wrong behavior. Customers may add seats while a small number of automated agents create most of the cost.
Pull the usage histogram before choosing 1,000, 5,000, and 10,000 because they look tidy in a deck.
Threshold placement changes customer behavior. A tier slightly above a meaningful customer cluster can encourage a larger commitment, but a tier far beyond normal usage may never influence a buying decision.
The histogram also shows engineering where boundary traffic will concentrate. If hundreds of accounts cross the same threshold near month-end, test that path under a burst.
Enterprise contracts add account hierarchies, custom periods, and negotiated allowances, and enterprise subscription management keeps those rules tied to the correct account state.
Write down what happens one unit below and one unit above every threshold.
Graduated tiers lower the rate for units inside the new band, while volume pricing may reprice the full quantity as soon as the threshold is reached.
The specification should answer a few direct questions:
If your own team calculates two totals from the same example, the rule needs another pass.
Run the proposed structure across real customer consumption before launch.
A pricing simulation shows which accounts would change tiers, how their blended rates move, and where margin becomes thin.
Average usage can hide the expensive cases. Replay end-of-month bursts, long agent runs, late events, and accounts using premium models, since those are the moments that put pressure on the pricing logic.
Localized list prices and global discount percentages can produce different effective margins across markets.
Cursor shows what regional pricing looks like once it leaves the spreadsheet. Its Start plan costs ₹649/month in India, billed in INR with tax included. Add a hypothetical 15% discount, and engineering must calculate ₹551.65 before applying the final rounding rule.
When you use price localization, store the regional base price and discount as separate versioned rules. This keeps the calculation traceable when currencies, tax treatment, or local prices change.
The pricing page gives customers the promise, but engineering needs the operational version.
Document mid-period crossings, concurrent requests, resets, refunds, plan changes, and late-arriving usage. Include the effective timestamp and pricing version in the same specification.
For live consumption, the system needs one defined moment when the new tier takes effect. Leaving that decision to request ordering means you get inconsistent bills.
Tiered discounts for credits, tokens, and usage connect the discount directly to variable product consumption.
A customer may buy a larger credit pack at a lower effective rate, and another contract may apply a discount after monthly token usage crosses a threshold. Promotional credits can also deliver the discount as an extra balance with its own expiry date.
A customer buys 100,000 paid credits and receives 15,000 promotional credits. The promotion expires first, but the wallet burns paid credits first. A month later, support has to explain why the bonus disappeared untouched.
Once discounts are expressed as credits, the runtime needs clear rules for:
Several requests can cross a tier boundary together. The request path needs current usage, a consistent tier decision, and an atomic debit or reservation before expensive work begins.
Tier thresholds are enforced by keeping usage current, defining the exact crossing point, and sending the new tier state to every service that needs it.
At 11:42 a.m., an account sits at 99,998 API calls. Three requests arrive together, and one earlier request is retried. Within milliseconds, the account may cross the threshold more than once unless metering, rating, and cached state follow the same rules.
Every usage event needs a stable ID, event time, billing period, and account owner. Those fields tell the system whether the event is new, when it occurred, and which customer’s total should change.
A duplicate near the boundary can move an account into the next tier early. The customer sees a strange rate change, while engineering has to trace the error back through ingestion and aggregation.
Late usage creates another wrinkle. An event received today may belong to yesterday’s billing period, where the customer was still one request below the threshold.
Once the meter confirms call 100,000, the rating service still needs a precise rule. Does the lower rate begin with call 100,000 or 100,001? Does it apply to earlier usage in the period? Which timestamp decides?
Graduated pricing can begin the next band at the threshold. An all-units model may need to reprice earlier usage from a consistent snapshot while new events continue arriving.
Each approach can work when the implementation matches the pricing promise. The event should carry the tier, rate, pricing version, and effective time used for the decision.
After the boundary changes, every node handling protected traffic needs the updated state. One stale cache can keep applying the earlier rate or allowance while the rest of the product uses the new tier.
The cache design should define:
A discount calculated on the next invoice can tolerate some delay, but a tier that changes credits, limits, or access needs fresher state before the next request runs.
This is the 2:17 p.m. crossing from the top of this article, the moment tier design meets live traffic and everything either lines up or leaks money.
Tiered discount structures break most often where usage crosses a boundary under live traffic.
A retry, late event, concurrent request, or stale cache can apply the wrong tier even when each service processes its own input correctly. Production needs one answer for which event crosses the threshold and when the new rate begins.
One test catches a surprising number of problems. Put an account one request below a threshold and send a burst of concurrent traffic.
Then inspect the tier, rate, and pricing version attached to every event. The result should stay consistent across repeated runs, even when request order changes.
If the total moves between runs, the pricing rule still has a concurrency problem.
Runtime controls support a tiered discount structure when tier position changes what the product can charge, allow, or consume during live traffic.
Billing keeps owning invoices and financial records. Stigg is the usage runtime for AI products, which means it enforces entitlements, meters usage, and governs AI spend in the request path.
Miro reported saving 5,000 engineering hours after moving credit and entitlement handling out of their own codebase and into a dedicated runtime. That’s the kind of number that only shows up once tier enforcement stops being a rolling engineering project.
For tiered discount structures, Stigg adds:
You can adopt each component as the need appears. Metering may come first, followed by credits or entitlement checks as the pricing model grows. The result is one consistent commercial decision across the live product and the final invoice.
Follow the full production path in the Stigg docs, from metering the first event to checking entitlements and debiting credits in real time.
A tiered discount structure lowers the effective price when a customer reaches set thresholds for usage, spend, units, or contract term. The pricing rule should define which units receive the discount and when the new rate begins.
The main difference between tiered and volume discounts is which units receive the lower rate. Graduated tiers price each band separately, while volume discounts apply the reached rate to the full quantity.
Choose tier thresholds using real customer distribution, marginal cost, and the behavior you want to reward. Test usage immediately below and above each boundary to catch price cliffs and weak margins.
Tiered discounts for usage-based or AI products are enforced by evaluating current cumulative usage against tier thresholds in the request path, synchronously, before the rate is applied or the credit is burned. Enforcing after the fact lets you observe overages without preventing mis-tiered charges.
Yes. Tiered discounts can hurt margins when the lower unit rate approaches the cost of serving heavy usage. Replay historical consumption and include model, compute, storage, and third-party API costs before launch.