Blog
/
Guides

Metered Usage: How It Works and Why AI Products Break It

Metered usage is the measurement layer under every usage-based bill. See how it works for AI products and why token consumption breaks seat-based pricing.

Sara NelissenSara Nelissen
Written by
Sara Nelissen
Last updated
October 9, 2026
read time
8
minutes
Metered Usage: How It Works and Why AI Products Break It

Table of contents

Your meter counts a support agent's conversation as one event, yet that conversation fired 30 model calls, 3 retrieval passes, and a retry loop, and none of it reaches an invoice until the cycle closes.

Metered usage is the measurement layer under every usage-based bill, recording what each customer consumed in tokens, API calls, or credits. If you build AI products, the harder question is where that record ends, and enforcement begins.

What metered usage is

Metered usage is the practice of measuring and recording how much of a product a customer consumes, so you can price on it. The unit is whatever maps to cost and value, such as gigabytes stored, API calls, or minutes of compute.

AI products use more specific units, including input and output tokens, model calls, inference seconds, and credits drawn from a shared balance.

The metered unit sets how honest your pricing stays as usage grows. When you pick a unit that tracks real cost, pricing holds up, but if you pick a vanity metric, customers overpay or learn to game it.

Usage metering in general follows the same idea, applied here to products whose cost per request varies. A seat-based plan charges the same whether a user runs one query or ten thousand, and metered usage ties the invoice to the work the system did.

Metered usage vs metered billing vs usage-based pricing

Three terms get used interchangeably, and the confusion costs you time when you choose a billing tool or debug an invoice.

Term What it decides Who owns it
Metered usage How consumption is measured and recorded Engineering
Metered billing How measured usage becomes charges Billing system and finance
Usage-based pricing Whether you charge by consumption at all Product and engineering

When a usage-based invoice comes out wrong, look at the metering layer before the pricing config. Dropped events, double counts, and late data all show up as billing errors while starting upstream at the meter.

If you treat the terms as one thing, you spend hours auditing prices while the fault sits in the pipeline.

How metered usage works, end to end

Behind every usage-based bill there’s a pipeline with five stages: capture, aggregate, rate, reconcile, and expose.

Component What it does Failure mode if missing
Event capture Records each usage event (a token count, an API call) at the source Usage goes unbilled
Aggregation window Rolls raw events into billing periods and deduplicates them Double counts or missing events corrupt every invoice
Rating engine Applies pricing rules (per-unit, tiered, volume) to aggregated usage Charges don't match the contract
Ledger Keeps an auditable, append-only record of what was consumed and billed Finance can't reconcile and disputes have no receipts
Customer visibility Shows live usage and alerts before thresholds hit Bill shock and a queue of support tickets

Each stage has to guarantee one property. Capture has to be idempotent because usage pipelines retry and the same event will arrive twice. Aggregation has to be exact, since even a small event drop under load breaks any tiered or high-watermark calculation.

Stripe's meter events API shows the constraint. It enforces identifier uniqueness within a rolling 24-hour window and accepts timestamps from the past 35 calendar days (up to 5 minutes in the future), so anything outside those windows is yours to handle.

For AI products, the pressure lands on speed. Spikes happen in minutes, and a nightly batch job reports the overspend after the compute is spent, so real-time metering lets you catch a runaway agent while it runs.

Why token usage breaks old pricing models

Token usage breaks seat-based and flat pricing because consumption is unbounded and varies per request, while those models assume one user, one seat, and a steady amount of work.

A short question and a 200-page document analysis both count as one use, yet their token cost can differ by orders of magnitude. With agents, a single user action fans out into dozens of model calls.

Every model call also carries a marginal cost in tokens, so a flat price absorbs whatever a heavy user consumes and light users pay for capacity they never touch.

This is why many AI products moved to credits and usage-based pricing.

A 2025 survey of 240 software and AI companies, cited in HubSpot's guide to credit-based AI pricing, found seat-based pricing falling from 21% to 15% in 12 months and hybrid pricing rising from 27% to 41%.

What makes metering AI usage harder than SaaS metering

Metering AI usage is harder than metering SaaS usage because it adds problems that counting gigabytes never had. The four below show up first:

  • Non-determinism shows up because token counts vary per request and vary again when you swap the underlying model, so last month's rate card can misprice this month's traffic.
  • Multi-dimensional units mean input tokens, output tokens, cached reads, tool calls, and retrieval steps each carry different cost, and the meter rolls them into one number a customer understands.
  • A low-latency requirement applies because the enforcement decision sits in the request path and cannot add meaningful latency to an inference call. Metering records what happened and enforcement decides what runs.
  • The credit abstraction hides this complexity behind one balance that customers spend across features at different rates.

A credit system that holds up in production issues credits in blocks with expiry dates and a cost basis, tags grants as paid or promotional, and sets a burn order.

It also applies hard or soft depletion and records every debit in an append-only ledger for reconciliation. A running total covers none of that.

Where metering ends and enforcement begins

Metering ends at the record of what happened, and enforcement begins at the request, where a synchronous decision allows or denies the call before compute runs.

Component What it does Failure mode if missing
Event capture Records each usage event (a token count, an API call) at the source Usage goes unbilled
Aggregation window Rolls raw events into billing periods and deduplicates them Double counts or missing events corrupt every invoice
Rating engine Applies pricing rules (per-unit, tiered, volume) to aggregated usage Charges don't match the contract
Ledger Keeps an auditable, append-only record of what was consumed and billed Finance can't reconcile and disputes have no receipts
Customer visibility Shows live usage and alerts before thresholds hit Bill shock and a queue of support tickets

A nightly billing run flags an overage after the compute is spent, so only enforcement can stop a spike. Enforcement applies one of two limits:

  • A hard limit denies the request when a balance is gone.
  • A soft limit lets the request through into a negative balance and reconciles later.

The decision needs low latency, so it sits in the request path behind a local cache of entitlement data. On a cache hit, entitlement checks resolve instantly from local Redis. On a cache miss, the Sidecar fetches from Stigg's Edge API at around 100ms, with a configurable timeout.

The Sidecar runs as a Docker container in your own cloud (BYOC), and persistent caching keeps reads available if the upstream service becomes unreachable, which covers reliability.

The metered ledger stays the system of record, and enforcement acts on it in time to be useful. Building this yourself means starting from limits, allocations, and budgets and working back to the runtime.

How entitlements turn metered usage into a decision

Entitlements turn metered usage into a decision by comparing a customer's running consumption against the limit their plan allows, then returning an outcome your code can enforce.

An entitlement is a commercial allowance for a feature, and it carries a limit plus a running count of consumption. A Boolean on/off flag carries neither, which is why the two behave so differently at runtime.

Three neighboring concepts get mixed up, so the table below separates them by the question each one answers.

Concept Question it answers Example
RBAC (role-based access control) Who inside the organization may use this feature? An admin can open the billing settings
Entitlements How much of this feature does the plan include? The Pro plan includes 10,000 tokens a month
Billing What does the usage cost after the fact? The invoice lists 12,400 tokens at the plan rate

A customer whose plan allows 10,000 tokens, whose active trial allows 20,000, and whose promotional override allows 15,000 resolves to 20,000, because the most generous value wins. The same resolution also reads the parent plan the customer inherits from and any add-ons.

The resolved result reaches your code as a small object, shown here in illustrative form:

json

{

  "hasAccess": true,

  "usageLimit": 20000,

  "currentUsage": 13250,

  "isUnlimited": false

}

One lookup gives your code the access decision, the cap, the running total, and whether a cap exists at all, so it never queries plan, usage, and overrides separately.

From that object, your code picks the outcome. It can block the request under a hard limit, let it overflow under a soft limit, or render an upgrade prompt that shows the customer a locked feature they can buy.

The engineering failure modes of AI metering

The engineering failure modes of AI metering are non-idempotent capture, dropped events, ledger drift, stale rate cards, and meter-dashboard mismatches, and production load is where all five appear.

The table below pairs each failure with the symptom you would see and the control that prevents it:

Failure mode Symptom What prevents it
Non-idempotent capture The same event counts twice, and customers get double-charged Unique event IDs and idempotent deduction
Dropped events under load Aggregates and high-watermark tiers come out wrong every cycle Back-pressure handling and reconciliation against source
Ledger vs finance drift Usage numbers don't match what finance recognizes An append-only ledger with real-time deductions as the single source of truth
Model price change mid-period Old usage gets repriced at the new rate Versioned rate cards tied to each event's timestamp
Meter and dashboard disagree Customers see one number, get billed another, and stop trusting you One usage source feeding both the invoice and the UI

The connective tissue is concurrency. Under real load, several requests read the same balance before any write lands, all pass the check, and the balance goes negative after the compute is spent.

Idempotency, consistency, and enforcement timing are one problem wearing different hats, which is why patching them one symptom at a time rarely holds.

The fix needs no exotic machinery, only an append-only ledger, idempotent capture, and a decision point in the request path, all in place before traffic gets real.

How to set up metered usage the right way

Setting up metered usage the right way comes down to six decisions, and you'll feel each one later if you skip it now.

  1. Pick the unit first: Choose something tied to real cost that your customer can predict. If you can't guess what a single action costs them, they can't either, and the metric is wrong.
  2. Make capture idempotent: Assume every event shows up twice, because in production it will. Design the deduction so a repeat changes nothing.
  3. Aggregate in real time: Batch is fine for your reporting dashboard, but enforcement needs numbers from this minute, since AI spikes don't wait for the nightly job.
  4. Keep an auditable ledger: Use an append-only record with cost basis per grant, so when finance asks where a number came from, you can answer without digging through spreadsheets.
  5. Decide limits up front: Pick hard or soft behavior per feature and enforce it in the request path. Overage handling is a pricing call and an engineering call, so you'll want both sides in the room.
  6. Show usage live: Dashboards and threshold alerts close most bill-shock tickets before they open, which saves you the support thread.

Design the pricing model first, then confirm your infrastructure can run it. A high-watermark or rolling-window model reads clean on a contract and needs a mature pipeline underneath.

Before you commit to the commercial shape, run your real event volume through the meter and see what breaks. It's cheaper to find out in a test than in an invoice dispute.

What metering leaves unsolved

Metered usage leaves the request-time decision unsolved, since a record of what a customer consumed cannot tell you whether the next call should run.

Stigg is the usage runtime for AI products, which means entitlements, credits, usage limits, and spend governance get enforced synchronously in the request path.

Your billing keeps doing its job, whether that's Stripe, Zuora, or a custom system, and Stigg holds steady performance at scale across large, complex setups. Startups often begin with a single SDK integration.

Each piece plugs in on its own terms:

  • Runaway agents stop at the request: The Sidecar runs as a Docker container beside your app and caches entitlement data in Redis. Checks resolve instantly on a cache hit, and a miss fetches from Stigg's Edge API at around 100ms with a configurable timeout.
  • Usage data stays in your cloud: BYOC (Bring Your Own Cloud) runs the same runtime inside your VPC, which covers data residency requirements.
  • Credits behave like a ledger: The credits engine handles block-level expiry, cost basis, paid and promotional categories, configurable burn order, hard or soft depletion, and an append-only record for reconciliation.
  • Every meter reading maps to a plan: Entitlements resolve limits across plans, add-ons, trials, and promotional grants, with the most generous value winning on conflict.
  • Budgets follow the org chart: Limits apply per agent, user, team, product, or department.
  • Pricing changes skip the deploy: Packaging moves through configuration, while engineering owns pricing-model changes.
  • Adoption happens in slices: Metering, entitlements, or the credits engine each run on their own, without touching your billing stack.

Your meter already records what happened. Stigg's docs walk through how the Sidecar, entitlement layer, and credits ledger decide what happens next in your own stack.

Frequently Asked Questions

1. What is metered usage?

Metered usage is the practice of measuring and recording how much of a product a customer consumes so you can bill on it. For AI products, the measured unit is tokens, model calls, or credits rather than seats or a flat fee.

2. What is the difference between metered usage and metered billing?

The main difference between metered usage and metered billing is that metered usage measures and records consumption, and metered billing applies prices to that measurement to produce an invoice. Accurate metered usage has to come first for metered billing to be correct.

3. Is metered usage the same as usage-based pricing?

No, metered usage is the measurement mechanism, and usage-based pricing (also called consumption-based pricing) is the commercial model built on top of it. You can't run usage-based pricing reliably without metered usage underneath.

4. How do you meter token usage for an AI product?

You meter token usage for an AI product by capturing each request's token and model-call events idempotently at the source, aggregating them in real time, and mapping them to a credit or dollar cost. Enforcement then checks limits before the next call runs.

5. How do you prevent overspend with metered usage?

You prevent overspend with metered usage by enforcing limits in the request path before each model call runs. Hard limits block usage when a balance is gone, soft limits allow a controlled negative balance, and threshold alerts warn customers before they hit a wall.

Latest news.

One email per month.
From engineers, for engineers.

Thank you! Your submission has been received.
Oops! Something went wrong while submitting the form.