Blog
/
Guides

Metering and Billing: What They Are and How They Work

Follow the metering and billing path from usage event to invoice, including ingestion, aggregation, rating, reconciliation, credits, and enforcement.

Sara NelissenSara Nelissen
Written by
Sara Nelissen
Last updated
October 7, 2026
read time
10
minutes
Metering and Billing: What They Are and How They Work

Table of contents

A customer opens a billing ticket with a simple question. “Why did we get charged for this twice?” Engineering traces event IDs, finance checks the invoice, and nobody knows where the bad number entered the system.

Metering and billing connect product usage to what the customer pays. The interesting part lives between those endpoints, across ingestion, aggregation, rating, reconciliation, and the runtime controls that act before more AI compute runs.

What metering and billing mean

Metering measures product consumption. Billing applies commercial rules to that measured usage and creates the financial record.

Metering starts inside the product. Each billable action emits an event tied to a customer, feature, quantity, timestamp, and other dimensions needed downstream.

Those events might represent:

  • Tokens consumed by a model
  • API requests against a billable endpoint
  • Agent actions completed inside a workflow
  • Compute time used by a workload
  • Credits consumed from a customer balance
  • Storage or other resources tied to product usage

AWS’s SaaS architecture guidance separates metering, operational metrics, and billing into distinct concerns, which is a useful separation when assigning ownership across services.

Billing picks up from the measured quantity. Rating applies the relevant price or rate card, then invoicing turns rated charges into the amount the customer owes.

The metered billing layer covers charging based on measured consumption. The broader metering and billing system also includes the event path that produces that measurement.

How metering and billing work end to end

Metering and billing work as a pipeline that turns raw product activity into a rated charge, invoice, and reconciled financial record.

Stage State produced Common failure
Event ingestion Durable usage events Missing or duplicate events
Aggregation Billable quantity Wrong window, dimension, or function
Rating Priced usage Wrong rate, tier, or contract term
Invoicing Financial record Invoice state diverges from rated usage
Reconciliation Verified usage-to-charge trail Errors survive into close or disputes

The important detail is state ownership between stages. Each layer should consume a defined output from the previous one without rebuilding its logic.

A duplicate event can pass cleanly through aggregation, rating, and invoicing. Every downstream calculation may be correct while the final bill is still wrong.

Event ingestion

Event ingestion captures each billable action as a durable, uniquely identifiable usage event before aggregation begins.

The event envelope needs enough context to support attribution, deduplication, replay, and downstream grouping. A typical event might carry:

  • tenant_id for the customer or account
  • event_id as a stable unique identifier
  • meter_id for the measured feature or resource
  • quantity for the consumed amount
  • timestamp for window assignment
  • model, region, or workload metadata when relevant
  • idempotency_key when the producer can retry

The distinction between event_id and business identity is important. Two valid requests may look identical while representing separate consumption. A retry of one request should still collapse to one event.

If your transport uses at-least-once delivery, duplicate delivery is part of the design space. Deduplication needs to happen before a retry can inflate the billable quantity, and replay deserves the same treatment.

If ingestion pauses, durable source events should let you rebuild the missing window without creating a second copy of usage already accepted.

Aggregation

Aggregation converts raw events into the exact quantity rating will price.

The aggregation rule depends on what the meter represents:

Meter type Aggregation Example use
Tokens Sum Total model input and output
API calls Count Requests to a billable endpoint
Active users Count unique Monthly active usage
Peak concurrency Maximum Highest simultaneous workload
Stored state Last value End-of-period recorded quantity

The grouping key matters as much as the function. A token meter might aggregate by:

tenant_id + model + billing_period

Add workspace, region, feature, or agent and the cardinality grows quickly. Every extra dimension gives you more attribution detail and more state to query, store, and reconcile.

Window semantics also need one owner. UTC boundaries, customer-local billing periods, late events, and backfills can all move otherwise valid usage into a different period.

Late-event policy should be explicit. You need to know whether a late event updates the closed period, creates an adjustment, or moves into the next billing cycle.

Rating

Rating turns an aggregated quantity into priced usage under the commercial terms that applied when the usage occurred. A typical path looks like this:

measured quantity → included allowance → applicable tier → contract override → overage rule → charge

The rating layer may need to resolve:

  • Per-unit rates
  • Tiered or volume pricing
  • Included allowances
  • Overage rates
  • Minimum commitments
  • Customer-specific terms
  • Effective dates

The architecture of a billing system gets harder to reason about once multiple versions of those rules can apply to the same customer over time.

Pricing configuration therefore needs versioning.

Usage generated on August 31 should resolve against the terms active on August 31, even if the customer upgrades on September 1.

A useful rating record should preserve the quantity, pricing version, rate, and resulting charge. That gives reconciliation something concrete to compare later.

Invoicing and payment

Invoicing converts rated usage into the financial record the customer receives and accounting systems recognize.

By this point, the meter has already answered how much was consumed, and rating has answered what that usage costs.

The invoice layer still has several pieces of state to preserve:

  • Invoice line item linked to the rated usage
  • Credits or adjustments applied after rating
  • Tax treatment
  • Currency
  • Billing period
  • Payment status
  • Contract references needed for audit or dispute handling

The useful debugging direction runs backward.

An engineer should be able to start from an invoice line and trace it to the rated charge, aggregated quantity, and source usage events.

Without that chain, a customer dispute turns into a search across logs, billing exports, and ad hoc SQL.

Reconciliation

Reconciliation verifies whether the usage that entered the pipeline matches the usage that became a charge. This is where the pipeline proves its own work. A useful reconciliation process should answer:

  • Which source events produced this aggregate?
  • Which aggregate produced this rated quantity?
  • Which pricing version and rate applied?
  • Which invoice line contains the charge?
  • Which events arrived after the billing window?
  • Which events were replayed or deduplicated?
  • Which adjustments changed the final amount?

Reconciliation becomes much more useful when mismatches identify a stage.

Mismatch Likely place to inspect
Source events > aggregate Ingestion, deduplication, or aggregation
Aggregate > rated quantity Allowance or rating rules
Rated charge ≠ invoice line Invoicing or adjustment logic
Closed period changes later Late-event or backfill policy
Same event appears twice Event identity or deduplication

“Usage and billing differ” gives you an alert. But “Seventeen late events entered after the August window closed” gives an engineer a starting point. That traceability is what turns metering and billing from a chain of counters into an operable production system.

The building blocks behind reliable metering and billing

Reliable metering and billing depends on stable event identity, durable ingestion, consistent aggregation rules, versioned pricing configuration, and reconciliation. The core pieces include:

  • Stable event identity for retries, deduplication, replay, and tracing
  • Durable ingestion before downstream processing begins
  • Aggregation rules for windows, functions, dimensions, and time zones
  • Versioned rate configuration with effective dates
  • Customer-visible usage for current consumption and thresholds
  • Reconciliation across source events, aggregates, charges, and invoices
  • Observability for late events, duplicate rates, ingestion errors, and processing latency

The usage metering layer owns the measurement side of this path. Billing should be able to consume that state without reconstructing product activity from application logs or analytics tables.

How metering and billing change for AI products

Metering and billing for AI products gets harder because one customer action can create several different kinds of usage across models, tools, and execution paths.

An agent request might call one model, query a vector store, invoke an external tool, then call another model before returning a result. The meter has to preserve enough context to connect those events to the same customer and workload.

The next choice is the commercial unit. Different units expose different parts of the workload:

  • Tokens capture model input and output with fine granularity.
  • Requests work well when endpoints have similar resource profiles.
  • Compute time fits workloads where runtime tracks resource consumption closely.
  • Agent actions package several internal steps into one measurable unit.
  • Credits give mixed workloads one customer-facing balance with different burn rates underneath.

Tokens are useful when model consumption itself drives the commercial rule, but their economics still vary by model, context length, caching behavior, and workload.

The AI token cost model becomes useful when those differences need to feed usage attribution, internal costing, or customer-facing pricing.

Requests need more care once workload size varies. A short extraction and a multi-step agent run may both count as one request while consuming very different resources.

Credits solve a different part of the problem, by letting several workload types draw from one balance while the metering layer keeps the lower-level events available underneath.

A production credit system therefore needs richer state than a single balance field:

  • Block-level expiry for individual grants
  • Cost basis attached to each block
  • Paid and promotional categories
  • Configurable burn order
  • Hard or soft depletion rules
  • Append-only ledger entries for every balance change

The usage store and credit ledger should have separate jobs. The usage store records what was consumed, while the credit ledger records how grants, debits, expirations, refunds, and adjustments changed the available balance.

Keeping those states distinct makes attribution, reconciliation, and request-time enforcement easier to reason about as AI workloads become more complex.

How pricing models change the metering path

Pricing models change the metering path by adding different state and decision rules to the same usage events.

Pay-as-you-go

Pay-as-you-go needs a measured quantity, an applicable rate, and a reproducible path from event to charge. You can rate usage continuously or aggregate first and rate later, but both approaches need stable events, deterministic aggregation, and versioned rates.

Included usage with overage

Included usage adds allowance state and an overage boundary. The system needs to track:

  • Included allowance
  • Reset period
  • Eligible usage
  • Current consumption
  • Overage rate

The usage-based billing layer then prices consumption beyond the included amount. Late events matter here because they can push an account across the allowance boundary after a period has been calculated.

Subscription plus credits

Subscription plus credits adds a consumable balance with its own lifecycle.

Each usage event can affect the credit deduction, remaining balance, burn order, depletion rule, and billing state.

Shared balances also introduce concurrency. Several requests can draw from the same balance at once, which calls for atomic debits, idempotency, and ledger-backed credit state.

The pricing model changes how much state the metering path has to carry forward before billing or runtime enforcement can act.

Where metering and billing break

Metering and billing break when different stages disagree about event identity, time, quantity, pricing configuration, or financial state.

Failure mode Root cause Architecture response
Duplicate usage Retries arrive without deduplication Stable event IDs and dedup before aggregation
Missing usage Events disappear before durable capture Durable ingestion, replay, and backfill
Aggregation mismatch Services use different windows or dimensions Shared definitions for units, windows, and time zones
Wrong charge Rating reads stale pricing Versioned rate configuration and effective dates
Late usage Events arrive after a period closes Late-event policy and reconciliation
Reconciliation mismatch Usage and billing are never compared Standing reconciliation across both systems

Duplicate usage is easy to miss because every downstream step can still behave correctly. Aggregation, rating, and invoicing simply preserve the wrong input.

When it comes to late events, you need a clear policy for whether they reopen the period, create an adjustment, or move into the next billing cycle.

Observability should surface both problems early. Useful signals include:

  • Duplicate-event rate
  • Late-event volume
  • Ingestion failures
  • Processing latency
  • Reconciliation drift
  • Unrated usage
  • Missing tenant attribution

Those metrics give you a chance to catch a bad usage trail while it is still an engineering problem, before it turns into a billing dispute.

Where metering and billing stop

Metering and billing stop at measurement and financial processing. Request-time enforcement handles the separate decision of whether the next protected workload can run.

For AI products, timing matters. Usage can keep accumulating while earlier events are still moving through aggregation, rating, and billing.

Concurrency adds another failure mode. Several requests can read the same allowance before any debit commits, which makes stale state look valid to every request.

A production request path needs:

  • Atomic debits against shared balances
  • Idempotency across repeated requests
  • Current entitlement state at decision time
  • Cache invalidation after package changes
  • Tenant isolation across shared infrastructure
  • Fallback behavior during dependency failure

Hard and soft limits define what happens when the allowance is exhausted. A hard limit fails closed and denies further usage. A soft limit allows the request, flags it as overage, and lets you track and bill it instead of blocking it.

The AI billing infrastructure layer connects those runtime decisions with the metering and financial systems downstream.

Once usage needs to affect the next request, the problem has moved beyond counting and charging into live state and enforcement.

What to do once you know the usage

Knowing that an account has 120 credits left is useful. However, deciding which of several concurrent requests gets to spend them is the harder infrastructure problem.

That decision needs current metering state, entitlement rules, and credit state in one request-time path. Stigg handles that runtime layer, enforcing entitlements, credits, usage limits, and spend governance synchronously in the request path.

That runtime holds p99 entitlement checks under 10ms while ingesting 1M+ events per second on BYOC across 1M+ entities per hierarchy root, with multi-region active-passive failover and up to 99.99% uptime. This is the scale budget most in-house builds run out of before the second product line ships.

For a metering and billing stack, the relevant pieces include:

  • Usage metering records consumption by customer, feature, product, or workload and supplies the usage state billing reads downstream.
  • Hierarchical tenancy: every usage event updates the full chain (org → department → team → user → individual agent) simultaneously, with dimension-scoped sub-budgets at each level, so a single meter can support per-agent spend caps, per-team monthly allowances, and org-level governance from the same event stream.
  • Entitlements resolve the allowances attached to plans, add-ons, trials, and promotional overrides before usage is permitted.
  • Credits engine tracks grants and debits with expiry, cost basis, burn order, paid and promotional categories, and hard or soft depletion.
  • Stigg Sidecar runs beside your application and caches entitlement data in-memory by default, with optional Redis for serverless runtimes or large container fleets that need entitlements to survive restarts and stay shared across instances.
  • Node.js applications skip the Sidecar entirely. The Node SDK provides the same local caching, real-time updates, and low-latency checks natively in-process.
  • Cache hits resolve instantly, misses reach Stigg's Edge API at around 100ms, and a configurable timeout keeps upstream latency from cascading into your application.
  • BYOC places the runtime in your own cloud account on AWS, GCP, or Azure, managed via Infrastructure-as-Code, so end-user and usage data never leaves your network perimeter. It’s the same product, not a stripped-down SKU, often a procurement requirement for regulated-industry enterprise.
  • Billing integrations pass settled usage and commercial state to the systems already handling invoices, payments, tax, and financial records.
  • Modular adoption lets you use metering, credits, or entitlements independently. You can meter usage without entitlements, add credits without changing billing, or start with one SDK integration and add more components later.

A useful architecture check is to follow one live request. Find where usage is measured, where the allowance lives, where credits are debited, and where execution is finally allowed or blocked.

The Stigg docs map metering, Sidecar checks, credits, entitlements, and BYOC onto that production request path.

Frequently Asked Questions

1. What is the difference between metering and billing?

The main difference between metering and billing is measurement versus charging. Metering captures and aggregates usage into a billable quantity, while billing applies pricing to that quantity and creates the customer’s financial charge.

2. Is metering and billing the same as metered billing?

No. Metering and billing describe the full path from usage event to financial record, while metered billing charges customers based on the usage measured by that path.

3. How do you meter usage for an AI product?

You meter usage for an AI product by tracking a defined unit such as tokens, requests, agent actions, compute time, or credits. Stable event IDs, attribution, and deduplication keep retries and concurrent activity from distorting the total.

4. What causes incorrect usage-based bills?

Incorrect usage-based bills often come from duplicate or missing events, inconsistent aggregation windows, stale pricing rules, late usage, or reconciliation errors. Durable ingestion, stable event identity, consistent windows, and reconciliation help catch those failures.

5. Can metering and billing prevent overages before they happen?

No. Metering and billing record consumption and turn it into a charge after usage occurs. Preventing further usage requires request-time enforcement against the current entitlement, credit balance, or usage limit before protected compute begins.

Latest news.

One email per month.
From engineers, for engineers.

Thank you! Your submission has been received.
Oops! Something went wrong while submitting the form.