Blog
/
Industry Insights

Inference gets 10x cheaper every year. Your price sheet updates once.

Inference costs fall about 10x a year while agent usage climbs. Why static price sheets leak margin, and what fast repricing takes to build.

Sagi Bracha Sagi Bracha
Written by
Sagi Bracha
Last updated
September 28, 2026
read time
5
minutes
Price sheet next to inference cost falling from $1.00 to $0.001 between 2023 and 2026

Table of contents

For a given level of capability, the cost of inference has been falling about 10x a year. A task that cost a dollar last year runs for cents now. If you sell an AI product, falling input costs look like free margin. They aren't. A price you set once and leave alone can't track economics moving at that speed. This shows up on the cloud bill you got this month, not in a forecast. A price set today is calibrated to today's cost basis and today's usage patterns. Both move within months. Your price sheet doesn't, and it breaks in two directions. Costs fall. Your cost to serve drops, your price stays put, and a competitor who updates their pricing undercuts you for no reason a customer can see. Usage climbs. Your customers, and increasingly their agents, burn far more compute than you budgeted for. Cost to serve passes your flat price, unit economics invert, and a deal you booked as profitable turns into a margin leak. Either way, you find out a quarter late.

The unit you bill stopped matching the work you do

SaaS pricing assumed one request in, one response out. One clean unit to charge for. An agent handling a workflow pulls context, reasons over it, calls external tools, validates its output, and retries on failure. One user intention triggers twenty model calls. You still bill for the one. Static pricing counts requests. It has no view of the compute each request sets off. Every one of those calls is a spending decision your system is making on a customer's behalf, and right now most systems make it blind.

Prices decay from the day you set them

The gap doesn't hold still. Work per request grows, usage patterns shift monthly, and infrastructure costs move in both directions: model efficiency improves while GPU rental rates spike on supply. Static pricing loses margin when costs drop and takes unit-economic damage when they spike. The drift is continuous, so no single correction catches you up. Every day, your price sheet is a little more wrong than it was. That makes repricing speed an operational capability. If changing a price takes two quarters of engineering cycles, you're permanently priced for a market that's already gone. If it takes an afternoon, you never fall behind.

Fast repricing is an architecture decision

Whether you can reprice in a day has nothing to do with the size of your finance team or the quality of its spreadsheets. It comes down to where pricing logic lives. In most companies it lives in the application. Plans, limits, meters, and feature flags sit in the main codebase, so changing a price takes a ticket, a code review, regression tests, and a release. Pricing moves at the speed of deploys, which means it barely moves.

The alternative is pricing as configuration. The commercial catalog sits outside the application, and the product asks an entitlement layer at runtime what this customer is allowed to do and how they're charged. Costs drop or agent traffic climbs, and a price change becomes a config update with no downtime, no ticket, no deploy. Someone makes a commercial call in the morning and it's live in production by the afternoon.

This is the point where teams reach for feature flags, because the shape looks familiar. A flag answers whether code is on. An entitlement answers whether this customer, on this plan, with this much budget left, gets this call right now. The second question has state and money behind it, and a flag system has nowhere to put either.

Knowing what to charge is a different problem than changing it fast

Pulling pricing out of the code solves half of it. You can move the number now. You still have to know which number, and against agent traffic that's harder than it sounds. Pricing on usage means measuring usage as it happens: a real count of what each customer consumed, at the unit you charge on. A loose meter makes every price built on top of it a guess, and meters that were fine at request volume start dropping events when each request fans out twenty ways.

Then there's the unit itself. Billing customers in raw tokens ties your price sheet to a model provider's SKU, so every model swap and every provider price change reprices your product for you. An abstraction in between, credits or tasks or whatever fits your product, decouples what you charge from what you spend and gives you somewhere to absorb the volatility. It also gives customers a number they can reason about, which raw token counts never were. Then limits. A customer on a capped plan, or an agent burning through its budget, has to be stopped in the moment. Find out in the monthly rollup and one runaway agent has already spent your margin, and there's nothing to claw back.

These are two different jobs, and the split is where the commercial question turns into an engineering one. Metering is measurement. It counts what each customer consumed, it can run alongside the request, and you can settle it after the fact. Enforcement is a decision on the request path: before the call runs, the system checks whether the customer is still inside their limits and either passes it or stops it. That answer has to come back while the request waits.

What this costs to build

Deciding to price on real usage signs your product up to meter and enforce every request as it happens, at whatever volume your customers and their agents produce. That's the half teams underestimate. Metering has to stay accurate under fan-out, which means an ingestion path that survives bursts and a ledger you can reconcile when finance asks why a number moved. Enforcement sits on the hot path, so its latency budget is small and the number that matters is p99, not the average: the check runs on every call an agent makes, and an agent makes a lot of calls. And balances have to be right under concurrency: parallel agent calls drawing on one budget is a double-spend problem, and it fails the same way. The pricing layer is what your team sees and touches. The layer underneath is what has to hold up in production, and if it doesn't, the fast pricing on top of it doesn't matter. Inference costs will keep falling and agent traffic will keep climbing whether or not your price sheet moves. Two things are worth building for that: a price you can change the day your costs move, and a meter underneath it you can trust.

Latest news.

One email per month.
From engineers, for engineers.

Thank you! Your submission has been received.
Oops! Something went wrong while submitting the form.