Take a pricing change that sounds trivial. An API company notices that a third of its traffic now hits a cache, so those calls cost almost nothing to serve. The sensible move is to price cached calls lower and leave the headline rate alone. Everyone agrees in the meeting. It's a one-line decision.
Then it meets the stack. The gateway team has to split a route, or read a header, so cached calls can be counted on their own. The application team has to add a field to the usage event. Someone creates a new price in the billing system, and a new version of every plan that uses it. Sales wants to know what happens to customers on annual commits, and whether renewals already in flight get the new rate. Finance asks whether cutting the rate on existing commits counts as a contract modification, and what it does to the breakage estimate now that those commits will burn down more slowly. If the change ships inside a quarter, people are pleased.
Since leaving Google in 2024, I've spent most of my time asking API and AI teams how they change their prices, and building what I wished they had. The details vary, but the sequence hardly ever does.
I run a company that sells into this problem, so weigh what follows accordingly. My argument is that usage-based billing and the API gateway were each built for a different job, and nobody owns the work that falls between them. When that work stalls, you get what I've come to call frozen pricing.
The Price Stopped Being a Fact
For most of the history of B2B software, price was something you negotiated once a year. The billing system's job was to remember the number and send the invoice on time. That arrangement rested on two assumptions that no longer hold for anything sold by the call: that the unit of sale stays put, and that the cost of serving it moves slowly.
The unit doesn't stay put. A request to a frontier model is metered on input tokens, output tokens and cache activity, and the cache alone can carry three prices. With Anthropic's prompt caching, for example, writing to the cache costs 1.25 or 2 times the input rate depending on how long you keep it, and reading from it costs a tenth of the input rate, or less on some models. Anyone building on top of that has to design their own unit on purpose, whether that's the request, the token, the tool call, the agent run or the resolved ticket. Most teams I talk to haven't settled on theirs, and the ones who have expect to change it.
The cost doesn't stay put either, and it doesn't move in one direction. For a model of a given capability, inference has been getting roughly ten times cheaper every year. But customers don't stay on last year's model, and the work they hand it keeps getting longer. Ethan Ding pointed out last year that tasks which used to return a thousand tokens now return a hundred thousand. Your cost per task can rise and fall inside the same quarter.
At the end of 2025, 37% of the companies in ICONIQ's survey of AI builders said they planned to change their pricing model within a year, and by mid-2026 the share using consumption-based pricing had gone from 35% to 42% in six months. A large part of the industry is repricing on infrastructure that was designed around an annual price change.
Where Usage Billing Stops
Usage-based billing was a genuine advance. It moved software pricing off the seat and closer to what customers actually consume, and companies like Twilio and AWS showed how far per-unit pricing can go.
Look at what a usage-billing system is asked to do, though. It receives events, aggregates them, applies a rate and produces an invoice. It assumes someone else has already decided what the billable unit is, what the rate should be, which contract a given call counts against and whether the customer is profitable at that rate. Those decisions get made upstream, in application code, in gateway config and, more often than anyone admits, in a spreadsheet that one person in finance fully understands. That was workable when prices changed once a year. It stops working when the unit and the rate move every quarter, because each change has to be made in several systems at once, by teams who don't report to each other.
Margin makes this harder, because AI isn't cheap to serve. At the 75 to 80 percent gross margins typical of SaaS, revenue was a decent proxy for gross profit. Bessemer's numbers from last year put the fastest-growing AI companies at an average gross margin of around 25%, often negative, and the steadier ones at around 60%. When every call carries a real cost, you need cost per call sitting next to revenue per call, and most billing systems only ever see the second.
This year two of the best-known independent usage-billing companies were bought by payment processors. Stripe closed its acquisition of Metronome in January, and Adyen closed on Orb in July. Both deals make sense for the buyers, and they say something about where the market thinks usage billing belongs, which is next to the money. For the companies doing the selling, the practical consequence is lock-in. If your rating logic lives inside your processor's billing product, changing processors means rebuilding every plan and every contract in someone else's schema.
What the Gateway Can't See
The API gateway has the opposite problem. It sits in exactly the right place, since every call goes through it, and it knows who made the call, what they called and how often. For request-level units, the gateway is the natural place to meter. For tokens, cache hits or outcomes it usually needs the service behind it to report them, which is why the application team turns up in almost every pricing change.
Gateways are also getting better at seeing what passes through them. The July 2026 revision of the Model Context Protocol copies the method of every request, and the name of the tool being called, into Mcp-Method and Mcp-Name headers so that gateways can route and inspect tool calls without parsing the body. There's a catch a platform team will spot quickly: only newer servers reject a header that doesn't match the body, so a gateway enforcing on those headers should check the protocol version first. Still, it tells you where the people designing the protocol expect enforcement to happen.
Seeing a request isn't the same as knowing what to charge for it. Most gateways don't know what the customer signed, meaning the plan, the prepaid balance or the discount someone agreed to at the end of last quarter. Apigee's monetization feature is the exception, and it only sees Apigee traffic, which is awkward for the many large companies that run more than one gateway. No gateway knows what a call cost you upstream in model spend, infrastructure and payment fees. And each call is judged on its own, so an agent task that fans out into forty calls across three services looks like forty unrelated requests, when the thing you probably want to price is the task.
The controls a gateway does have were built to protect the service. Rate limits and quotas exist to stop one customer from hurting everyone else, and AWS says as much in its API Gateway documentation, where usage plan quotas and throttling are described as "applied on a best-effort basis" and not something to rely on for controlling costs. That's the right design for a gateway. It just means a quota isn't a price.
Agents Raise the Stakes
None of this started with agents, but agents make it much harder to ignore.
An agent can make thousands of calls in the time a person makes a handful. It doesn't read your pricing page, and it won't notice when it runs past an included quota. Entitlements, limits and margin checks now have to hold on traffic that nobody is watching in real time.
Agents are also starting to buy. Since September 2025 we've had OpenAI and Stripe's Agentic Commerce Protocol, Google's AP2 with more than sixty partners, a foundation for x402 from Coinbase and Cloudflare, and Stripe and Tempo's Machine Payments Protocol. They deal with how an agent authorizes a payment and how the money settles, and they push the pricing problem closer to the request. With x402, the server answers a paid request with a 402 response that states what the client has to pay. Whatever produces that response needs the customer's rate, discount and remaining balance at that moment. That's the gateway's blind spot again, now on the critical path of every sale.
I'd still be wary of building a monetization stack around agents specifically. Most of metering and billing doesn't care who sent a call. What agents change is attribution and grouping: which customer an agent is acting for, and which forty calls add up to one task. Get those right and the rest of the pipeline is the one you'd want for a plain API anyway.
Frozen Pricing
Put all of this together and you get the condition I mentioned at the start. You know your price is wrong. Inference got cheaper, or a heavy customer found a way to run expensive calls on a cheap plan, or a competitor moved first. You know roughly what the new price should be, and you can't ship it for a quarter or two, because the change touches the gateway, the application, the billing system and finance's spreadsheet, and nobody owns all four.
It never shows up as an incident. It shows up as margin that leaks every month the price lags the cost, and as pricing ideas that never get proposed because everyone already knows how long they'd take. It's as true for a company selling plain REST APIs to insurers as for anyone selling tokens.
There's an opposite failure, and it's just as instructive. In June 2025 Cursor moved its $20 Pro plan from 500 requests a month to $20 of frontier-model usage at API prices. Within three weeks its CEO wrote that the changes "were not communicated clearly", and the company refunded unexpected charges. A lot of people took the lesson to be that you shouldn't touch pricing. I think that's the wrong lesson. Cursor's costs had moved, and the price had to follow.
Both failures come from the same gap. There's no single place where a price is defined, run against last month's real usage customer by customer, and rolled out with each customer able to see what their next bill will be. Without that, you either ship and surprise people, as Cursor did, or you don't ship at all.
Five Questions Worth Asking
If you sell an API, an MCP server or an AI product, these are the questions I'd put to your team. You don't need our product to answer them.
- When did you last change your pricing, and how long did it take from decision to live?
- Can you tell me yesterday's cost to serve for your largest customer, before the invoice runs?
- If you changed your billable unit tomorrow, how many teams would have to ship code?
- When a customer on last year's contract starts using something you launched last month, does their bill come out right without anyone checking it by hand?
- Could you switch payment processors, or add a second gateway, without re-implementing your pricing?
If your answers are "last year", "no", "three or more", "someone checks" and "no", you're in good company.
What We're Building
Those five questions are, more or less, the spec for Aforo. We describe it as an API and agentic monetization platform, which in plain terms means it meters usage on the gateway you already run, then prices it, bills it and recognizes the revenue, with the invoice and the revenue schedule built from the same events.
The meter runs as a plugin on Kong, Apigee, AWS API Gateway, Azure API Management or MuleSoft. It reports usage asynchronously and fails open, so no traffic moves and an Aforo outage never blocks a call.
We're early. We're signing our first design partners now and don't have customer logos to show, so I'd rather you judge the product than my description of it. The live demo gives you a full workspace with sample products, pricing and usage.
And if you've lived through a pricing change that took longer to ship than the product it was pricing, I'd like to hear about it, particularly if you think I've got part of this wrong. I'm at jay@aforo.ai.
— Jay