AI pricing is understood now but the token is still just a cost.

Insights

AI pricing is understood now but the token is still just a cost.

AI pricing is understood now but the token is still just a cost.

Read time: 3 min

Subscribe to our newsletter

Arnon Shimoni

✓ Expert opinion

Open pretty much any AI pricing guide published since 2025 (including some of my old ones) and it starts the same way: pricing AI is the hardest, most unsolved problem in software and no one has figured it out.

Well, I disagree by this point. The reason it still feels unsolved is one substitution people keep making, over and over and over again.

Because you look at a token and that's considered the price.

Yes, a token is what the model costs you to run an input. The output costs you more (output tokens run 3-8x input) and the outcome, the thing your customer wants from you, isn't measured in tokens at all. Once you stop pricing the cost of your own input and start pricing the value on top of it, the "playbook" if you can call it that is becoming a bit clearer.

So what does work now?

What's working right now

I see three main moves which you can make, but a combo of all three is also possible.

Bundle the token into a plan or seat to buy adoption

Notion folded AI into its higher-priced plans instead of selling it as a separate add-on. Cursor sells seats with included usage (Pro at $20 a month, Ultra at $200 with far more headroom). Perplexity does the same across its $20 to $325 tiers. You eat the token cost and bet that adoption pays it back through retention and expansion. It's a bet on your own product. Not my favourite structure, but for land-and-expand it works.

Price on the value the tokens produce

Obviously we all know Fin by now with their $0.99 per resolution, on top of a $49 base that includes the first 50 - we also have Zendesk, AgentForce, Justt, Chargeflow and so many more doing this.

The tokens are the vendors' cost to manage, invisible to the buyer.

When you can define the outcome and measure it, this tracks value so much better than anything else on the list.

Gate it so it can't run away from you

This is the move that makes the other two safe, and it's one that's easy to skip over.

When you create a bundled plan with no ceiling and no entitlement gating, you're inviting users to consume much more than what they pay for.

Cursor's tiers are gating in practice with each seat having a usage threshold and the next tier up is priced for the accounts that outgrow it.

The market runs all of these at once. Chargebee's 2025 data puts hybrid pricing (a base plus usage) at 43% of companies, heading for 61% by the end of 2026. Our read of PricingSaaS data shows the same restlessness on the ground: add-on launches ran hot through 2025, peaking around 21.6 per 100 tracked companies in Q2, as teams bolted new shapes onto their core plans. (The tracked set jumped to 669 companies by 2026Q2, so treat the later quarters as a bigger, noisier sample, not a slowdown.)

[Diagram: the decoupling] token cost (your COGS) → metering + entitlements → price (bundle or outcome). Hand-drawn / Excalidraw.

Pricing KPIs you should know

If AI pricing is understood, these are the numbers that prove you understood it. Six, not forty.


KPI

What it tells you

Rough benchmark

Token cost as % of revenue

Your AI gross margin, inverted

AI GM rose from 41% (2024) to 52% (2026), floor now around 60-65%

Gross margin per customer

Which accounts actually make money

Watch the spread, not the average

Bundle utilization rate

Breakage on one end, overrun on the other

Aim for a band, not a ceiling

Consumption / overage as % of revenue

How much expansion your usage layer drives

Rising is usually the healthy sign

Net revenue retention

Whether the model compounds

Usage/hybrid 115-130%+ vs 95-105% for flat

Effective rate vs list

What caching and batch actually leave you

Recompute quarterly

Gross margin per customer is the one almost everyone skips, and it's the one that bites. The blended average is the most comforting number on your dashboard and the most useless. Your top 5% of users are either your best case studies or your margin leak, and the average will never tell you which... you find that out at month-close, three months too late, in a meeting nobody enjoys.

How to measure it

The teams who run these numbers well share one habit. They meter every billable event in real time and attribute the cost to the customer and the task, instead of reconciling it after the invoice ships.

That habit is what the whole playbook stands on. You can't bundle what you can't track. You can't charge for an outcome you haven't defined and measured. And gating a limit means seeing it live, before the overrun, not reading about it at quarter-close. Gross margin on this customer this month is only a real number when the events, the costs, and the entitlements sit in one place.

That's the work Solvimon does underneath. Metering turns raw events into billable, attributable quantities per customer per task. Entitlements enforce the ceiling before it's breached. Together they make every KPI above computable in real time. The team built this at Adyen, where the meter had to be right across €1T+ in volume, so the boring part (getting the number right, every time) is the part we treat as solved. For the deeper version on the billing side, we wrote about what usage-based billing actually requires from your stack.

The token price is the number you can look up in two seconds. Whether you built a business on top of it is the number that takes real instrumentation to see... worth knowing which one you've been watching.

FAQ

Is AI pricing actually a solved problem? The models are understood: bundle, meter, value/outcome-based, hybrid. What's unsolved for most teams is the plumbing underneath, i.e., metering and gating consumption per customer accurately enough to trust the margin. The strategy is known. The instrumentation is where people are still stuck.

Should I charge my customers per token? Only if the token is genuinely your product, e.g., a foundation-model API or raw inference. For applications built on top, the token is your cost of input, and billing it straight to the customer passes your supplier's meter through at a markup. Bundle it or price the outcome instead.

What's a healthy token-cost-to-revenue ratio? It maps inversely to gross margin. AI-native gross margins climbed from about 41% in 2024 to 52% in 2026, with the durable floor landing around 60-65%, and hybrid SaaS-AI products sitting higher. If token cost eats more than 40% of revenue, look at the pricing model before you blame the model provider.

What is outcome-based pricing? You charge when the AI delivers a defined result, e.g., a resolved ticket or a completed task. Intercom's Fin ($0.99 per resolution) is the clearest live example. It tracks value closely, and it asks you to define the outcome precisely and handle the cases where the AI half-succeeds.

How do I measure gross margin per customer? Attribute token and infra cost to each account, not just aggregate COGS, then set it against what the account pays. This needs per-customer metering. The blended average hides your loss-making whales, which are usually the accounts you're proudest of.

Ready for billing v2?

Solvimon is monetization infrastructure for companies that have outgrown billing v1. One system, entire lifecycle, built by the team that did this at Adyen.