AI pricing is understood now but the token is still just a cost

Insights

AI pricing is understood now but the token is still just a cost

AI pricing is understood now but the token is still just a cost

Read time: 8 min

Subscribe to our newsletter

Arnon Shimoni

✓ Expert opinion

Open pretty much any AI pricing guide published since 2025 (including some of my old ones) and it starts the same way: pricing AI is the hardest, most unsolved problem in software and no one has figured it out.

Well, I disagree by this point. The reason it still feels unsolved is one substitution people keep making, over and over and over again.

Because you look at a token and that's considered the price.

A token isn't a rpice.

Yes, a token is what the model costs you to run an input. The output costs you more (output tokens run 3-8x input) and the outcome, the thing your customer wants from you, isn't measured in tokens at all. Once you stop pricing the cost of your own input and start pricing the value on top of it, the "playbook" if you can call it that is becoming a bit clearer.

The standard advice we see all over makes the substitution for you, as every explainer on token pricing lands on the same instruction: work out what a request costs to run, then set the token price to cover it, plus a margin.

A variety of token pricing advice

Cost-plus. It's clean, it's defensible, and it quietly assumes the token is the thing you sell. For a foundation-model API, fine... the token is the product. For an application built on top, cost-plus on the token caps your price at your own COGS curve, right when the customer is buying an outcome worth many times that.


So what does work now?

What's working right now

I see three main moves which you can make, but a combo of all three is also possible.

Bundle the token into a plan or seat to buy adoption

Notion folded AI into its higher-priced plans instead of selling it as a separate add-on. Cursor sells seats with included usage (Pro at $20 a month, Ultra at $200 with far more headroom). Perplexity does the same across its $20 to $325 tiers. You eat the token cost and bet that adoption pays it back through retention and expansion. It's a bet on your own product. Not my favourite structure as a consumer, but for land-and-expand it works.

Price on the value the tokens produce

Obviously we all know Fin by now with their $0.99 per resolution, on top of a $49 base that includes the first 50 - we also have Zendesk, AgentForce, Justt, Chargeflow and so many more doing this.

The tokens are the vendors' cost to manage, invisible to the buyer.

When you can define the outcome and measure it, this tracks value so much better than anything else on the list.

Gate it so it can't run away from you

This is the move that makes the other two safe, and it's one that's easy to skip over.

When you create a bundled plan with no ceiling and no entitlement gating, you're inviting users to consume much more than what they pay for.

Cursor's tiers are gating in practice with each seat having a usage threshold and the next tier up is priced for the accounts that outgrow it.

Our partner PricingSaaS's data shows restlessness on the ground: add-on launches ran hot through 2025, peaking around 21.6 per 100 tracked companies in Q2, as teams bolted new shapes onto their core plans.

Our partner PricingSaaS's data shows restlessness on the ground: add-on launches ran hot through 2025, peaking around 21.6 per 100 tracked companies in Q2, as teams bolted new shapes onto their core plans.

When done correctly, gating add-ons protects both sides: you hold your margin, and the customer gets a balance they can watch and budget against instead of a surprise invoice.

Unfortunately, wone badly it's an invisible wall that trips mid-workflow with no warning, which customers experience as a trap and they eventually churn.

It's not a ceiling per se, it's a live balance that the customer shouldn't have to get stuck on.

Pricing KPIs you should know

If AI pricing is understood, these are the numbers that prove you understood it.

Luckily, there are just six of them

KPI

What it tells you

Rough benchmark

  1. Token cost as % of revenue

Your AI gross margin, inverted

AI GM rose from 41% (2024) to 52% (2026), floor now around 60-65%

  1. Gross margin per customer

Which accounts actually make money

Watch the spread, not the average

  1. Bundle utilization rate

Breakage on one end, overrun on the other

Aim for a band, not a ceiling

  1. Consumption / overage as % of revenue

How much expansion your usage drives

Rising is usually the healthy sign

  1. Net revenue retention

Whether the model compounds

Usage/hybrid 115-130%+ vs 95-105% for flat

  1. Effective rate vs list

What caching and batch leave you

Recompute quarterly

That second one is the one I like the most, and it's one that gets neglected a bunch. You may find more comfort in a blended average but your top 5% of users are either your best case studies or your margin losers, and the average will never tell you which...

How to measure it

One big habit we noticed at Solvimon is some companies meter every billable event in real time and attribute the cost to the customer and the task, instead of reconciling it after the invoice gets finalized - and that's a winning activity.

We believe you can't effectively bundle what you find difficult to track and it gets harder as you want to align value. It's damn near impossible to charge for an outcome you haven't defined and measured.

That's the work Solvimon does underneath. Metering turns raw events into billable, attributable quantities per customer per task. Entitlements enforce the ceiling before it's breached. Together they make every KPI above computable in real time.

The token price is the number you can look up in two seconds. Whether you built a business on top of it is the number that takes real instrumentation to see... worth knowing which one you've been watching.

FAQ

Is AI pricing actually a solved problem?

I'd say yes because the models are understood: bundle, meter, value/outcome-based, hybrid. What's unsolved for most teams is the plumbing underneath, i.e., metering and gating consumption per customer accurately enough to trust the margin.

The strategy is known, doesn't mean it looks identical for everyone.

Should I charge my customers per token?

Only if the token is genuinely your product (e.g., a foundation-model API or raw inference).

For 99% of applications built on top, the token is your cost of input, and billing it straight to the customer passes your supplier's meter through at a markup which feels bad for the customer.

Bundle it or price the outcome instead.

What's a healthy token-cost-to-revenue ratio?

It maps inversely to gross margin. AI-native gross margins climbed from about 41% in 2024 to 52% in 2026, with the durable floor landing around 60-65%, and hybrid SaaS-AI products sitting higher. If token cost eats more than 40% of revenue, look at the pricing model before you blame the model provider.

What is outcome-based pricing?

You charge when the AI delivers a defined result, e.g., a resolved ticket or a completed task. Intercom's Fin ($0.99 per resolution) is the clearest live example. It tracks value closely, and it asks you to define the outcome precisely and handle the cases where the AI half-succeeds.

How do I measure gross margin per customer?

Attribute token and infra cost to each account, not just aggregate COGS, then set it against what the account pays. This needs per-customer metering. The blended average hides your loss-making whales, which are usually the accounts you're proudest of.

Ready for billing v2?

Solvimon is monetization infrastructure for companies that have outgrown billing v1. One system, entire lifecycle, built by the team that did this at Adyen.