Rafay and Solvimon are partnering for Neocloud billing

Announcements
Read time: 6 min

Arnon Shimoni
✓ Expert opinion
Solvimon is partnering with Rafay. Rafay runs the control plane on top of GPU fleets and token farms, we run the ledger for the commercial model so an operator can get from raw capacity to a billable service without writing billing code for it.

Why are we so serious about this??
An IDC CIO Playbook 2026 commissioned by Lenovo found 84% of organizations expect to run AI on-prem or at the edge alongside cloud. The reasoning is, as you'd expect: sensitive inference stays inside for data residency, and burst work goes out to the cloud.. .But nobody wants their model access governed by someone else's rate limits during an incident.

What tends not to get planned for is that this makes you a cloud provider.
Recently, TELUS started running their own sovereign AI studio in Canada while banks, hospital groups, government ministries, and regional hosters are doing versions of the same thing, with build-outs across India, Indonesia, across Africa, Australia, and even Naver in Korea.
We wrote about the public-market version of this in neoclouds owe customers years of compute. Neoclouds get paid in one of 3 ways, and the prepaid capacity commits are the ones that fund the hardware and collateralize the debt.
Why doesn't the AWS way of metering transfer to GPUs?
Nearly everything the industry knows about charging for compute traces back to the EC2 of the early 2000s, billed on machine time. That fit a general-purpose box running a web server, and it survived two decades because the fit was genuinely good. On a GPU it's a much worse approximation that doesn't fit well.
Cast AI's 2026 State of Kubernetes Optimization Report looked at roughly 23,000 clusters across AWS, Azure, and GCP and put average GPU utilization at 5%, while Gartner expects AI infrastructure to add $401 billion in new spend this year. That 5% figure is going to appear in ohg so many vendor decks for the next months (treat it as an indication, not the ground truth) but even if we triple that number, the value gets really detached from the wall-clock billing.
A tenant holding a card at 30% pays what a tenant driving it at 90% pays, and the operator eats the difference, which happens to be most of what the operator is actually selling.
What does a private AI cloud actually have to charge for?
Five things the fleet already produces that can be measured. Most of them don't appear on any bill in the market today.
Layer | What gets measured | Why it's chargeable |
|---|---|---|
Compute | Utilized GPU time, priority and preemption tiers | The efficient tenant gets rewarded, and the operator resells reclaimed headroom instead of giving it away |
Memory and movement | Data in and out of GPU memory, warm model state | Long-context and multimodal work drives real time and energy here. A warm start is a feature people pay for |
Output | Tokens in and out per model, response size, completed requests | The honest unit for anything delivered as a finished service |
Power | kWh drawn | The layer underneath already bills this way. A flat rate means the operator carries variability it isn't paid for |
Guarantees | Attested location of inference, air-gapped operation, per-request audit trail, model access scope | A jurisdiction guarantee carries a number a bank will pay that a commercial endpoint abroad cannot offer at any price |
A hospital buying inference capacity isn't buying GPU minutes, it's buying confidence that 20,000 scans clear the SLA. A ministry is buying reviewed permits, and a telco is buying a latency floor. The GPUs underneath are identical and the commercial products aren't, which is difficult to express in dollars per hour no matter how the rate card is structured.
How do Rafay and Solvimon fit together?
Rafay is already scheduling GPUs, enforcing multi-tenancy, and applying policy across bare metal, virtualized, and Kubernetes environments, which means it's already sitting on the signals that make these meters possible. It meters at tenant, team, and workload level, keeps entitlements separate from provisioning, and ships a token-metered serverless inference API for operators selling further up the stack. What it deliberately doesn't do is run pricing logic or move money.
That part is ours. Commit structures, drawdown against mixed SKUs, GPUaaS billing, and the revenue recognition that falls out of both. We run billing for fintechs, banks, and telcos today, and AI infrastructure is the most complicated pricing environment we've come across!
The way we set things up now is:
Rafay meters consumption and exposes it through the control plane
Solvimon pulls it, applies the operator's rate cards, and handles invoicing, settlement, and collection
Solvimon pushes rate cards and spend caps back into Rafay so entitlements stay in sync with what the customer actually bought
No human clicks through a dashboard anywhere in that loop. For the operator, going from GPU capacity to a live billable service turns into a configuration exercise rather than an engineering project.
Why does this end up as a billing and monetization problem?
Basically, the accounting depends on it. A three-year commit sits on the balance sheet as deferred revenue until the capacity is actually delivered, and AI factories energize in tranches rather than on a calendar. Under ASC 606 the recognition schedule has to be computed from what was committed, what was delivered, and what was drawn-down, per contract, on a rolling basis, etc.
Two of those three data streams come from the meter, so recognition is only really ever as-good-as the metering underneath it.
If you're thinking of standing up your own GPU cloud or token factory - you may only be planning for capacity and not how much it costs…
Read the full piece from the Rafay team or talk to us about GPUaaS billing.
Ready for billing v2?
Solvimon is monetization infrastructure for companies that have outgrown billing v1. One system, entire lifecycle, built by the team that did this at Adyen.



