Home Stocks Nebius Raised Inference Prices 18.3%, Not Cut Them
Stocks

Nebius Raised Inference Prices 18.3%, Not Cut Them

Share

Everyone knows the cost of intelligence falls every quarter. It is the single most repeated claim in AI infrastructure, and from 1 October it will be wrong at Nebius. The AI cloud company is lifting prices on Token Factory, its managed inference platform, by an average of 18.3% on pay-as-you-go dedicated endpoints, with B300 capacity rising 20% from $8.10 to $9.70 an hour, according to customer communications posted publicly and reported by Stocktwits on 23 September. Nebius has not formally announced the change. The stock closed at $243.48 on 24 September, within sight of a $299.86 52-week high against a $73.52 low.

The reason this matters is not the percentage. It is what the percentage reveals about the shape of the business. Compute is being sold the way ships are chartered, not the way software is licensed — and Nebius has effectively published its own charter rate curve. In its second-quarter shareholder letter the company said its core AI cloud contracts now carry annual contract value of $20–25m per megawatt, while the short-term and auction deals it began piloting in the third quarter show “a price opportunity in the $40-50 million per MW range.” That is a two-times spread between term and spot, disclosed by the operator, in writing. The Token Factory increase is simply the retail end of the same curve finally repricing.

The insight: compute is a charter market, and it just went into backwardation

Dry-bulk shipping spent a century learning this lesson. A shipowner signs multi-year time charters to finance the hull and keeps a slice of the fleet on the spot market to capture the spikes. When spot voyage rates run at twice the time-charter equivalent, two things happen in sequence: the owner shortens contract duration on new business, and the market builds a forward curve so everyone else can hedge the difference. Nebius is at step one. The industry is already at step two — the CME is about to list futures on the price of renting a GPU, and Kalshi has been building forward curves for compute for months.

Read the Nebius letter with that framing and the moves stop looking opportunistic and start looking systematic. The company closed four landmark deals in Q2 averaging more than $1bn of total contract value each. It said pricing on older-generation GPUs was “more than 30% higher” than in Q1. It ran its first capacity auction pilot and “secur[ed] the highest price we have cleared for NVIDIA Blackwell chips to date.” It signed its first short-duration deal — typically three to six months, “priced at a significant premium” — for a customer with an acute, time-bound need. Every one of those is a yield-management decision that a software company would never have to make and a shipowner would recognise instantly.

Key facts: Nebius pricing and Q2 2026 economics

  • Token Factory pay-as-you-go dedicated endpoints rise an average 18.3% from 1 October: H100 $4.05 to $4.70 (+16%), H200 $4.70 to $5.60 (+19%), B200 $7.40 to $8.70 (+18%), B300 $8.10 to $9.70 (+20%) — Stocktwits, 23 Sep 2026; not formally announced by Nebius
  • Q2 2026 group revenue $582.3m, up 454%; first-half revenue $981.3m, up 529% — Nebius Q2 2026 results, 12 Aug 2026
  • Adjusted EBITDA of $236.2m, against a $21.0m loss a year earlier; adjusted net loss narrowed to $33.2m; GAAP net loss from continuing operations was $190.4m — Nebius Q2 2026 release
  • AI cloud revenue $574.9m, up 514%, about 98% of group revenue; ARR of $3.0bn at 30 June, up 598% year on year and 56% from $1.9bn at 31 March — Nebius Q2 2026 shareholder letter
  • Annual contract value of $20–25m per MW on landmark deals; a stated price opportunity of $40–50m per MW on auction and short-term capacity — Nebius Q2 2026 shareholder letter
  • $40bn of customer commitments; more than $9bn of customer prepayments expected in 2026; roughly 70% of Q2 deals included prepayments covering 50–60% of the associated capex — Nebius Q2 2026 shareholder letter
  • Expected payback on Q2 deals shortened to 1 year 10 months from a historical two-to-three-year range; Q2 capital expenditure was about $5.7bn — Nebius Q2 2026 shareholder letter
  • Year-end contracted power guidance raised to 5 GW, with more than 1 GW of deployment per year planned from 2027 — Nebius Q2 2026 shareholder letter

What is actually happening, and why the deflation story broke

The “cost per token collapses” thesis is not wrong about silicon. Each generation of accelerator does deliver more tokens per dollar of hardware. The thesis fails because the price a customer pays for inference is not set by the cost of the chip. It is set by the scarcity of the megawatt the chip sits in, and megawatts are not getting cheaper. They are the binding constraint across the entire buildout — which is why data-centre power contracts have become financial instruments in their own right, with Vistra taking a 5% equity stake alongside a 207 MW Odessa power agreement and Fermi’s $6.5bn TensorWave lease turning on closing deadlines rather than chip supply.

Nebius is one of a small number of operators able to demonstrate that scarcity in its own numbers rather than assert it. Revenue growing 454% is a demand statement. Pricing on older generation GPUs rising more than 30% quarter on quarter is a scarcity statement, and a much more informative one: when last year’s silicon gets more expensive, there is no supply.

The Token Factory increase extends that logic into managed inference, which had been the one part of the stack where competitive price cuts were routine. Inference workloads on Token Factory more than tripled in the second quarter. Nebius has also been buying its way up the stack — the letter credits acquisitions of Eigen AI and Clarifai with bringing “industry-leading inference optimization in-house,” and the Tavily acquisition pairs Token Factory’s inference with real-time retrieval. An operator that has just tripled volume, improved efficiency and acquired the optimisation layer is raising list prices anyway. That is not a cost pass-through. That is pricing power.

A caveat that deserves more attention than it has received: this increase has reached customers but not the tape. Nebius has issued no press release and made no filing, and the reporting rests on screenshots of customer communications. For a Nasdaq-listed issuer whose entire equity story is pricing power, that is an odd place to leave a material commercial change — and it means the first authoritative confirmation will probably arrive on the third-quarter call rather than in a disclosure document. Treat the specific per-hour numbers as well-sourced but unconfirmed until then.

Market impact: the spread is the whole thesis, in both directions

Combine the two disclosures in the shareholder letter and you get a number nobody is quoting. Nebius’s contracted book earns $20–25m per MW. Its spot and short-duration book is being priced at $40–50m per MW. Against a year-end contracted power target of 5 GW, every percentage point of the fleet that migrates from term pricing to spot pricing is worth roughly $200m–$250m of annualised revenue at the midpoint of the gap. That is the bull case, stated arithmetically rather than rhetorically.

It is also the bear case, because the arithmetic runs backwards just as cleanly. Headline spot rates are not fleet rates. FinanceFeeds made exactly this point about a rival last week: CoreWeave’s $40m per megawatt is not what the fleet earns, because a single marquee contract at the top of the market tells you nothing about the weighted average across older capacity, older GPUs and older terms. Nebius’s own letter is honest about this: the $40–50m figure describes a “price opportunity” and one signed deal, not a book rate.

Metric What the bulls see What the bears see
18.3% inference price rise Pricing power in the highest-margin, stickiest part of the stack List prices, not realised prices; large customers negotiate, and none of this is company-confirmed
$20–25m vs $40–50m per MW A 2x repricing runway as term contracts roll The spread is a cycle signal; spot converges to term when capacity lands
70% prepayment, 50–60% of capex Customers financing the buildout; payback cut to 1yr 10m Prepayment is the market standard in a shortage — it disappears the moment supply loosens
$40bn commitments, 5 GW target Contracted visibility well into 2027 Q2 capex was $5.7bn in a single quarter, funded partly by a $5.75bn convertible note

The financing detail belongs in the bear column and rarely gets there. Nebius closed a private offering of convertible senior notes with gross proceeds of approximately $5.75bn in late August, having launched the deal at $4.5bn and upsized it. Combined with quarterly capex of $5.7bn, that tells you the prepayments are necessary rather than merely flattering. The business is now large enough that a pricing cycle and a financing cycle are running simultaneously, and they do not have the same duration.

The regulatory and structural tension

Two pressures are building on this model at once. The first is sovereignty. Nebius’s site map spans the United States, France, the United Kingdom, Spain and Israel, and on 8 September it announced a partnership with Palantir to deliver a complete sovereign AI stack to Palantir’s customers. Sovereign requirements fragment capacity by jurisdiction, and fragmented capacity is less fungible — which supports pricing but complicates the yield-management strategy that makes the $40–50m per MW opportunity real in the first place. You cannot auction a megawatt in Frankfurt to a customer who is contractually required to run in Virginia.

The second is accounting scrutiny. The market has begun interrogating hyperscaler depreciation assumptions, and a business that capitalises billions of dollars of GPUs each quarter is exposed to the same question: over how many years is last year’s accelerator being written down, and does a 30% price increase on older-generation GPUs justify a longer useful life or merely postpone the reckoning? Nebius is arguing, in effect, that its older assets are appreciating in rental value. That is a defensible claim in a shortage and an indefensible one in a glut, and auditors will ask which is being assumed.

What happens next

One: the Token Factory rise sticks, and at least one major rival follows before year-end. Price increases in a shortage are tested by whether customers leave. Inference volumes tripled in a quarter, switching costs on a managed platform with in-house optimisation are real, and the rival neoclouds face identical power economics. The tell will be whether a competitor matches on dedicated endpoints rather than on headline per-token rates, which are easier to discount quietly.

Two: the disclosure gap closes on the Q3 call, and the number Nebius chooses to highlight will be ARR, not price. ARR went from $1.9bn at end-March to $3.0bn at end-June. A company with pricing power and tripling inference volume has every incentive to keep the market focused on the run-rate rather than on a list-price schedule that invites competitive response. If management confirms the 18.3% figure directly, take it as a deliberate signal to the buy side.

Three: the $20–25m to $40–50m spread narrows through 2027, and the compression is the risk nobody is hedging. Nebius plans more than 1 GW of deployment a year from 2027. So does everyone else. Spot premiums exist because capacity is late, and the entire industry is racing to make capacity early. When the CME contract lists, the forward curve will price that convergence long before the income statement does — which is precisely why a compute futures market is arriving now rather than in five years.

For now, the practical takeaway for anyone budgeting AI spend is unglamorous and immediate: the line item goes up on 1 October. The deflation everybody planned around is real at the level of silicon and absent at the level of the invoice, and the gap between those two facts is where the entire neocloud equity story currently lives.

FAQ

Is Nebius raising inference prices?
Yes. Token Factory pay-as-you-go dedicated endpoint prices rise an average of 18.3% from 1 October, according to customer communications reported by Stocktwits on 23 September. Hourly rates move from $4.05 to $4.70 for H100, $4.70 to $5.60 for H200, $7.40 to $8.70 for B200 and $8.10 to $9.70 for B300. Nebius has not issued a formal announcement, so treat the specific figures as well-sourced but unconfirmed.

Why are inference prices rising when compute is supposed to get cheaper?
Because the binding constraint is power and capacity, not silicon efficiency. Tokens per dollar of hardware do improve each generation, but the price a customer pays reflects scarcity of the megawatt the hardware occupies. Nebius disclosed that pricing on older-generation GPUs was more than 30% higher in Q2 than in Q1 — a clear scarcity signal, since last year’s silicon should be getting cheaper, not dearer.

What is Nebius Token Factory?
Token Factory is Nebius’s managed inference platform, which lets customers run open-weight and fine-tuned models in production without operating the infrastructure. Production inference workloads on the platform more than tripled in the second quarter of 2026, and Nebius has brought optimisation capability in-house through the acquisitions of Eigen AI and Clarifai.

How much revenue does Nebius make per megawatt?
The company’s Q2 2026 shareholder letter states annual contract value of $20–25m per megawatt on its four landmark deals, each with total contract value above $1bn. It separately describes a “price opportunity in the $40-50 million per MW range” for auction and short-term capacity, and says it signed its first such deal in Q3. The second figure is an opportunity and one contract, not a fleet-wide rate.

How is Nebius funding its capacity buildout?
Through a mix of customer prepayments, secured facilities and convertible debt. Roughly 70% of Q2 deals included prepayments covering 50–60% of associated capex, the company expects more than $9bn of prepayments in 2026, and it closed a convertible senior note offering with approximately $5.75bn of gross proceeds in late August. Q2 capital expenditure was about $5.7bn.

What are the main risks to the Nebius pricing thesis?
Convergence and financing. The gap between term pricing and spot pricing exists because capacity is scarce, and Nebius plans more than 1 GW of annual deployment from 2027, as do its competitors. Prepayment terms that currently de-risk the buildout are a feature of shortage and would fade in a glut. Separately, a business capitalising billions of dollars of accelerators each quarter carries real depreciation-assumption risk if rental values for older GPUs stop rising.

Share
Related Articles

Google’s Project Suncatcher Puts 4 TPUs in Orbit on…

Project Suncatcher is not a space story. On 1 October, Google will...

IonQ Superion 256: A Sale Today, Revenue in Late 2027

A quantum computer was sold this week, and almost nothing about that...

The Loonie Slips Past 1.4150 as US Yields Overpower Oil

The Canadian dollar weakened beyond 1.4150 per US dollar on Friday as...

Ethereum Price at $2,717 as ETF Buyers Return – Bull…

Updated 25 September 2026, 13:10 UTC. Ethereum trades at $2,717.02, up about...