Nvidia has told some of its largest customers that prices for servers containing its AI chips will rise by more than 15 percent on systems delivered from early 2027, according to Bloomberg — the second major price increase in under a year, layered on top of hikes of roughly 30 percent across almost all product lines in July.
The warning, communicated through server manufacturers that assemble systems for large data center operators, lands at a delicate moment for the AI industry: enterprises are wrestling with the cost of massive infrastructure buildouts, and the memory market — the stated driver of the increase — shows no sign of cooling. Nvidia is scheduled to report quarterly results on August 26, where analysts expect fresh detail on demand, supply, and costs. For more context on this story, see our ongoing AI news.
Which systems are affected
According to the reporting, the increase will apply to servers built around Nvidia's flagship Grace Blackwell platforms and its upcoming Vera Rubin generation. The hikes will not be uniform: the size of the adjustment depends on the generation of Nvidia silicon involved and the amount and configuration of memory in each system.
Server manufacturers serving major technology companies — including Microsoft, Google, and Oracle — have reportedly begun notifying customers about the forthcoming price changes for systems shipping in early 2027.
Memory is the culprit
The primary driver is the surging cost of memory. High-performance AI systems require enormous quantities of high-bandwidth memory, and demand from data center expansion has sent prices for HBM and LPDDR5 climbing. Memory manufacturers Samsung Electronics, SK Hynix, and Micron Technology are all facing strong demand as semiconductor capacity is increasingly directed toward high-bandwidth memory and other advanced computing products.
Gaurav Gupta, a vice president and analyst at Gartner, said the increases reflect pressures across the whole AI supply chain, not just RAM. "Memory prices are going up, especially HBM and LPDDR5, but there are other aspects, like leading-edge foundry wafers, advanced packaging, and other component shortages," he told CIO.com, adding that those issues generate longer lead times, "which typically translates to higher prices."
And he does not expect relief soon. "In the current environment of strong demand and limited supply, we expect this situation to continue in the near to mid-term," Gupta said — meaning higher costs both for organizations deploying their own servers and for those renting compute in the cloud.
Little choice for buyers
For CIOs, the calculus is unappealing. Consultants and analysts say enterprises have little choice but to accept that AI infrastructure pricing will keep rising, and switching suppliers — even if possible — would not sidestep the increases, because memory and component costs affect the entire market.
Scott Bickley, an advisory fellow at Info-Tech Research Group, argued that Nvidia is actually absorbing part of the blow. His calculations suggest the company is likely eating some of its own rising costs and passing along only a fraction to its largest customers. But he was blunt about the leverage buyers have: "The workload cost is going down per token while the underlying hardware and infrastructure costs are going up," he said. "If you are directly building out your own clusters, this is an automatic uplift to an already egregiously expensive solution. If you are buying your own hardware, you're going to have to suck it up. You are not going to negotiate your way out of this."
There is one silver lining Bickley pointed to: per-token prices for AI models appear to be falling, which gives buyers a way to blunt hardware inflation by architecting workloads that rely less on raw memory capacity.
Squeezing more from what you already own
That strategy — getting more mileage from existing hardware — is where several analysts converge. Bickley pointed to techniques such as model routing, compression, and batch processing as ways to raise utilization of clusters already in service.
Flavio Villanustre, CISO for the LexisNexis Risk Solutions Group, agreed that the squeeze is fundamentally a supply-and-demand problem that should plateau once memory production ramps to meet AI-driven demand — but he does not see that happening in the next few months. In the meantime, he noted, some AI model vendors are adapting their models to run better in memory-constrained environments. He pointed to Google's Gemma E4B and similar models, which use a hierarchical tiered approach that pulls only the necessary parts of a model into memory rather than holding the entire model in RAM.
Mike Wilkes, enterprise CISO at Aikido Security, argued the price hikes could do some good if they force more disciplined architectures. Enterprises have spent the past few years treating frontier-model tokens as an "infinitely elastic utility," he said — routing workloads to the biggest model regardless of whether the task required frontier-level reasoning. The right architecture is increasingly hybrid: reserving roughly 10 percent of consumption for frontier models on problems that genuinely need them, while pushing classification, extraction, summarization, and routine agent actions toward small language models and open-weight models running on infrastructure the enterprise controls.
"That gives CIOs leverage against price gouging or unilateral price setting," Wilkes said.
An inflation signal for the AI economy
The increase is also a signal about where the AI boom's costs are heading. As memory, wafers, and packaging all rise in price, the inflation is spreading from chipmakers to server vendors to cloud providers and ultimately to enterprises building on AI. With Nvidia's earnings report due August 26, buyers and investors alike will be watching for how much of that cost pressure the company absorbs — and how much it passes downstream.
---
Stay Ahead of AIGet the latest AI news, analysis, and breakthroughs — all in one place.
Read more AI news →