One-day data center outage could cost tens of millions in lost compute fees
A handful of on-site power projects built to feed the artificial intelligence boom are already showing cracks, in both their equipment and the financial assumptions underwriting them, as hyperscalers race to bring capacity online before their turbines and engines have proven they can hold up under the volatile demand patterns AI workloads create.
Three of four U.S. data centers operating with off-grid or partially grid-connected power have experienced reported equipment failures, according to data provider Cleanview. The four sites are xAI’s Colossus and Colossus II in Memphis, Crusoe’s Stargate Abilene in Texas and Vantage Data Centers’ Vantage VA 2. Cracks have formed at gas-fired turbines powering xAI’s Colossus data center in Memphis, Tennessee, and combustion engines at other sites have suffered broken cranks, Bloomberg reported earlier this month. In July, a critical power failure at Vantage Data Centers’ Loudoun County, Virginia campus forced the operator to switch to diesel backup power for about a day, a spokesperson confirmed. Crusoe has also run into turbine glitches and other technical problems at its Stargate Abilene site, according to the Information; Oracle runs servers for OpenAI at the Stargate campus.
The failures appear tied to the unusual demands AI workloads place on power systems. AI computations cause massive power demand spikes that drop off as results are gathered, according to a paper from Schneider Electric. As data centers scale up, those oscillations could reach hundreds of megawatts or even gigawatts in size. Such swings can stress and shorten the life of engines and turbines and, in extreme cases, cause shaft fractures, according to energy research and consulting firm Wood Mackenzie.
The same dynamic has surfaced on the broader grid. When grid-connected data centers suddenly drop power usage, grid operators have had to scramble to take power supply offline to keep a sudden excess of supply from damaging power plants and infrastructure. Troy Patton, general manager at on-site power provider AlphaStruxure, said in an interview that one potential hyperscaler client required power response times measured in tens of milliseconds rather than the hundreds of milliseconds the industry typically plans for, a threshold Patton said no rotating equipment or lithium battery has had to meet.
The financial stakes are rising alongside the technical ones. Anthropic is paying xAI a monthly fee of $1.25 billion for computing capacity on Colossus and Colossus II. While the underlying contract terms are not public, that figure implies a single day of downtime could be worth tens of millions of dollars, either in lost fees for xAI or in costs to Anthropic for going without that computing capacity.
Power providers face parallel exposure. Industry contracts typically require providers to refund customer fees and cover repair or replacement costs when they fail to deliver power, and to absorb millions of dollars in equipment replacement after extended outages, according to industry experts and public filings. Nina Sadighi, founder of Eradeh Power Consulting and an electrical engineer who formerly worked at Amazon and Tesla, said one of her former client’s on-site data-center power systems saw a provider replace equipment worth millions of dollars following a multiday outage. Cracked turbines and broken engine cranks both qualify as major failures that require replacement, Sadighi said.
Solaris Energy Infrastructure, which operates turbines for the Colossus campus, must refund customer fees and cover equipment repair or replacement under its contracts, but cannot be held liable for customers’ lost revenue from downtime, according to a public filing. Solaris co-CEO Bill Zartler, in an emailed statement, said the company has “not seen any extraordinary operating issues or failure rates on our generation equipment.” On an earnings call, Solaris did not directly confirm or deny the existence of turbine cracks at the Colossus campus but said its turbines were “in great shape.”
Hyperscalers carry highly profitable business models and strong balance sheets, but many of the newer power providers do not. In addition to early players such as Solaris and VoltaGrid, several oil-and-gas companies, including Liberty Energy, Williams Companies, Atlas Energy Solutions and Kodiak Gas Services, have announced business units dedicated to on-site power for data centers. Investors have been willing to pay up for those with contracts: Solaris, which signed two additional customers after xAI, trades at an enterprise value of roughly 10 times forward earnings before interest, taxes, depreciation and amortization, about 19% higher than its oil-field service peer group.
The scale of planned off-grid builds is set to grow. Amazon is planning an AI data-center campus in Pecos County, Texas, that would use natural-gas-powered turbines to generate as much as 7.65 gigawatts of power, according to Cleanview, equivalent to seven or eight nuclear power plants.
Whether the early failures are fixable learning pains or structural problems is not yet clear. Sadighi said nobody has enough operating history yet to predict how the equipment will hold up. What is clear, she added, is that data-center power loads at this scale are challenging enough to cause problems even for experienced, large-grid operators.