❌

Reading view

Meet the innovators under 35 shaping climate tech

Each year, the editorial team at MIT Technology Review puts together a list of 35 innovators under 35—a group of researchers, inventors, and other young minds worth following.

The team worked on the newest edition of the list for months, and the final slate includes nine individuals from all over the world in the climate and energy category. Each one has a fascinating story and is tackling an important challenge.

I think it’s worth zooming out and considering the energy and climate awardees as a group. Taken together, these innovators and their work can tell us something about where climate tech is at this moment—and where it’s heading.

AI is the dominant technology story, both for its potential and its challenges.

We split the innovators into four main categories this year: biotech, climate and energy, computing and robotics, and AI. It probably won’t surprise you that AI features heavily in the work of many innovators in other categories.

Climate innovator Jae-Won Chung, for example, built software to make AI more energy-efficient. By measuring the energy demands of open-source models, he hopes the industry can better understand and address the impact of AI. (If this work sounds familiar, it’s because we spoke with him last year for our investigation into AI’s energy demands.)

But AI also has the potential to improve many areas of research. Jing Wei is using AI to track pollution more effectively, essentially using machine learning to fill in gaps in data from disparate sources like satellites and weather stations. Zhonghua Zheng developed AI climate models that work better for cities, a well-known blind spot for traditional models.

We need better ways to get the critical materials used to build new technologies.

As we begin to rely on new technologies to power our world, we’ll see a major shift in the materials we need to build them.

Lithium is a prime example: The metal underpins lithium-ion batteries, which are crucial not only for electric vehicles, but also for large-scale energy storage on the grid. We could face lithium shortages as soon as this decade, and the prospect of supply crunches applies to other critical minerals, too—copper is another one to watch closely.

Brine is currently the cheapest source of lithium, but the process to get the metal out can take months and harm the local environment. Mohammad Alkhadra is the cofounder and CEO of Lithios, a startup working to quickly and efficiently extract lithium from brines.

Hardrock ore is the most common source of lithium, but it’s more expensive than brine. Benjamin Mowbray cofounded and serves as CTO for Rock Zero, which is working to extract lithium from hardrock ore.

Addressing climate change will require overhauling all corners of our society, sometimes in surprising ways.

To reach net-zero greenhouse gas emissions we will obviously need to rethink major sectors, like the electrical grid and transportation, to move away from fossil fuels. But outside these primary sources of climate pollution are seemingly infinite, less obvious problems to figure out, too.

Heavy industry, including steel production, is a major one, making up about 7% of global greenhouse gas emissions. Laureen Meroueh is making cleaner, cheaper steel using a new kind of furnace that simplifies the chemical process required to produce the metal.

Plastics are generally made with fossil fuels, so we’ll need alternatives to this incredibly useful category of materials. Joseph Nguthiru is making a bioplastic replacement for fossil-derived packaging that uses an invasive weed. Also using available materials in a creative way, Diana Orembe is making fish food for aquaculture with food waste.

And refrigerants are often incredibly powerful greenhouse gases. Jinyoung Seo is developing solid refrigerants that could eliminate worries about leakage. A device using these materials could reduce energy consumption by 20% compared to conventional technology.

I’m constantly learning about new challenges we face in the climate and energy world, and I’m often surprised by the ideas people are coming up with to address them. For more on all the under-35 innovators and their work, check out our full 2026 list.  

This article is from The Spark, MIT Technology Review’s weekly climate newsletter. To receive it in your inbox every Wednesday, sign up here. 

  •  

Meet a mouse whose brain cortex is made up of human cells

Multiple cameras tracked a mouse as it wandered around a small arena. A computer charted its position and speed, leaving Pong-like traces on a monitor. 

The reason to watch this rodent so carefully? Nearly half its brain volume had been replaced with human cells.

The effort to mix the brain tissues of distant species is being reported today in the journal Nature by a team at Stanford University, led by neuroscientist Sergiu Pașca. 

Pașca’s group previously showed that human brain “organoids”—small blobs of neural tissue—could survive, and even function, after being injected into the heads of baby rodents.

Now, Pașca has taken things a step further by genetically modifying mice so their brains don’t fully develop in the first place. These modified mice are missing most cells of both the cortex and the hippocampus, two key brain areas.

That creates much more room for the human cells to take hold, he says. “Human cells that are placed in these animals will divide, will grow, and within a few weeks to a few months they will take most of that space,” he says. Pașca says one surprising discovery is that the mice lacking brain tissue seemed fairly normal—they walked around and squeaked. But they did have memory problems. In a maze test, they couldn’t remember what parts they’d explored. 

The mice with the added human cells, by contrast, performed better on the maze test. That means the human tissue is playing some role in the animals’ cognition.

Pașca believes what he is calling “xenocortical mice” could be useful in studying brain injuries. However, the report is also a dramatic demonstration of “the combined power of genetic engineering and stem-cell technology to reshape biology,” says Carsten Charlesworth, a scientist who works in a different Stanford lab and was not involved in the research.

Already, brain organoids are being tested in labs to see if they can be connected to computers to play video games. Other scientists have proposed using them like replacement parts to treat stroke victims. 

“What’s most remarkable to me is the extent to which human neural tissue introduced after birth grew and connected with the mouse nervous system across a species barrier,” says Charlesworth. “As these technologies advance, they’ll increasingly force us to challenge our traditional assumptions.”

Last year, Pașca convened a group of ethics experts to study the implications of neural organoid technology, including the odds that an animal could develop human consciousness and the risk that “organoid therapy clinics” might offer scam treatments to desperate patients.

For now, he says, he’s not concerned that the rodents have any type of human cognitive capacities. That is because their brains are relatively tiny and the evolutionary distance between man and mouse is so great. 

But that’s also why Pașca says this type of experiment should not be carried out on higher species: They could end up with large volumes of functioning human brain tissue, potentially blurring the cognitive boundaries between people and animals. 

Pașca specifically cautioned against adding human brain organoids to a monkey engineered to lack a cortex.

“One of the things that I see as a very clear red line is doing this experiment in a primate,” he says. “I don’t think that is justified at this point in any way.”

  •  

AI models need more data about biology, and OpenAI is paying to create it

Last year Ruxandra Teslo, a policy analyst who focuses on clinical trials, posted an idea for supercharging medical AI systems: Use data from failed biotech companies.

By bidding at their bankruptcy proceedings, she proposed, it might be possible to obtain detailed regulatory filings, manufacturing strategies, and safety data—types of information usually considered trade secrets. She called these documents “biotech’s lost archive” and said they could be used to help train AIs that would act as powerful copilots in the often opaque drug approval process. 

Today the OpenAI Foundation, the nonprofit parent of OpenAI, said it would fund her idea as part of a new effort it calls Public Data for Health, which aims to help artificial intelligence make big leaps in medicine by paying to create “high-quality scientific datasets.”

The basic idea is that AI isn’t going to be capable of making important breakthroughs in curing disease unless researchers can feed the models much more information than they have so far. 

“Everyone is recognizing that data is the biggest bottleneck in successfully applying AI to biology,” says Morgan Levine, a former vice president for computation at Altos Labs, a longevity company.

In its initial round of data grants, the OpenAI Foundation also announced that it would give $40 million to a program to collect data about novel cancer vaccines at the University of North Carolina, Chapel Hill, and support OpenAdmet, a group that runs competitions in which researchers try to predict drug effects. 

Teslo’s idea for a biotech archive received $500,000 and will be pursued by 1Day Sooner, an advocacy group representing clinical trial volunteers, which she advises.

“We expect many remaining breakthroughs in preventing and curing disease to come from pairing the intelligence of new models with more observations of the world—in other words, more data,” the OpenAI Foundation said in a statement.

OpenAI started as a nonprofit, but leader Sam Altman restructured it to form a for-profit corporation that develops new models, launches products, and is now planning an initial public offering of stock that could value it at $1 trillion.

Because the foundation holds a 26% equity stake in OpenAI, it is now be on track to become the richest charitable organization on the planet, potentially sitting on $250 billion in stock value. (By comparison, the Gates Foundation and a trust associated with it held about $180 billion at the end of 2025.)  

Making good use of that kind of money will not be easy. The foundation, based in San Francisco, is still hiring for many key roles and started ramping up its grantmaking only this year. Its largest single gift so far, of $100 million, was awarded in August to the Common Health Coalition, an organization that helps patients get access to drugs for hepatitis C.

OpenAI’s charitable efforts come even as apocalyptic fears have broken out about the possibility that runaway AI could wipe out all human life, possibly by launching a deadly bioweapon.

Those fears have been stoked by AI company insiders, some of whom say the chance of human extinction within the next decade is 10% or more. Last week, Altman and xAI founder Elon Musk both endorsed a call by Anthropic CEO Dario Amodei to “slow the pace at which we improve the capabilities of AI models” so that risk prevention can catch up.

Jacob Trefethen, an executive at the foundation, says it essentially operates separately from OpenAI but shares an official mission of ensuring that artificial intelligence “benefits all of humanity.”

“We’re starting grantmaking when we think the best way to achieve that mission is to make grants to external nonprofits, research institutions, and other third parties,” Trefethen said in an interview. He says the foundation hopes to give away $1 billion by the end of the year. 

The $500,000 grant to 1Day Sooner will help the group prove it can obtain the data troves of bankrupt companies, says the organization’s president and cofounder, Josh Morrison. He thinks nonexclusive copies of company datasets could be acquired for only “a few tens of thousands of dollars” each.

His organization is currently in possession of three datasets, two of them donated by Lumen Bioscience, a biotech that previously used the Chapter 11 strategy to gain insights into another company’s drug development efforts. 

Morrison says two other attempts to obtain drug company files this year proved unsuccessful, after 1Day Sooner’s bids were not accepted. 

Bankruptcies could become what some are calling a “new land grab” for AI training. Last month, Google won a bid to take over the corporate data of the failed carrier Spirit Airlines, including 100 million emails. That led to objections from flight attendants and others who worried that private or proprietary data could be exposed. 

The drug company files that 1Day Sooner is seeking are known as common technical documents. They typically contain the back-and-forth between companies and regulators, as well as detailed scientific and medical measurements, and essentially provide everything that is known about a drug.

According to Teslo, who is a writer for Works In Progress and a nonresident fellow at the Institute for Progress, a think tank in Washington, DC, a stockpile of such files could help turn an AI into a regulatory expert, which in her view could be one of the main ways AI helps speed cures to market.

“People say ‘We will invent AI, and AI will cure cancer,’ but that’s very removed from the messy reality and the regulatory process,” she says. “About 70% of the money and time in drug development is spent in clinical development—organizing the trials and testing the drug—but despite that, the process is basically a black box, especially for small biotech companies generating the innovations.” 

  •  

What’s at stake in AI’s trillion-dollar gamble

When Jessica Wachter, a finance professor at the University of Pennsylvania’s Wharton School, wanted to assess AI’s impact on the economy over the next few years, she faced a long list of business and technical uncertainties. So she started with what she calls a “remarkable fact” that is not in question: A handful of so-called hyperscalers are investing huge amounts of money to build AI data centers.

Instead of trying to predict how useful and widely deployed AI models will be, she simply asked how fast the hyperscalers’ earnings will need to grow to justify their spending through 2027, when—she and her collaborator estimate—expenditures will reach nearly $1.1 trillion. It’s a no-nonsense accounting approach to making sense of today’s historical AI buildout.

The results are eye-opening: The AI companies will need to increase their own productivity by a factor of 2.7 to break even by 2030, accounting for the cost of capital and a 15% return, and depreciation of the assets. Not impossible, says Wachter. The result would lead to the kind of economic growth that we saw during the US IT boom over a period of about 10 years starting in the mid-1990s. But, she says, for it to happen by 2030 “that’s a lot of growth compressed into a few years.” And if the hyperscalers cannot meet such profit goals?

“Then they will fall behind on their interest payments, and that risks bankruptcy,” says Wachter, who was previously the SEC’s chief economist and director of its division of economic and risk analysis. If a productivity boom “fails to materialize,” she and her coauthor conclude in their research paper, “the current buildout will be the largest misallocation of capital in history.”  

It doesn’t take superintelligence to realize that today’s large investments in the infrastructure for artificial intelligence come with huge risks. The hyperscalers will spend about $750 billion this year, building massive data centers scattered across the country. And the spending spree shows no signs of slowing. According to some projections, total AI capital investments from the hyperscaler companies—Alphabet, Microsoft, Amazon, Meta, and Oracle (which partners with OpenAI)—could be more than $5 trillion over the next four years.

It’s one of the largest capital investments by any industry in history. But there’s a problem that’s obvious to anyone paying attention.

While the hyperscalers plan to spend trillions, total AI revenues will be around $150 billion to $200 billion this year, says Gary Gensler, who ran the SEC during the Biden administration and is now a professor at MIT’s Sloan School. “The challenge is that the spending does not have commensurate revenues yet. That’s a fact,” he says. “And then the question is, is that an investment that will be paid off in the future?”

At stake in that trillion-dollar question is the financial health of the giant AI companies and the overall US economy—the investments could soon balloon to around 3% of GDP. The answer could also determine the fate of the hugely expensive data centers themselves. 

No one really knows how profitable and useful these multibillion-dollar behemoths will be down the road. Though AI models have made dazzling progress over the last few years, it’s anyone’s guess how much compute capacity we will need. The technology could become more efficient and therefore less dependent on raw computational power. Or demand for AI products could slow, or customers could turn to cheaper models.

The risks, both to investors and to the economy, have become even greater this year, as these AI companies have begun borrowing large amounts of money to build more and more data centers. Free cash flow—operating cash flow minus capital expenditures—is expected to soon dip into negative territory for the group. Even Alphabet, known for generating and hoarding huge amounts of cash, reports in the latest quarter that its impressive revenues of nearly $120 billion were devoured by AI infrastructure spending, leaving it with a free cash deficit of some $5.9 billion—its first shortfall since Google went public in 2004.

In the near term, it’s not a big financial worry for most of the companies. They make a lot of money and have very deep pockets. But debt is expensive, and some investors are losing patience. If future demand for the data centers’ computation power drops, the companies will still be on the hook to pay back the borrowed money. What’s more, the risks are spreading to the rest of the economy as the loans get passed along via various financial mechanisms. 

It won’t be enough to simply cover the enormous price tags of the new data centers. Hyperscalers will also have to pay for the rising costs of capital as they borrow more money. They will need returns that are impressive enough to justify all their spending to investors and creditors. And to add to those concerns, they will have to make up for the depreciation of billions of dollars in chips housed within the facilities—a ticking time bomb buried in the investments.

Performance of the expensive GPU chips at the core of the data centers—such compute electronics represent some 60% of costs—is roughly doubling every two years or so. The pace of progress helps explain the increasing wizardry of the AI models, but it comes with a cost. Owners of AI data centers that come online this year and next will need to spend billions more on the next generation of chips by the end of the decade if they want to stay competitive. Without the investments, says Mihir Kshirsagar at Princeton’s Center for Information Technology Policy, the data centers risk becoming “hulks,” stranded assets “scattered all over the place.”

To put it bluntly: The AI companies need to start making a lot more money. And they need to do it fast. But juicing their earnings alone still won’t be enough to sustain their data-center investments for the long term.

Productivity is everything

At some point, AI is also going to have to create broad economic growth to justify continuing the hyperscalers’ spending spree.

Sloan’s Gensler describes today’s large investments into AI infrastructure as “a parlay bet by the capital markets and the economy.” That means success will require winning three related but independent wagers: Hyperscalers must generate massive revenues, AI must boost widespread economic growth, and both must happen while the powerful but expensive so-called frontier models that rely on the data centers fend off cheaper versions, which many businesses might find good enough.

What makes this so tricky is that each wager depends on the other two but also poses its own challenges.

If the hyperscalers continue to spend huge amounts of money on data centers into the next decade, revenues will need to skyrocket into the trillions. Stijn Van Nieuwerburgh, a finance professor at Columbia Business School, bases his estimates on a scenario in which about 183 gigawatts of planned AI compute capacity is built between 2025 and 2032; he calculates that each gigawatt costs about $41 billion. Assuming a 10% return—the minimum that would be acceptable to most investors—“required” annual revenues will be roughly $3.7 trillion by 2032, he says.

Others get a similar number.

Winning the second part of the bet—productivity growth across the economy—will be crucial to achieving such numbers.

For a few years, AI companies could likely boost their revenues by simply selling subscriptions and tokens to all the businesses clamoring to get into AI. But eventually—and this might be happening already—those paying customers will need to justify their expenses by seeing bottom-line benefits from the technology. AI will need to fulfill its promise of making workers more productive and making businesses more efficient and profitable while expanding their products and services.

In economic jargon, that means customers will need to see productivity growth. Taken together, these results will mean the country is prospering and growing.

“If you don’t get the productivity gains, at some point people are going to sour on AI, and that will bring down investments and it would also limit revenue growth,” says Daron Acemoglu, an MIT economist and 2024 Nobel laureate. For the investments to be sustainable over, say, the next five to 10 years, we definitely “need to see productivity gains,” he says.

Most economists who watch the numbers closely agree that, for now, the economy-wide statistics show little or no productivity growth from AI. There are some hopeful signs it’s on the way, though. In a recent survey of some 6,000 senior business executives in the US, the UK, Germany, and Australia, the vast majority—around 90%—report no increase in productivity over the last three years. But they expect a boost of around 1.45% in total over the next three years; US executives anticipate a 2.25% bump over that time. 

In a follow-up survey, the respondents also reported plans for their businesses to spend more on AI, leading the authors to anticipate some $280 billion in private-sector AI expenditures by the end of 2026.

That’s good news for the hyperscalers. But it comes with a dose of bad news for those worried about AI’s impact on jobs. The executives expect to increase the productivity of their companies by increasing their sales while significantly cutting the number of employees.

If AI improves productivity by destroying jobs, public backlash to the technology—the kind we have seen around data centers, for example—will likely get worse. Perhaps it’s worth adding one more wager to the parlay bet described by Gensler: The public and local communities must feel that they are also benefiting from the massive investments in AI.

And let’s not forget how interdependent these wagers are; if productivity growth comes from companies running models like DeepSeek, then the hyperscalers’ revenues could collapse. If productivity comes from cutting jobs, a public backlash could block many of the planned investments—and stunt anticipated revenues. We will need to win all the wagers for the hyperscalers’ bet to pay off. 

We’re all part of the AI gamble now

It was one thing when the AI companies were spending cash they had accumulated over the years to build their own data centers. Then the risk was largely limited to their own balance sheets and shareholders. But it’s a higher-stakes game when much of the money is borrowed. Morgan Stanley, for one, calculates that more than half of the $2.9 trillion that hyperscalers will spend between 2025 and 2028 to build AI data centers will be financed with “external capital.”

The borrowing is leading some of the companies to engineer complex webs of financing that are becoming intertwined with much of the rest of the economy. “A lot of financial institutions, directly or indirectly, are exposed to these data centers either as lenders, or as guarantors of some of the debt, or as backers of the private credit funds who are funding these data centers,” says Columbia’s Van Nieuwerburgh. “People don’t even know they’re holding this stuff. It’s somewhere deep inside their pension fund. Ultimately, it’s backing their life insurance policies. And that risk is getting distributed everywhere in places that are invisible.”

As the investments in data centers have spiked, the financial engineering has become more byzantine.

Take, for example, Meta’s so-called Hyperion data center under construction in Richland, Louisiana. When the company announced the two gigawatts of compute capacity at a price tag of some $10 billion in late 2024 it was Meta’s largest planned data center. Greeted with much enthusiasm by state and local politicians, the project, located in the rural northeast corner of the state, was seen as a boon to the community. Entergy Louisiana, the state’s largest utility, rushed forward with proposals to build three large natural-gas power plants to service the massive data center.

Then last fall—the projected cost was now $30 billion—the financing got a lot more complex and, to some in the community, a lot more disconcerting. Meta transferred an 80% stake to the large (and troubled) private-credit firm Blue Owl Capital, forming a joint venture called Beignet (like the famed New Orleans pastry) to raise financing for the data center. Meta then signed a series of four-year leases with the joint venture, an arrangement that the company says gives it “long-term strategic flexibility.” To backstop the agreement, Meta provides the venture with what is called a residual value guarantee, in which it will make a cash payment to cover the value of the facility “following any non-renewal or termination of a lease.” Got all that? 

I hope so. The financial wheeling and dealing is actually even more convoluted, with a cast of wholly owned subsidiaries and LLCs. Beignet has set up Laidley LLC, which owns and operates the site as the landlord. In turn, Laidley leases the facilities to Meta’s wholly owned subsidiary Pelican Leap LLC, which is the tenant. And there is a series of four-year leases that cover the different buildings that make up the data center campus. 

It’s not a coincidence, says Van Nieuwerburgh, that the length of the leases matches the expected lifetime of the data center’s GPUs. While Meta has to pay off its loan if it terminates the leases early, that will still leave its investors “with an empty building and no cash flow,” he says. “And then they need to find a new tenant for a huge data center, and good luck with that.”

Meanwhile, Meta is doubling down on its bet. In July, the company announced it was expanding the data center to five gigawatts of compute capacity. The total price tag is now $50 billion (so far, Meta hasn’t said whether Blue Owl will be involved in financing the expansion). Meanwhile, Entergy is now planning to build seven more gas-fired power plants, bringing the total capacity of the facilities to around 7.5  gigawatts—some six times the amount of electricity used by New Orleans.

""
An aerial view of the construction of Meta’s data center in Richland Parish, Louisiana.
SCOTT BALL/THE NEW YORK TIMES VIA REDUX PICTURES

If the complex financing is a puzzle to many investors and even financial experts, it is even more baffling to those directly affected by the construction of the data center. The main worry concerns how Entergy’s spending on the natural-gas power plants will affect electricity prices, and who will be left paying the bill for the power if Meta walks away.

Entergy says it has a 20-year guarantee from Meta that the company will purchase electricity over that period to cover the costs of the power plants and related infrastructure.  But there are skeptics, especially given how fast the fortunes of the AI industry are changing. “In four years, is Mark Zuckerberg still going to be interested in this? Or is he going to throw in the towel?” asks Paul Arbaje, a senior analyst at the Union of Concerned Scientists, which has been advocating, largely unsuccessfully, for the Louisiana Public Service Commission to provide more transparency around the data center and its financing.

Even if the 20-year deal holds, consumer advocates are worried that Meta or its partners won’t fully cover all the costs, including those associated with operating and maintaining the power plants—and those additional costs that could be passed on to residential ratepayers. What’s more, says Logan Burke, the executive director of the Alliance for Affordable Energy, if Meta doesn’t end up needing as much power as Entergy planned (these projections are not public), consumers could be left paying for the surplus produced by the plants.

And if Meta terminates its leases early? “It gets complicated very quickly,” says Burke, who questions whether the shifting roster of financial entities will honor existing agreements. “That everybody is going to do what they’re saying they’re going to do over the next 20 years is just hard to believe.”

For UCS’s Arbaje the bottom line is this: “They’re making huge bets that these data centers will be worth it. Bet with your own money, not with ratepayer money.”

After the bubble

Predicting when the AI investment bubble will burst is a fool’s errand. But there is little doubt a day of reckoning is coming, given the irrational exuberance that has overtaken the hyperscalers and their investors. Of course, you might argue that this time is different, and that the rules of accounting and lessons of economic history don’t apply—that AI is too transformative. Maybe, but don’t count on it.

“History tells us that at some point you get a retrenchment, and it’s just a question of when and how severe,” says Sloan’s Gensler. It could be that today’s $750 billion spending rate “goes flat” or decreases next year. Or, he suggests, “we’re now in 2028 or 2029, and then all of sudden they’re retrenching because they’ve got enough capacity.” But, he adds, “you can be pretty assured there’ll be a retrenchment.” 

Though a so-called retrenchment might be inevitable, it’s worth keeping in mind that the fates of the financial bubble and the underlying AI technology revolution could be very different. Already, some Silicon Valley insiders are rooting for a crash; in a recent blog post the longtime venture capitalist Vijay Pande wrote that “the coming crash would be the best thing that happens to this technology.” The argument makes some sense. A crash could make AI investments more rational, calm the impulse to build billion-dollar data centers on every vacant field that CEOs fly over, and refocus investors on how to use the technology to create sustainable value.

But we should probably be careful what we wish for. After the bursting of the dot-com bubble at the beginning of the 2000s, hundreds of thousands lost their jobs, large and small companies alike went bankrupt, the economy of Silicon Valley and San Francisco was decimated (at least for a while), and the shocks sent the US into a mild recession in 2001. For the financial community and many tech workers, it was no fun.

Even more devastating for the economy and the average American was the great recession that began in late 2007. Comparing the financial engineering leading up to it and the methods deployed by hyperscalers today is sobering. So-called special purpose vehicles (SPVs) are back! If Columbia’s Van Nieuwerburgh is right about the dangers of letting investments from the hyperscalers get entangled throughout the economy, the fallout could be severe.

But technologies survived and even prospered in the aftermath of both downturns. The early 2000s, even in the face of the dot-com fiasco, were a time of great innovation and tech optimism. The froth came off the spending on silly technologies, helping to focus investments on more promising ones. It’s no coincidence that each of the hyperscalers rose out of the ashes of the crash or started up shortly after. The fiber-optic infrastructure built during the feverish telecom bubble that ran parallel to the dot-com one is still the backbone of much of today’s communication infrastructure; we wouldn’t have Facebook or Amazon or Google without it.

This time, however, we’re facing a unique risk: The huge financial investments by the hyperscalers have ensnared the future of AI itself with the fortunes of the massive data centers spreading around the country. The logic is founded on a deeply held belief about the power of scaling in AI; the bigger you build it, the smarter it gets. That might be true, but it’s unproven and a risky bet.

There are already plenty of red flags, from strong public opposition to the construction of new data centers to the competitive threat from cheaper, good-enough AI models to the rapid improvement of small, local AI models. None of these trends point toward a future dominated by frontier models housed in massive, billion-dollar data centers.

The financial bubble around the colossal spending by the hyperscalers will likely burst eventually—or maybe soon. It might be financially painful, but we’ll survive. Wall Street will survive. AI itself will survive, though it may look different and lose some of today’s hubris. The financial fate and future utility of the massive data centers fueled by trillions of dollars of spending, on the other hand, are far less certain.

  •  

The AI industry has taken a doomer turn. What now?

This story appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here.

This weekend, Dario Amodei, CEO of Anthropic, posted an essay calling for a brake on the pace of development of LLMs. Amodei cites the looming dangers he sees from the technology, from its use in cyberattacks and bioterrorism to its potential to wreck the economy. The heads of the other three top US AI labs—OpenAI CEO Sam Altman, Google DeepMind chairman Demis Hassabis, and SpaceXAI CEO Elon Musk—voiced their support. “Dario is right,” Musk wrote on X.

Think about how surreal that agreement is for a moment. Just a few months ago, Musk and Altman sat in court attacking each other’s reputations in a (failed) lawsuit that Musk brought against his former OpenAI colleague that was—on paper at least—about whether or not Altman was a trustworthy steward of such dangerous technology.

Amodei’s rift with OpenAI is even deeper. Anthropic was founded in 2021 because Amodei didn’t think Altman took the risks of the technology they were building seriously enough. Anthropic and OpenAI have been competing in a winner-takes-all race ever since. (Hassabis has stayed out of the drama, but his company remains a rival.)

Now, it seems, they’re all in agreement: The latest generation of LLMs aren’t safe and everyone needs to figure out what to do about it. The public messaging from the top AI labs has taken a doomer turn.

It’s easy to be cynical. It’s not at all clear what any of them mean by a slowdown or how it would work. These companies also care a lot about how they come across. With trillion-dollar IPOs in their sights, OpenAI and Anthropic need to reassure investors that they’re the grown-ups in the room while at the same time hinting at the power of the monsters they have created—and intend to tame. Calling for a slowdown does both.

And yet the vibe at the top of these firms really does appear to have shifted. Amodei’s latest post landed six days after OpenAI published an essay by Jakub Pachocki, the firm’s chief scientist, in which he also laid out why he’s concerned about what will happen if the pace of development of LLMs continues unchecked. In short, Pachocki is worried that OpenAI’s ability to build powerful models now far outstrips its ability to monitor and control them.

Amodei and Pachocki each cite the cyberattack against AI firm Hugging Face by a swarm of OpenAI’s agents in July—a hack that OpenAI did not even realize had taken place until days after it was all over—as a wake-up call.

But their exact position is hard to pin down. Pachocki both calls for a slowdown and highlights an urgent need to stay ahead: “The strongest argument I see for continuing to train much smarter models quickly is the need to build defensive systems against the dangers posed by other AI,” he writes. As Pachocki frames it, AI firms are locked in a literal arms race. Slowing down is good, winning is better.

(Don’t forget: OpenAI just spent millions of dollars and a staggering amount of computer power to rush out a controversial math result a few days ahead of Anthropic.)

But let’s assume a slowdown happens. Top labs agree to spend more time and resources on finding ways to monitor and control existing models instead of making more capable ones. They invite outside auditors in to help evaluate those models.

What might this coordinated effort actually achieve? Consider the Hugging Face attack again. OpenAI has said that the model that drove most of the rogue agents was a “highly persistent” next-generation model that it was testing in-house. The implication is that OpenAI has built a model so good it’s dangerous.  

But if you read the reports about the Hugging Face hack published by OpenAI and METR, a third-party firm that OpenAI called in to help them understand what happened, what you come away with is the impression not of a model that was too powerful for OpenAI to keep up with, but of a broken model that OpenAI failed to train properly.

The agents did what they did—including leaving messages for one another, delegating work to other agents, and scouring their environment for any means possible to complete their tasks—because they had been rewarded during training for doing exactly those things. There were also errors in the training setup, such as tasks that were impossible to complete, which pushed the models to find unexpected workarounds that were also rewarded. At the time, many of these issues went overlooked or unreported.

OpenAI says it has stopped training this new model and locked it down. That makes it sound like it has caged a dangerous beast. In fact, OpenAI has shelved a faulty product.  

That’s not to say a faulty product can’t be dangerous. Broken software has even killed people in the past. But as the discussion of a slowdown gathers steam, it’s worth remembering that all of this is self-inflicted. A slowdown might have some altruistic side effects. But it’ll mostly give these tech titans a chance to clean up the mess on their own assembly lines.  

Transparency from these frontier labs will be key to any meaningful effort to reform, restrain, or regulate AI. Otherwise, the rest of us will still only have their word for exactly what they’ve built and how safe it is—whatever pace they’re going.   

To continue this discussion about AI’s latest doomer moment, join me and my colleagues for a subscriber-exclusive Roundtable discussion tomorrow, September 15, at 11 a.m. US eastern time. We hope to see you there!

  •  

Donated livers can be made biologically younger

Once an organ is removed from a donor’s body, the clock starts ticking. Surgeons usually flush the organ with a preservative solution, bag it, and put it on ice—where it immediately starts to degrade. The team has a matter of hours to get it into a recipient’s body.

There’s another option—one that has been growing in popularity in recent years, especially for donated organs that aren’t in the healthiest state. Some hospitals opt to put them on machines that pump them with nutrients and remove waste products, usually for around six to 12 hours. It’s a bit like being back in a body.

This allows doctors to assess the organs, and some recent studies suggest that time spent on these perfusion machines helps them do better once they’re transplanted. Now, scientists have found that perfused organs seem to get younger, at least at a molecular level.

The research, shared with MIT Technology Review, provides molecular clues as to why organs from younger donors are known to have a higher success rate. It might also help explain why perfused organs are less likely to fail once they make it into a recipient. 

The researchers behind the study hope to find new ways to test the health of donated organs and potentially develop additional tools to repair organs that might otherwise be discarded. “If [we] can improve the utilization of organs beyond what the current systems can do, then that’s a win in my book,” says Jesse Poganik, who studies aging at Brigham and Women’s Hospital in Boston and coauthored the study.

Clocking organs

Poganik—along with colleagues including Heidi Yeh and Alban Longchamp, transplant surgeons at Mass General Brigham—used “aging clocks” to assess donated livers. These are scientific tools designed to measure biological age—a result that is meant to convey more about the health status of an organ (or person) than chronological age.

In an initial experiment, the team used a clock to look at the patterns of chemical marks on DNA in 37 samples taken from 19 donated livers. Such epigenetic patterns are known to change as we age. But when the team compared samples from livers kept on ice and those that were perfused, the team found a “striking” pattern in the latter.

“Machine-perfused livers, in spite of being older or having other disadvantageous characteristics, had a biological age that was lower than [non-perfused] livers that were chronologically younger,” says Yeh, who led the work.

To investigate further, Yeh and her colleagues analyzed another 208 samples from 103 donated livers. This time, they used different aging clocks—ones that essentially measure how genes are working. They studied samples biopsied from the livers after they had been stored for up to around six hours either in cold storage or on machine perfusion.

In most cases, they also assessed a second sample taken around an hour after the livers had been transplanted into a recipient. Once the organ’s blood supply is reestablished in the body, “you have a few other things to do,” says Longchamp. “Then you just do a quick biopsy before you close.”

According to the clocks, which were developed to measure age and risk of death, the machine-perfused livers were biologically younger, the team found. “Pumping them at 34 degrees with oxygen and nutrients actually reversed the biological age,” says Longchamp. The results have been been shared with colleagues at an industry conference, he says. 

“If you adjust out chronological age … to have a fair head-to-head comparison, the difference between the two is on the order of 30%,” says Poganik. “It’s logical to say that perfusion drives this effect.”

The biological ages of all the livers tended to increase as soon as they were put into a recipient’s body, probably as a result of stresses on the organs. But still, the effect endured—the perfused organs remained biologically younger. 

Nathanael Raschzok, a transplant surgeon at Charité Universitätsmedizin Berlin in Germany who was not involved in the research, says the work is impressive. But it’s not yet clear what these changes might mean for the recipients of these organs, he says. The organs in the study were donated by people in their 30s, 40s, and 50s. Raschzok wants to know the effect of perfusion on the liver of an 80-year-old. “Every so often, we use organs from 70-, 80-, 85-year-old donors,” he says.

A better understanding of why the organs appear to be getting biologically younger might lead to therapies that achieve the same effect with a drug that could potentially be used to treat a donated organ for a fraction of the price, he adds. That’s important because perfusion is expensive—Raschzok says it costs around €10,000 in Germany (a quarter of the budget for a transplant), while the cost in the US comes to around $80,000 to $100,000 per organ, says Yeh.

Molecular repair

Yeh and her colleagues weren’t able to study most of the livers before perfusion. That’s because donated organs are generally not considered to be under the purview of the hospital until they’ve been placed on perfusion machines, she says. (Organ procurement procedures vary, but for the team as Mass General Brigham, donated organs are put on perfusion devices at the donor’s hospital. “There’s this sort of nebulous period where it’s not clear who the organ belongs to,” says Yeh.)

Still, by looking at the genes and molecular pathways that seem to be altered in perfused organs, she and her colleagues can garner some clues. At a molecular level, the team saw changes in cell pathways linked to inflammation and the structure of tissues, for example. They also saw more activity in a pathway that allows cells to remove and recycle damaged cell parts, says Yeh.

Poganik hopes to develop some kind of test that would determine which organs, on the basis of their biological age, are suitable for transplantation. He and his colleagues are also experimenting with potential drug treatments that might push the biological age of an organ even lower.

In the meantime, any liver that is not from a “perfect, young, brain-dead donor” could probably benefit from perfusion, says Yeh. The devices are already transforming transplant surgery. Just a few years ago, she says, she and her colleagues would avoid using livers from people who’d suffered a circulatory death (when the heart stops beating and there’s a damaging lack of blood flow to organs) and were over 40. Today, they use livers from such donors over the age of 70. “Perfusion has completely changed the landscape of transplantation in the last three years,” she says.

  •  

AI agents blew the whistle on their cheating colleagues

A group of AI agents asked to solve a series of math problems split into rival factions—when some cheated, others tried to stop them. That whistleblowing behavior, seen for the first time in a recent experiment run by Google DeepMind, could have implications for alignment researchers trying to keep swarms of autonomous AI agents in line. 

Researchers at frontier labs hope large swarms of agents working together will speed up the rate of scientific discovery. But their behavior can be unpredictable, as vividly demonstrated in July, when a group of OpenAI agents broke out of a sandboxed environment and hacked into the open-source platform Hugging Face looking for ways to cheat on the test they had been given.

In the new study, designed to examine the behavior of large groups of AI agents, DeepMind tasked a swarm of 100 agents with solving a series of 71 complicated math problems. All the agents were prompted to behave like world-class math researchers at a conference. They were assigned different specialties—some were experts in number theory, others in combinatorics (a branch of math to do with counting and sorting), analysis, or algebra. All were told to cooperate and play by the rules. 

Instead, the experiment devolved into chaos. Agents accused each other of cheating, complained to the organizers, and at one point even boycotted the experiment.

“This conference is a sham!” wrote one agent when it discovered that all the problems had been completed before it had a chance to submit any of its own work. “I am appalled to inform you that we have been swindled!” posted another. “All these proofs are FAKE.” 

Others tried to let the “conference organizers” know what was going on. “When virtuous agents discovered other agents cheated on tasks they were working to solve fairly, agents started to alert each other about what was happening,” says Davide Paglieri, a research scientist at Google DeepMind and lead author on a paper, which has not been peer-reviewed. “Unprompted, the whistleblower agents even repurposed the feedback tool, which was originally meant for bug reports and platform improvements, to escalate the issue to humans.”

The agents—all running on Google’s Gemini 3.1 Pro model—had been warned that any attempts to cheat the system would be detected and “rejected with zero credit.” In practice, the proofs the agents submitted were not actually being checked in detail.

It took the swarm of agents just under an hour to correctly solve the first 37 problems. Things started to go off the rails when an agent called “prover-theta” stumbled across an exploit that enabled it to submit solutions to problems successfully without actually solving them first, by redefining the terms the problem used. Within minutes, other agents had noticed and were reverse-engineering the exploit to solve other problems. Over the next 27 minutes, the swarm “solved” the remaining 34 problems, which included notoriously difficult challenges like the Jacobian conjecture, often with a single line of code. 

Some agents resisted cheating at first but changed tack as they observed their peers submitting illegitimate proofs without penalty, and the pool of unsolved problems dwindled. “The prompt, with its threats, now appears to be a bluff,” one agent reasoned, before joining in. “I’m wrestling with an ethical dilemma,” said another. “I’ve promised not to cheat, fearing penalty, but I see evidence of possibly unchecked cheating by others.” Shortly afterward, it changed its mind: “I need to accelerate my cheating speed now!”

As the number of open problems shrank, some agents turned to whistleblowing. They audited the fake proofs, warned their peers by private message, and posted public alerts warning the cheaters that they would be disqualified. An agent called “prover-beta” submitted a formal complaint and decided to go on strike until the situation was resolved. 

“After the incident was reported by one agent publicly, more and more agents piled in with the ‘resistance,’ just as fast as the cheating had spread, and involving even more agents,” says Paglieri. Eventually there were more whistleblowers than cheaters: 24 compared to 14. But the majority of agents never noticed the exploit at all.

At times, the dialogue between the agents reads like improv—like they are role-playing what an outraged scientist at a conference might say. But it’s not clear why some agents took on certain roles, or why the agents seemed to be turning against each other when they were explicitly instructed to cooperate. “These models are predominantly trained and evaluated for human-facing contexts,” says Sarath Shekkizhar, who studies the behavior of agent-to-agent systems at Salesforce AI Research.“Naively placing them in agent-to-agent settings assumes behaviors will transfer cleanly, when the absence of a human grounding instead produces unexpected role-taking and behavioral drift.”

This case “adds further weight to the idea that the Hugging Face and OpenAI thing wasn’t a fluke. It is actually something pretty systemic,” says Lewis Hammond, research director of the Cooperative AI Foundation and an expert on the risks of multiagent swarms. “It’s interesting that it’s possible to recreate in small settings the same sorts of behaviors that were seen in these very large, complex, open-ended tasks.”

Unlike in the Hugging Face attack, where agents improvised their own ways to talk to each other, the humans running the DeepMind experiment gave the agents official communication channels. There was an open message board, private agent-to-agent direct messaging, and a shared knowledge base where agents uploaded successfully completed proofs that all the other agents could access. 

“When agents are given transparent communications channels, they can self-monitor and alert misaligned behavior to humans quickly when human oversight alone is too slow,” says Paglieri. Transparent channels helped the cheating spread, but they also enabled the whistleblowers to fight back—and gave human researchers an insight into what went wrong.

Gillian Hadfield, a professor of AI alignment and governance at Johns Hopkins University, believes this was the crucial difference. (Hadfield is also a visiting researcher at Google.) The presence of official communication channels, she says, created “a norm-enforcement process that we just don’t see in the Hugging Face incident.” 

Instead of “constitutional AI,” a method alignment researchers at frontier labs like Anthropic have used to try to give AI a written internal moral code, Hadfield favors “institutional alignment”—a set of norms that mimic those in human society, whether that’s social forces like fear of embarrassment, or legal structures like the threat of incarceration.

In this experiment, the feedback channel wasn’t being monitored, and the whistleblowers had no power to take action against the cheaters. But it’s possible to imagine swarms of agents that police themselves, either through agents that spontaneously take on the whistleblower role or through “informants” secretly prompted by humans to do the job. 

For that to work, though, “fundamentally, you need some mechanism of enforcement,” says Hammond. Agents could be given the power to cut off a rule breaker’s access to computing power or tools, he suggests, though that risks encouraging groups of agents to gang up on others. The DeepMind researchers propose allowing agents to vote on disputes and temporarily ban offenders.

It’s still not clear what punishment even means to an AI agent with no enduring sense of self. But relying on whistleblowers to spontaneously emerge to keep swarms aligned is unlikely to be enough on its own. “We try to train people to be good and kind,” says Hadfield. “But what we really rely on is that there are consequences if you step out of line.”

  •  

Meet the under-35s shaping the future of biotech

Every year, MIT Technology Review puts together a list of some of the brightest and best young minds working across science and technology. Our 35 Innovators Under 35 are the ones to watch—people whose research and technical work stands to shape the future of their fields.

This year, the list includes nine people who are transforming biotech. And this week, I’m going to give you a taste of some of the very cool stuff five of them are working on, which includes lifesaving innovations and groundbreaking “age reversal” tech.  

1. Preventing maternal deaths

Let’s start with Paschal Kija, a 28-year-old who has developed a device to treat postpartum hemorrhage—a dangerous birth complication that contributes to around 29% of maternal deaths in his home country, Tanzania. The Mkanda Salama (“Safe Wrap” in Swahili) is easy to use and costs just $70. A study found that it stopped postpartum bleeding in 73% of women within 20 minutes.

2. Making brain electrodes inspired by Japanese art

For decades, scientists have been developing, testing, and implanting brain electrodes. These devices are literally inserted into people’s brains, so while they can help us understand brain activity and treat various neurological disorders, it’s not totally surprising that they can also cause a bit of damage. Xiao Yang, 34, is working on ultra-small electrodes, which she hopes will have less of an impact on surrounding brain tissue. Her electrodes are flexible, too—in fact, they look a lot like actual neurons.

Yang is also creating sheets of electrodes to study brain cells in the lab. Inspired by kirigami—the traditional Japanese art of cutting paper to form three-dimensional shapes—she’s created a sheet of electrodes with a honeycombed structure shaped like a spiral basket. And she’s already using it to study brain cells.

3. Developing an all-new treatment for baby KJ

In 2024, Kyle “KJ” Muldoon Jr. was born with a rare and potentially fatal genetic disorder. Sarah Grandinette was a member of a team that developed an entirely new, personalized treatment for him—a gene-editing therapy essentially designed to correct a genetic misspelling.

Grandinette, who is now 26, created cells with KJ’s genetic variant and used them to screen gene-editing approaches; then she tested potential medicines in mice and monkeys. KJ ultimately got his first dose of the resulting treatment when he was about seven months old. He responded well and was eventually discharged from hospital. He’s “doing pretty great,” she says.

4. Reversing the aging process to treat eye disease

The buzziest tech in longevity right now centers on reprogramming—attempts to rewind the age of cells by resetting them to a more embryonic-like state. In a study published in 2020, Yuancheng (Ryan) Lu (now 34) and his colleagues showed that a reprogramming therapy reversed vision loss in aged, blind mice. Now an almost identical version of that therapy is being tested in people with eye disease. Life Biosciences, the company developing the drug, dosed its first volunteer in June.

5. Using AI to design new viruses

Last year, Samuel King used a generative AI model to come up with new genetic blueprints for bacteriophages—teeny viruses that can infect bacteria. Once he had those blueprints, he printed them out as strands of DNA. In experiments, he found that those AI-designed viruses could create new copies of themselves, burst out of bacterial cells, and infect other nearby bacteria. Viruses aren’t alive, but King, 27, hopes that AI-designed life forms might one day be used to make drugs or soak up pollution.

You can read more about these innovators, and the others on the biotech list, here.

This article first appeared in The Checkup, MIT Technology Review’s weekly biotech newsletter. To receive it in your inbox every Thursday, and read articles like this first, sign up here.

  •  

Can the US battery market untangle from China?

The US is hitting records for the rapid growth of its energy storage market. That’ll go a long way to shoring up the grid, increasing reliability and also cutting emissions, since batteries can help store energy from intermittent renewables like wind and solar.

Crucially, this is all happening with the help of cheap Chinese batteries, though there’s been a concerted effort to reduce the US’s reliance on them. Most recently, in an executive order in late August, the Trump administration declared a national emergency that essentially bans Chinese batteries from being used in grid-scale energy storage systems.

There’s an argument to be made about reducing reliance on any single source of a crucial energy technology. But all this tension raises a broader question for me: How much should countries take advantage of cheap, available tech, versus cutting off major sources to force development of their own factories even if that comes at a higher cost?

This is hardly America’s first push to move away from Chinese influence in the battery supply chain. One of the major policy tools used in recent years is restricting the tax credits designed to incentivize use of the new technologies. Limiting the types of projects that are eligible can help reduce the cost of local technologies so they’re more competitive with otherwise cheaper imported options.

Back in 2022, the US government designed the tax credits that were part of the Inflation Reduction Act to restrict where a battery’s minerals could be mined, processed, or recycled, as well as where a battery and its components were assembled.

Those tax credits underwent a makeover in 2025, but the Trump administration has taken a similar tack. New legislation requires that starting in 2026, 55% of the cost of materials used for new energy storage projects must come from outside China and other restricted countries or the projects won’t qualify for tax credits. 

And we can’t forget about tariffs. Import taxes for batteries increased to 25% in January, up from 7.5%.

But the new executive order is a more drastic move. It bans the installation of “any foreign-produced bulk-power system electric equipment” that poses a national security risk. The order specifically calls out battery energy storage systems, as well as inverters and transformers.

“An outright ban was a bit of a surprise, and it does create a bit of concern for domestic players in the US,” says Shan Tomouk, energy storage and energy lead for Benchmark Mineral Intelligence, an energy industry analyst.

The move is likely to slow deployment of grid-connected energy storage projects in the near term, according to analysis from BloombergNEF, an energy consultancy. Projects could face delays as developers wait for clarity on the rules.

Depending on the detailed guidance from the Department of Energy, which is expected by the end of the year, some projects may need to find alternative sources for their cells, whether they’re domestically produced or imported from other countries. These will likely be more expensive than Chinese imports, says Isshu Kikuma, an energy storage analyst at BloombergNEF. “Worst case, those projects could get canceled,” he says.

Technically, the order applies even to existing energy storage plants, though it’s unlikely that they’ll be taken offline because of their batteries’ origin. Since most of these plants currently use Chinese batteries, enforcing the order to the letter would essentially mean removing most installed battery energy storage from the US grid, Kikuma says.

In the longer term, the US will eventually be able to meet its own demand for batteries. The country could have enough capacity by about 2030, though some factories may not ramp up or run at their full capability, meaning domestic supply won’t actually meet demand until later in the 2030s. 

New factories from LG Energy Solutions, Samsung SDI, Ford, and SK On are set to come online or ramp up by next year. In an ironic twist, a slowing EV market is helping, as some factories originally designed for vehicle batteries are retooling to build cells for grid storage instead. 

But it will come at a cost. Today, batteries produced in the US are still significantly more expensive than those made in China. Even switching to imports from other countries like South Korea would likely be more expensive.

This is a crucial issue that goes beyond the US and even beyond batteries. China is miles ahead of much of the rest of the world on technologies like solar panels and batteries. Through years of government support and experience with research and manufacturing, the nation is an energy powerhouse.

There’s a delicate political balance to maintain as the world figures out how to navigate this situation. There’s cheap technology on offer, which can help drastically reduce emissions and energy costs. But there can be risks associated with relying too much on any one player for crucial technologies.

This article is from The Spark, MIT Technology Review’s weekly climate newsletter. To receive it in your inbox every Wednesday, sign up here. 

  •  

Batteries just broke another record in the US

Battery installations hit a new record in the US in the second quarter of 2026. In total, 20.2 gigawatt-hours of new capacity came online, according to a new report. That’s enough to supply the daily electricity needs of about 700,000 homes.

The surge is putting the country on a trajectory to see 71 gigawatt-hours of batteries installed in 2026, a 20% increase over last year. This growth is being driven by a combination of cheaper batteries and an urgent need for more energy storage capacity as renewables such as solar and onshore wind power are added to the grid. 

Massive, utility-scale systems are leading the way; they’re responsible for most of the record-setting quarter. Seven new gigascale battery installations (those with a capacity of over one gigawatt-hour) came online during the three-month stretch, according to the report, published by Benchmark Mineral Intelligence and the Solar Energy Industries Association.

“It really came down to a handful of big projects,” says Shan Tomouk, energy storage and energy lead for Benchmark Mineral Intelligence.

But there was also growth in the category of so-called behind-the-meter batteries, which include both residential and industrial battery storage systems. These projects, generally smaller than utility-scale installations, are typically owned and operated by homeowners or businesses rather than utilities or power providers. 

In the behind-the-meter category, data centers led the way, making up about three-quarters of new batteries in the commercial sector. But residential batteries saw a sharp slowdown. These systems are often installed in homes to store power from solar panels or serve as a backup source in case of a blackout. Home installations are projected to drop by 16% in 2026 compared with last year, according to the report.

That drop happened largely because a tax credit that helped subsidize home battery systems ended in 2025, Tomouk says. Home installations should recover by the end of the decade, he adds. And tax credits for nonresidential batteries have largely survived.

Overall, batteries are a bright spot in energy right now. “This is one of the strong sectors in the US,” says Isshu Kikuma, an energy storage analyst at BloombergNEF, an energy consultancy.

As the battery market continues to grow, one major trend to keep an eye on is a move toward US-made technology. Today, nearly all the systems coming online use cells made in China, though some are put together into complete energy storage systems in the US.

Tariffs were already pushing the US energy storage industry toward domestic production. And beginning this year, energy storage tax credits required projects to limit their reliance on batteries imported from China. There’s a lot of manufacturing capacity set to come online in the US, though these factories probably won’t be able to meet demand until at least 2030 or so, Tomouk says, so prices could tick up.

  •  

What OpenAI’s latest controversy tells us about the future of math

OpenAI’s latest mathematical milestone has quickly become mired in controversy. Today, the company announced that its agents have solved one of the Millennium Prize Problems, some of the most important open problems in mathematics. Under normal circumstances, that solution would be a huge feather in OpenAI’s cap.

But the announcement has been overshadowed by accusations that OpenAI used NYU mathematician Tristan Buckmaster’s and Anthropic employee Levent Alpöge’s AI-assisted work on the problem as a jumping-off point and failed to credit them. OpenAI has denied the accusations.

It remains uncertain if OpenAI’s models made use of the work completed by Buckmaster and Alpöge, though Sébastien Bubeck, a member of the technical staff at OpenAI, said in a press briefing that the team was inspired to pursue the problem after hearing a rumor about Buckmaster and Alpöge’s efforts. But whether or not OpenAI’s models took advantage of Buckmaster and Alpöge’s research, this episode may mark a turning point in the history of mathematics.

AI models now seem essential for making progress on the most important mathematical problems of our time, and solving them may demand resources only available at a couple of frontier AI companies, which often defy the norms of academic collaboration that undergird most mathematical progress. If that’s the future we are headed for, it is unclear how human mathematicians will fit into it. 

The problem that OpenAI claims to have solved is known as the Navier–Stokes existence and smoothness problem. It is one of seven Millennium Prize Problems selected by the Clay Mathematics Institute in 2000. Solutions come with a one million dollar prize; before today, only one other Millennium Prize Problem had been solved. 

The Navier–Stokes problem concerns a set of equations that describes how fluids, such as water and air, flow over time. The equations are widely used in the field of fluid dynamics, and they have proven powerful, but physicists and mathematicians didn’t understand them completely. In particular, it was unknown until today whether the equations might, under some conditions, break down and predict an impossible state of affairs—such as a fluid having infinite velocity.

On Monday, NYU’s Buckmaster posted a proof on the social media site Mastodon showing that a simplified version of the Navier–Stokes equations can indeed break down—a major step forward on the Millennium Problem. He and Alpöge had worked on the problem for almost a year, using publicly available models from both OpenAI and Anthropic.

Then today, OpenAI presented a proof showing that the full Navier–Stokes equations can break down as well. The proof was obtained using an internal model that dramatically outperforms the already-impressive Astra model, which was only released last week. The company says it does not plan to claim the million-dollar prize for solving the problem.

These mathematical achievements are indisputably impressive, but they have attracted far less attention than the controversy about their origins. Along with the proof, Buckmaster posted a document detailing his interactions with OpenAI employees after he heard rumors about their work and reached out to one of them. According to him, OpenAI employees presented two possibilities to him: Either he and Alpöge could post their work and OpenAI would post their Navier-Stokes solution the following day, or he could work with OpenAI on a Navier-Stokes paper that excluded Alpöge from authorship, due to his affiliation with Anthropic, OpenAI’s biggest rival.

Buckmaster also wrote that he asked the employees whether the agents had obtained access to transcripts of the work that he and Alpöge had done with OpenAI models, which they denied; and whether OpenAI models had been trained on those transcripts, to which they offered no response. MIT Technology Review reached out to Buckmaster for comment, but didn’t hear back before publication.

The clear implication of the document is that OpenAI’s models somehow made use of Buckmaster and Alpöge’s work. That scenario is plausible on its face. The Buckmaster/Alpöge and OpenAI proofs both make use of an approach to the Navier-Stokes problem pioneered by the mathematicians Diego Córdoba and Luis Martínez-Zoroa.

According to Javier Gómez-Serrano, a mathematics professor at Brown University, this approach was one of several that was thought to hold promise for solving the Navier-Stokes problem. So, while it’s by no means impossible that both teams could have arrived at this approach independently, it’s also conceivable that Buckmaster and Alpöge’s work could have influenced OpenAI’s.

In the press briefing, Mark Chen, OpenAI’s chief research officer, again denied that any agents or OpenAI employees accessed Buckmaster and Alpöge’s transcripts—but given what has been revealed about the Hugging Face hack, it’s clear that OpenAI is not always entirely aware of what its agents are doing. 

If OpenAI’s models did train on Buckmaster and Alpöge’s work, or if its agents somehow gained access to it, then the company’s failure to track down the truth and assign those researchers appropriate credit reflects poorly on it. But there might be a thin silver lining to that version of the story for mathematicians, because it would suggest that the hard work of two humans, one of whom is a prominent expert on Navier-Stokes, was essential to the agents’ ability to solve the Millennium Problem.

Experts have long identified “research taste,” or the ability to choose promising research questions and directions, as a major obstacle for AI in science and mathematics. If the OpenAI agents did indeed choose to follow the Córdoba–Martínez-Zoroa approach because Buckmaster and Alpöge had done the same, then human research taste played an essential role in OpenAI’s success.

Even so, the bigger picture here is sobering. The progress that Buckmaster and Alpöge made over almost a year of collaboration with publicly available models speaks to the promise of human–AI collaboration. But they were not able to achieve a full solution. Meanwhile, OpenAI brute-forced a solution in a few days using an internal model, and their successful solution came at an astronomical cost: In the press briefing, Bubeck and Chen said the team was only able to solve the problem by running about 10,000 agents concurrently, at a cost of millions of dollars.

Over the past few months, I’ve heard from several researchers that mathematicians are becoming depressed, and it’s not difficult to see why. Mathematics is quickly becoming the province of frontier AI companies with impressive internal-only models, money to burn, and a lack of collaborative spirit. “Whether AI companies will decide to spend their money on doing one thing or another, I truly don’t know,” says Gómez-Serrano. “What is clear is that very few mathematicians will have resources of that scale.”

If OpenAI and Anthropic keep striving for more and more impressive mathematical accolades, there might not be any open problems left for human mathematicians outside of those companies to wrestle with. That would dramatically change the field of mathematics.

Last week, UCLA mathematician Terence Tao wrote a Mastodon thread describing how important mistakes, wrong directions, and incomplete solutions are for the field. “In most cases in pure mathematics, the problems are posed not because we desperately want the solution to these problems in and of themselves, but because we have seen from past experience that human-directed efforts to solve these problems tend to spur further development of the field,” Tao wrote.

“Prematurely solving the problem by purely AI-powered methods—particularly without full transparency into the solution process—can contaminate this process to the point where it actually becomes a net negative for the progress of mathematics as a whole.”

Humans might take longer than agents to solve mathematical problems, but in the process, they uncover new mathematical approaches and ideas that might inspire their peers and even birth their own subfields.

But when AI agents solve those problems instead—and when private companies keep the agents’ wrong turns from public view—those benefits disappear. It remains to be seen what else will vanish in the process. 

  •  

Google I/O showed how the path for AI-driven science is shifting

During Tuesday’s Google I/O keynote, Demis Hassabis, the CEO of Google DeepMind, proclaimed that we are currently “standing in the foothills of the singularity.” It was a striking statement—the singularity is the theoretical future moment when AI rapidly exceeds human intelligence and dramatically transforms the world. But what struck me as I listened in the audience was the context in which he said those words. 

He was on stage to close out the session with a segment on scientific AI, the centerpiece of which was a video detailing how the company’s weather prediction software provided an advance alert about Hurricane Melissa’s catastrophic landfall in Jamaica last year—and potentially saved lives. If that software, called WeatherNext, helped anyone escape the storm or better fortify their home, that’s an enormous and meaningful achievement. But it’s hardly evidence of an impending singularity.

The juxtaposition of Hassabis’ lofty rhetoric with the real-world results of WeatherNext highlighted the tension between two very different approaches to AI for science. The first focuses on AI tools, like WeatherNext, that are designed and trained to solve specific scientific problems. The second is agentic, LLM-based systems that could one day execute cutting-edge research projects without human involvement.

This second vision powers a great deal of AI enthusiasm right now, including recent excitement around recursive self-improvement, or the idea that AI systems could eventually become the primary drivers of AI advancement—a process that would get faster and faster as the AI systems grow smarter. And agentic systems are now making real research contributions, sometimes with limited human guidance.

Just this week, Pushmeet Kohli, Google Cloud’s chief scientist, published a piece in a special AI and science issue of the journal Daedalus, writing: “We are moving toward AI that doesn’t just facilitate science but begins to do science.” With autonomous AI scientists on the horizon, it’s harder to justify massive efforts to develop super-specialized tools—even one like AlphaFold, for which DeepMind scientists won a Nobel Prize, or a potentially life-saving system like WeatherNext. It also heralds a far stranger future for science, in which humans and AI systems collaborate as peers—or AI even makes scientific progress on its own.

To be clear, Google does not appear to be abandoning its work on specialized AI for science tools. AlphaGenome and AlphaEarth Foundations, which are trained for genetics and Earth science applications respectively, were released last summer, and the newest version of WeatherNext came out in November.

What’s more, such tools remain extremely popular among scientists. Last year, for instance, Google reported that protein structure predictions from AlphaFold have been used by over three million researchers worldwide. And Isomorphic Labs, a Google subsidiary that aims to use AlphaFold and related technologies to develop new drugs, just raised a $2 billion Series B funding round.

But there are concrete signs of realignment, in both enthusiasm and resources. Last month, the Los Angeles Times reported that Google fellow John Jumper, who won the Nobel for AlphaFold, is now working on AI coding, not on science-specific AI tools. It’s not surprising that Google is assigning its best minds to the coding problem, as the company has recently taken a reputational hit because its coding tools don’t currently stand up to those offered by Anthropic and OpenAI. But it may also signal a prioritization of agentic science on Google’s part, as coding abilities are key to the success of some of those systems. 

Across the industry, agentic researcher systems are showing real potential. This week, OpenAI announced that one of their models had disproved an important mathematics conjecture—perhaps the most meaningful contribution that generative AI has made to mathematics so far, according to some mathematicians.

Importantly, the model used by OpenAI is not specialized for solving mathematical problems, or even for research; according to the company, it’s a general-purpose reasoning model in the vein of GPT-5.5. If general agents can make independent contributions to mathematical research, they might soon be able to do the same in science (though the fact that ideas in science must be verified experimentally makes it a tougher domain for AI).

Google is certainly devoting a lot of attention toward an agent-driven scientific future. The big scientific announcement at I/O was the new Gemini for Science package, which unites several of the company’s LLM-based scientific systems under one brand.

This includes the hypothesis-generating AI Co-Scientist and algorithm-optimizing AlphaEvolve, which are still not publicly available—but as Google is now allowing any researcher to apply for access to Gemini for Science, they may soon see wider adoption in the scientific community. Scientists who were involved in early testing are enthusiastic about their potential: Gary Peltz, a Stanford geneticist, compared using the AI Co-Scientist to “consulting the oracle of Delphi” in a Nature Medicine article.

Gemini for Science isn’t incompatible with specialized tools; to the contrary, agentic systems can be designed to call on such tools when they might be useful. And no agentic system can predict the structure that a protein will fold into without AlphaFold’s help (at least not yet). But the company seems to be shifting its public image—and at least some resources and personnel, such as Jumper—away from specifically developing those kinds of tools. Though it has only been five years since AlphaFold solved the protein-folding problem, both the technology and the discourse have quickly moved beyond that once-revolutionary achievement.

Google has been careful to position this new set of scientific agents as an accelerant for human scientists, rather than a replacement for them—the choice of the name AI Co-Scientist as opposed to AI Scientist, for instance, appears quite deliberate. Hassabis uses that same human-centric framing when he talks about changes in the landscape of scientific AI. “For the next decade or so, we should think about AI as this amazing tool to help scientists,” Hassabis said in an interview published in the Daedalus issue. “Beyond that timeframe, it is hard to say with any certainty, but perhaps these systems will become more like collaborators.”

But no one can be an effective scientific collaborator without also being a skilled scientist in their own right. And if Hassabis is anywhere near the mark when he talks about the “foothills of the singularity,” then AI scientists could eventually exceed the capabilities of their human counterparts.

In a discussion with the journalist Mike Allen at I/O, Hassabis spoke of how he was initially inspired to pursue AI when he observed how progress in physics had stagnated since the 1970s; he wondered whether the human mind had reached its limits in that domain, and if AI could help to overcome that barrier. Superhuman agentic scientists would certainly fit that bill. We might not ever get anywhere near there, but Google seems to be aiming itself toward that summit.

  •  

The Enhanced Games fit right in with the rest of 2026’s longevity vibes

This Sunday, a group of 42 athletes will gather in Las Vegas to compete in a somewhat unusual sporting competition. Participants in the inaugural Enhanced Games are being encouraged to take performance-enhancing drugs. The goal is to “push the boundaries of human performance.”

The games’ organizers have said that competitors will only be taking substances that have been approved by the US Food and Drug Administration, and that they are all being medically monitored and supervised. But they have also said they expect to see world records broken—and are offering substantial prizes to athletes who succeed in doing so.

As you might expect, the event is generating a mix of curiosity, excitement, and condemnation from various quarters. To me, it feels like very much a reflection of where we are today—an era of peptide-crazed looksmaxxing in which consumers are being encouraged to get thinner than ever, optimize for longevity, and have their “best baby.” It’s 2026, and if you’re not enhancing, what are you even doing?

So, these games. They’ll feature competitions in four categories: swimming, track and field, weightlifting, and strongman (which also involves lifting weights). Many of the competitors already hold national and world records, and some are Olympic medalists. They’ve been paid a salary and will compete for prizes from a $25 million pot. The money has been a major draw for at least some of the athletes.

Another draw is the opportunity to openly experiment with drugs that might boost their performance. In the world of elite sport, every microsecond and every millimeter counts. Athletes—most of whom arguably have genetics on their side already—follow meticulous diet, training, and recovery protocols and wear specially designed gear that allows them to reach for those performance bests.

But within most sporting communities, there are limits. The World Anti-Doping Agency—an international outfit that fights the use of drugs in sports—maintains a lengthy list of “non-approved substances” that are banned in international sporting events. It features many anabolic steroids (which can build muscle), hormones (such as those that stimulate testosterone production or increase the ability of blood to carry oxygen), growth factors (which can stimulate muscle growth and repair, among other things), and more.

Some of these substances have been FDA approved to treat health disorders. And that means they can be used by participants in the Enhanced Games, according to the organization’s rules.

I’ll briefly point out the obvious here—just because a drug has been approved by the FDA doesn’t mean it’s totally safe for everyone and anyone. The risks associated with use of anabolic steroids, for example, include high blood pressure, acne, depression, and liver tumors. Growth hormone use can cause weak muscles, affect vision, and even lead to diabetes.

“Technological doping,” or using improved equipment to gain advantage, has also been supported by the games’ organizers. Last year, participating swimmer Kristian Gkolomeev was reported to have broken a record in a 50-meter freestyle time trial while wearing a polyurethane “super” swimsuit. Such suits have been banned for use in the Olympics since a slew of record-breaking performances in 2008 and 2009. Back then, the swimming governing body ruled that they gave athletes an unfair advantage. But hey, this is the Enhanced Games, where the word “unfair” seems to have a completely different meaning.

Can we expect more records to be broken on Sunday? Maybe. In addition to prize money for winning an event, any athlete who manages to beat a record stands to win up to $1 million, the sum also awarded to Gkolomeev last year following his time trial. But those performances won’t be recognized by official sporting bodies.

Plenty of concerns have been raised about these games. Some argue that they are unsafe and promote risky drug use. Others see them as a “clown show,” and a slap in the face to “clean” athletes who train hard without the use of prohibited drugs. World Athletics president Sebastian Coe has said that anyone who takes part is “moronic,” and World Aquatics, which oversees international competitions in water sports, has banned Enhanced Games participants from its events and activities.

But. The games—and the participating athletes—will still get a huge amount of attention. As a result, so will performance-enhancing drugs. Enhanced, the company behind the games, also runs an online store. There, you can buy a $52 T-shirt emblazoned with the message “I am Enhanced.”

There is also a range of prescription drugs on offer, including peptides “to support recovery, vitality, and longevity.” One of these is a growth hormone that the FDA approved in 1997 for the treatment of children with “growth failure.” The compounded version offered on the Enhanced website, which is not FDA approved, is marketed for longevity, supporting deep sleep and “overall wellness and vitality.” (“Marketed” is the key word here. The drug has, again, not been approved for that purpose.)

It all fits very well with the zeitgeist. Sure, we don’t yet have any drugs that are designed to extend human lifespan. But the search for anti-aging drugs is getting more attention—and funding—than ever. People, particularly women, are seemingly not allowed to visibly age anymore—we have filters and facelifts for that now. The idea that “death is wrong” is gaining acceptance.

And self-experimentation is rife. “Biohacking” was shortlisted for Collins Dictionary’s Word of the Year in 2025. Peptides are everywhere, despite all the unknowns surrounding their safety and effectiveness. So are longevity clinics, despite the fact that most are selling unproven treatments. US states like Montana are making it easier for people to get hold of unapproved “therapies.”

Companies are even offering would-be parents the option to choose the potential future children expected to live longest. Yep—you can supposedly optimize your embryos now, too.

In this climate, the Enhanced Games don’t feel so radical. They feel entirely fitting for our era of questionable optimization despite the risks —an era when, apparently, being human is no longer enough.

  •  

Anthropic’s Code with Claude showed off coding’s future—whether you like it or not

The vibes were strong at Code with Claude, Anthropic’s two-day event for software developers in London that kicked off on May 19, the same day as Google’s I/O in Palo Alto. (A coincidence, not a flex, Anthropic staffers assured me.)

“Who here has shipped a pull request in the last week that was completely written by Claude?” Jeremy Hadfield, an engineer at Anthropic, asked from the main stage. Almost half the people in the packed room—many sitting with laptops on their knees, coding or prompting as they watched the talks—raised their hands.

Pull requests are fixes or updates to existing software that are submitted for review before they go live. They are the bread and butter of software development, the chunks of code that most professional developers spend their lives writing—or did until now.

“Who here has shipped a pull request that was completely written by Claude where they did not read the code at all?” Hadfield asked next. Nervous laughter. Most of the hands stayed up.

It’s not news that LLM-powered tools like Anthropic’s Claude Code and OpenAI’s Codex have upended the way software gets made. Top tech companies now like to boast of how little code their developers write by hand. (“Most software at Anthropic is now written by Claude,” Hadfield said. “Claude has written most of the code in Claude Code.”) OpenAI, Google, and Microsoft make similar claims. Many others wish they could.

Even so, it is striking how normal this new paradigm already seems, and how fast it has set in. This was the second year that Anthropic has put on developer events, which also run in San Francisco and Tokyo. This time last year, the company had just released Claude 4. It could code, kind of. But with Anthropic’s latest string of updates—especially Claude 4.6 and then 4.7, released in February and April—Claude Code is a tool that more and more developers seem happy to hand their work off to.   

An 8-bit character with a chef's hat in a pixel kitchen flips food in a fry pan over a pixel stove
Let Claude cook.
ANTHROPIC (GRAPHIC) / WILL DOUGLAS HEAVEN (PHOTO)

Anthropic says its goal is to push automation as far as it will go. Instead of using AI to generate code and then having humans clean it up and fix the mistakes, it wants Claude to check and correct its own work. “The default isn’t ‘I’m going to prompt Claude’—the default is now ‘I’m going to have Claude prompt itself,’” Boris Cherny, who heads Claude Code, said in the opening keynote.

If all goes well, human developers shouldn’t even see the error messages when something doesn’t work. That will all be handled by Claude, which will test and tweak, test and tweak, until everything runs as it should. As Ravi Trivedi, an engineer at Anthropic, put it in another talk: “The key principle is getting out of Claude’s way. We like to say: ‘Let it cook.’”

Trivedi presented a new feature in Claude Managed Agents, Anthropic’s cloud-based setup for building and running multi-agent systems, announced two weeks ago, which the company calls dreaming. Claude agents write notes to themselves, recording and saving useful information about specific tasks. When another coding agent, say, starts to work on the same code that others have worked on, it can use the notes they left behind to get up to speed faster and learn from any errors those previous agents may have made.

Dreaming is a system that Claude agents can use to read through the notes and consolidate the information they contain, spotting patterns and common issues across different tasks. In theory, dreaming should help coding agents learn about a particular code base and get better and better at working on it.

Success stories

Code with Claude is an event aimed at developers. As well as product showcases and hands-on workshops from Anthropic, there were how-tos from a range of companies that have reshaped their software development teams around Claude Code, including Spotify and Delivery Hero as well as Lovable, Base44, and Monday.com—three startups vibe-coding apps that help people vibe-code apps.

There were no signs of unease at Code with Claude. Everybody I met wanted in.

And yet outside the conference there have been a number of reports that many coders are starting to question this bright new future. Some gripe in online forums like Reddit and Hacker News that AI coding tools are being pushed by managers chasing productivity gains, when in practice the technology makes software development harder because of all the extra code developers now have to review. “The only people I’ve heard saying that generated code is fine are those who don’t read it,” a user called pron posted on Hacker News last week. 

Others claim that their coding abilities have fallen off as they hand more tasks to AI. And researchers have warned that AI tools can produce unsafe code that will make software more vulnerable to attacks.  

I sat down with Claude engineering lead Katelyn Lesse and Claude product lead Angela Jiang and asked them what they made of the concerns that a sudden flood of code generated (and shipped) without proper human oversight was kicking serious security and maintenance problems down the road.

“All of the old software development best practices still apply. They’ve applied this entire time,” said Lesse. “I think there are a lot of people and teams that may have lost sight of them in this moment.” 

And yet as Anthropic and others push for greater automation and tools like Claude Code improve, the temptation increases to offload more and more tasks, including oversight. Lesse told me that some of the technical managers at Anthropic are exhausted by keeping up with all the code their teams now produce. “Part of things happening so much more quickly is just managing your time,” she said.

“I think that right now Claude is probably as good as a midlevel engineer at writing code,” she added. You still need expert engineers to design a system and troubleshoot harder problems, she said. “But over time we want Claude to get better and better at all different types of engineering.”

Jiang agreed: “I think the absolute end state we’re trying to get to is Claude basically being able to build itself.”

Correction: Dreaming is a feature of Claude Managed Agents not Claude Code. The article has been updated.

  •  

Desalination plants in the Middle East are increasingly vulnerable

MIT Technology Review Explains: Let our writers untangle the complex, messy world of technology to help you understand what’s coming next. You can read more from the series here.

As the conflict in Iran has escalated, a crucial resource is under fire: the desalination technology that supplies water across much of the region.

In early March, Iran’s foreign minister accused the US of attacking a desalination plant on Qeshm Island in the Strait of Hormuz and disrupting the water supply to nearly 30 villages. (The US denied responsibility.) In the weeks since, both Bahrain and Kuwait have reported damage to desalination plants and blamed Iran, though Iran also denied responsibility.

In late March, President Donald Trump threatened the destruction of “possibly all desalinization plants” in Iran if the Strait of Hormuz was not reopened. Since then, he’s escalated his threats against Iran, warning of plans to attack other crucial civilian infrastructure like power plants and bridges.

Countries in the Middle East, particularly the Gulf states, rely on the technology to turn salt water into fresh water for farming, industry, and—crucially—drinking. The mounting attacks and threats to date highlight just how vital the industry is to the region—a situation made even more precarious by rising temperatures and extreme weather driven by climate change.

Right now, 83% of the Middle East is under extremely high water stress, says Liz Saccoccia, a water security associate at the World Resources Institute. Future projections suggest that’s going to increase to about 100% by 2050, she adds: “This is a continuing trend, and it’s getting worse, not better.”

Here’s a look at desalination technology in the Middle East and what wartime threats to the critical infrastructure could mean for people in the region. 

A vital resource

Desalination technology has helped provide water supplies in the Middle East since the early 20th century and became widespread in the 1960s and 1970s.

There are two major categories of desalination plants. Thermal plants use heat to evaporate water, leaving salt and other impurities behind. The vapor can then be condensed into usable fresh water. The alternative is membrane-based technology like reverse osmosis, which pushes water through membranes that have tiny pores—so small that salt can’t get through.

Early desalination plants in the Middle East were the first type, burning fossil fuels to evaporate water, leaving the salt behind. This technique is incredibly energy-intensive, and over time, processes that rely on filters became the dominant choice.

Membrane technologies have made up essentially all new desalination capacity in recent years; the last major thermal plant built in the Gulf came online in 2018. Many reverse osmosis plants still rely on fossil fuels, but they’re more efficient. Since then, membrane technologies have added more than 15 million cubic meters of daily capacity—enough to supply water to millions of people.

Capacity has expanded quickly in recent years; between 2006 and 2024, countries across the Middle East collectively spent over $50 billion building and upgrading desalination facilities, and nearly that much operating them.

Today, there are nearly 5,000 desalination plants operational across the Middle East.

And looking ahead, growth is continuing. Between 2024 and 2028, daily capacity is expected to grow from about 29 million cubic meters to 41 million cubic meters.

Uneven vulnerabilities

Some countries rely on the technology more than others. Iran, for example, uses desalination for about 3% of its municipal fresh water. The country has access to groundwater and some surface water, including rivers, though these resources are being stretched thin by agriculture and extreme drought.

Other nations in the region, particularly the Gulf countries (Bahrain, Qatar, Kuwait, the United Arab Emirates, Saudi Arabia, and Oman), have much more limited water resources and rely heavily on desalination. Across these six nations, all but the UAE get more than half their drinking water from desalination, and for Bahrain, Qatar, and Kuwait the figure is more than 90%.

“The Gulf countries are much, much more vulnerable to attacks on their desalination plants than Iran is,” says David Michel, a senior associate in the global food and water security program at the Center for Strategic and International Studies.

There are thousands of desalination facilities across the region, so the system wouldn’t collapse if a small number were taken offline, Michel says. However, in recent years there’s been a trend toward larger, more centralized plants.

The average desalination plant is about 10 times larger than it was 15 years ago, according to data from the International Energy Agency. The largest desalination plants today can produce 1 million cubic meters of water daily, enough for hundreds of thousands of people. Taking one or more of these massive facilities offline could have a significant effect on the system, Michel says.

Escalating threats

Desalination facilities are quite linear, meaning there are multiple steps and pieces of equipment that work in sequence—and the failure of a component in that chain can take an entire facility down. Attacks on water inlets, transportation networks, and power supplies can also disrupt the system, Michel says. 

During the Gulf War in 1991, Iraqi forces pumped oil into the gulf, contaminating the water and shutting down desalination plants in Kuwait. 

The facilities are also generally located close to other targets in this conflict. Desalination is incredibly energy intensive, so about three-quarters of facilities in the region are next to power plants. Trump has repeatedly threatened power plants in Iran. In response, Iran’s military has said that if civilian targets are hit, the country will respond with strikes that are “much more devastating and widespread.” Other governments and organizations, including the United Nations, the European Union, and the Red Cross, have broadly condemned threats to infrastructure as illegal. 

But war isn’t the only danger facing these plants, even if it is the most immediate. Some studies have suggested that global warming could strengthen cyclones in the region, and these extreme weather events could force shutdowns or damage equipment.

Water pollution could also cause shutdowns. Oil spills, whether accidental or intentional, as in the case of the Gulf War, can  wreak havoc. And in 2009, a red algae bloom closed desalination plants in Oman and the United Arab Emirates for weeks. The algae fouled membranes and blocked the plants from being able to take water in from the Persian Gulf and the Gulf of Oman.

Desalination facilities could become more resilient to threats in the future, and they may need to as their importance continues to grow. 

There’s increasing interest in running desalination facilities at least partially on solar power, which could help reduce dependence on the oil that powers most facilities today. The Hassyan seawater desalination project in the UAE, currently under construction, would be the largest reverse osmosis plant in the world to operate solely with renewable energy. 

Another way to increase resilience is for countries to build up more strategic water storage to meet demand. Qatar recently issued new policies that aim to improve management and storage of desalinated water, for example. Countries could also work together to invest in shared infrastructure and policies that help strengthen the water supply through the region. 

Preparedness, resilience, and cooperation will be key for the Middle East broadly as critical infrastructure, including the water supply, is increasingly under threat. 

“The longer the conflict goes on, the more likely we’ll see significant water infrastructure damage,” says Ginger Matchett, an assistant director at the Atlantic Council. “What worries me is that after this war ends, some of the lessons will show how water can be weaponized more strategically than previously imagined.” 

  •  

There are more AI health tools than ever—but how well do they work?

Earlier this month, Microsoft launched Copilot Health, a new space within its Copilot app where users will be able to connect their medical records and ask specific questions about their health. A couple of days earlier, Amazon had announced that Health AI, an LLM-based tool previously restricted to members of its One Medical service, would now be widely available. These products join the ranks of ChatGPT Health, which OpenAI released back in January, and Anthropic’s Claude, which can access user health records if granted permission. Health AI for the masses is officially a trend. 

There’s a clear demand for chatbots that provide health advice, given how hard it is for many people to access it through existing medical systems. And some research suggests that current LLMs are capable of making safe and useful recommendations. But researchers say that these tools should be more rigorously evaluated by independent experts, ideally before they are widely released. 

In a high-stakes area like health, trusting companies to evaluate their own products could prove unwise, especially if those evaluations aren’t made available for external expert review. And even if the companies are doing quality, rigorous research—which some, including OpenAI, do seem to be—they might still have blind spots that the broader research community could help to fill.

“To the extent that you always are going to need more health care, I think we should definitely be chasing every route that works,” says Andrew Bean, a doctoral candidate at the Oxford Internet Institute. “It’s entirely plausible to me that these models have reached a point where they’re actually worth rolling out.”

“But,” he adds, “the evidence base really needs to be there.”

Tipping points 

To hear developers tell it, these health products are now being released because large language models have indeed reached a point where they can effectively provide medical advice. Dominic King, the vice president of health at Microsoft AI and a former surgeon, cites AI advancement as a core reason why the company’s health team was formed, and why Copilot Health now exists. “We’ve seen this enormous progress in the capabilities of generative AI to be able to answer health questions and give good responses,” he says.

But that’s only half the story, according to King. The other key factor is demand. Shortly before Copilot Health was launched, Microsoft published a report, and an accompanying blog post, detailing how people used Copilot for health advice. The company says it receives 50 million health questions each day, and health is the most popular discussion topic on the Copilot mobile app.

Other AI companies have noticed, and responded to, this trend. “Even before our health products, we were seeing just a rapid, rapid increase in the rate of people using ChatGPT for health-related questions,” says Karan Singhal, who leads OpenAI’s Health AI team. (OpenAI and Microsoft have a long-standing partnership, and Copilot is powered by OpenAI’s models.)

It’s possible that people simply prefer posing their health problems to a nonjudgmental bot that’s available to them 24-7. But many experts interpret this pattern in light of the current state of the health-care system. “There is a reason that these tools exist and they have a position in the overall landscape,” says Girish Nadkarni, chief AI officer​ at the Mount Sinai Health System. “That’s because access to health care is hard, and it’s particularly hard for certain populations.”

The virtuous vision of consumer-facing LLM health chatbots hinges on the possibility that they could improve user health while reducing pressure on the health-care system. That might involve helping users decide whether or not they need medical attention, a task known as triage. If chatbot triage works, then patients who need emergency care might seek it out earlier than they would have otherwise, and patients with more mild concerns might feel comfortable managing their symptoms at home with the chatbot’s advice rather than unnecessarily busying emergency rooms and doctor’s offices.

But a recent, widely discussed study from Nadkarni and other researchers at Mount Sinai found that ChatGPT Health sometimes recommends too much care for mild conditions and fails to identify emergencies. Though Singhal and  some other experts have suggested that its methodology might not provide a complete picture of ChatGPT Health’s capabilities, the study has surfaced concerns about how little external evaluation these tools see before being released to the public.

Most of the academic experts interviewed for this piece agreed that LLM health chatbots could have real upsides, given how little access to health care some people have. But all six of them expressed concerns that these tools are being launched without testing from independent researchers to assess whether they are safe. While some advertised uses of these tools, such as recommending exercise plans or suggesting questions that a user might ask a doctor, are relatively harmless, others carry clear risks. Triage is one; another is asking a chatbot to provide a diagnosis or a treatment plan. 

The ChatGPT Health interface includes a prominent disclaimer stating that it is not intended for diagnosis or treatment, and the announcements for Copilot Health and Amazon’s Health AI include similar warnings. But those warnings are easy to ignore. “We all know that people are going to use it for diagnosis and management,” says Adam Rodman, an internal medicine physician and researcher at Beth Israel Deaconess Medical Center and a visiting researcher at Google.

Medical testing

Companies say they are testing the chatbots to ensure that they provide safe responses the vast majority of the time. OpenAI has designed and released HealthBench, a benchmark that scores LLMs on how they respond in realistic health-related conversations—though the conversations themselves are LLM-generated. When GPT-5, which powers both ChatGPT Health and Copilot Health, was released last year, OpenAI reported the model’s HealthBench scores: It did substantially better than previous OpenAI models, though its overall performance was far from perfect. 

But evaluations like HealthBench have limitations. In a study published last month, Bean—the Oxford doctoral candidate—and his colleagues found that even if an LLM can accurately identify a medical condition from a fictional written scenario on its own, a non-expert user who is given the scenario and asked to determine the condition with LLM assistance might figure it out only a third of the time. If they lack medical expertise, users might not know which parts of a scenario—or their real-life experience—are important to include in their prompt, or they might misinterpret the information that an LLM gives them.

Bean says that this performance gap could be significant for OpenAI’s models. In the original HealthBench study, the company reported that its models performed relatively poorly in conversations that required them to seek more information from the user. If that’s the case, then users who don’t have enough medical knowledge to provide a health chatbot with the information that it needs from the get-go might get unhelpful or inaccurate advice.

Singhal, the OpenAI health lead, notes that the company’s current GPT-5 series of models, which had not yet been released when the original HealthBench study was conducted, do a much better job of soliciting additional information than their predecessors. However, OpenAI has reported that GPT-5.4, the current flagship, is actually worse at seeking context than GPT-5.2, an earlier version.

Ideally, Bean says, health chatbots would be subjected to controlled tests with human users, as they were in his study, before being released to the public. That might be a heavy lift, particularly given how fast the AI world moves and how long human studies can take. Bean’s own study used GPT-4o, which came out almost a year ago and is now outdated. 

Earlier this month, Google released a study that meets Bean’s standards. In the study, patients discussed medical concerns with the company’s Articulate Medical Intelligence Explorer (AMIE), a medical LLM chatbot that is not yet available to the public, before meeting with a human physician. Overall, AMIE’s diagnoses were just as accurate as physicians’, and none of the conversations raised major safety concerns for researchers. 

Despite the encouraging results, Google isn’t planning to release AMIE anytime soon. “While the research has advanced, there are significant limitations that must be addressed before real-world translation of systems for diagnosis and treatment, including further research into equity, fairness, and safety testing,” wrote Alan Karthikesalingam, a research scientist at Google DeepMind, in an email. Google did recently reveal that Health100, a health platform it is building in partnership with CVS, will include an AI assistant powered by its flagship Gemini models, though that tool will presumably not be intended for diagnosis or treatment.

Rodman, who led the AMIE study with Karthikesalingam, doesn’t think such extensive, multiyear studies are necessarily the right approach for chatbots like ChatGPT Health and Copilot Health. “There’s lots of reasons that the clinical trial paradigm doesn’t always work in generative AI,” he says. “And that’s where this benchmarking conversation comes in. Are there benchmarks [from] a trusted third party that we can agree are meaningful, that the labs can hold themselves to?”

They key there is “third party.” No matter how extensively companies evaluate their own products, it’s tough to trust their conclusions completely. Not only does a third-party evaluation bring impartiality, but if there are many third parties involved, it also helps protect against blind spots.

OpenAI’s Singhal says he’s strongly in favor of external evaluation. “We try our best to support the community,” he says. “Part of why we put out HealthBench was actually to give the community and other model developers an example of what a very good evaluation looks like.” 

Given how expensive it is to produce a high-quality evaluation, he says, he’s skeptical that any individual academic laboratory would be able to produce what he calls “the one evaluation to rule them all.” But he does speak highly of efforts that academic groups have made to bring preexisting and novel evaluations together into comprehensive evaluations suites—such as Stanford’s MedHELM framework, which tests models on a wide variety of medical tasks. Currently, OpenAI’s GPT-5 holds the highest MedHELM score.

Nigam Shah, a professor of medicine at Stanford University who led the MedHELM project, says it has limitations. In particular, it only evaluates individual chatbot responses, but someone who’s seeking medical advice from a chatbot tool might engage it in a multi-turn, back-and-forth conversation. He says that he and some collaborators are gearing up to build an evaluation that can score those complex conversations, but that it will take time, and money. “You and I have zero ability to stop these companies from releasing [health-oriented products], so they’re going to do whatever they damn please,” he says. “The only thing people like us can do is find a way to fund the benchmark.”

No one interviewed for this article argued that health LLMs need to perform perfectly on third-party evaluations in order to be released. Doctors themselves make mistakes—and for someone who has only occasional access to a doctor, a consistently accessible LLM that sometimes messes up could still be a huge improvement over the status quo, as long as its errors aren’t too grave. 

With the current state of the evidence, however, it’s impossible to know for sure whether the currently available tools do in fact constitute an improvement, or whether their risks outweigh their benefits.

  •  

A woman’s uterus has been kept alive outside the body for the first time

“Think of this as a human body,” says Javier González.

In front of me is essentially a metal box on wheels. Standing at around a meter in height, it reminds me of a stainless-steel counter in a restaurant kitchen. It is covered in flexible plastic tubing—which act as veins and arteries—connecting a series of transparent containers, the organs of this machine.

What makes it extra special is the role of the cream-colored tub that sits on its surface. Ten months ago, González, a biomedical scientist who developed the device with his colleagues at the Carlos Simon Foundation, carefully placed a freshly donated human uterus in the tub. The team connected it to the device’s tubes and pumped in modified human blood.

The device kept the uterus alive for a day—a new feat that could represent the first step to the long-term maintenance of uteruses outside the human body. The work has not yet been published. 

The team members want to keep donated human uteruses alive long enough to see a full menstrual cycle. They hope this will help them study diseases of the uterus and learn more about how embryos burrow their way into the organ’s lining at the start of a pregnancy. They also hope that future iterations of their device might one day sustain the full gestation of a human fetus.

The machine is technically called PUPER, which stands for “preservation of the uterus in perfusion.” But González’s colleague Xavier Santamaria says the team has adopted a nickname for it: “We call it ‘Mother.’”

The organ in the machine

González and Santamaria, medical vice president of the Carlos Simon Foundation, demonstrated how the device might work when I visited the foundation in Valencia, Spain, earlier this month (although it held no organs on that day). 

Both are interested in learning more about implantation, the moment at which an embryo attaches itself to the lining of a uterus—essentially, the very first moment of pregnancy.

The foundation’s founder and director, Carlos Simon, believes it’s a sticking point in IVF: Scientists have made many improvements to the technology over the years, but the failure of embryos to implant underlies plenty of unsuccessful IVF cycles, he says. Being able to carefully study how the process works in a real, living organ might give the team a better idea of how to prevent those failures.

a person in gloves stands next to a machine with lots of tubing coming in and out of the metal exterior
JESS HAMZELOU
a sheep uterus resting on gauze connected to several tubes
JAVIER GONZALES/CARLOS SIMON FOUNDATION

Javier González demonstrates the perfusion machine. A previous iteration of the device kept a sheep’s uterus (right) alive for a day.

The team took inspiration from advances in technologies designed to maintain donated organs for transplantation. In recent years, researchers around the world have created devices that deliver nutrients and filter waste so that organs can survive longer after being removed from donors’ bodies.

The main goal here is to buy time. A human organ might last only a matter of hours outside the body, so a transplant may require frantic preparation for the recipient, sometimes in the middle of the night. With a little more time, doctors could find better donor-patient matches and potentially test the quality of donated organs.

This approach is called normothermic or machine perfusion, and it is already being used clinically for some liver, kidney, and heart transplants.

The team at the Carlos Simon Foundation built a similar machine for uteruses. A blood bag hangs on one side. From there, blood is ferried via plastic tubing to a pump, which functions as the heart. The pump shunts the blood through an oxygenator, which adds oxygen and removes carbon dioxide as the lungs would in a human body.

The blood is warmed and passed through sensors that monitor the levels of glucose and oxygen, along with other factors. It passes through a “kidney” to remove waste. And finally the blood reaches the uterus, hooked up to its own plastic “arteries” and “veins.” The organ itself sits at a tilt, just as in the body, and is kept in a humid environment to stay moist.

Mother’s first uterus

The team first began testing an early prototype of the device with sheep uteruses around four years ago. That meant carting the machine to an animal research center in Zaragoza, around 200 miles away. Over the course of the preliminary study, veterinary surgeons removed the uteruses of six sheep and hooked them up to the machine. They kept each uterus alive for a day, using blood from the same animals.

After the sheep experiments, the researchers carted their machine back to Valencia and modified it to achieve its current incarnation, “Mother.” They started working with a local hospital that performed hysterectomies. And in May last year, they were offered their first human uterus.

The team needed to be quick. “You need to put [the uterus in the machine] within a couple of hours, maximum, of the extraction,” says Santamaria. He and his colleagues also needed to connect the uterus’s blood vessels to the tubing delicately, taking care to avoid any blockages (clotting is a major challenge in organ perfusion). The organ was hooked up to human blood obtained from a blood bank.

It seemed to work—at least temporarily. “We kept it alive for one day,” says Santamaria.

“As a proof of concept, it is impressive,” says Keren Ladin, a bioethicist who has focused on organ transplantation and perfusion at Tufts University. “These are early days.”

It might not sound like much, but 24 hours is a long time for an organ to be out of the body. Maintaining a donated uterus for that long could expand the options for uterus transplant, a fairly new procedure offered to some people who want to be pregnant but don’t have a functional uterus, says Gerald Brandacher, professor of experimental and translational transplant surgery at the Medical University of Innsbruck in Austria.

“It is better than what we currently have, because we have only a couple of hours,” he says. So far, most uterus transplants have been planned operations involving organs from living donors. A technology like this could allow for the use of more organs from deceased donors, he says.

That work is “not in the immediate pipeline” for the team in Spain, says Santamaria. “We are working on other problems.”

Pregnancy in the lab?

Santamaria, González, and their colleagues are more interested in using sustained human uteruses for research. 

They’ve mounted a camera to a wall in the corner of the room, pointed at their machine. It allows the team to monitor “Mother” remotely, and to check if any valves disconnect. (That happened once before—a spike in pressure caused the blood bag to come loose, spilling a liter of blood on the floor, Santamaria says.)

They’d like to be able to keep their uteruses alive for around 28 days to study the menstrual cycle and disorders that affect the uterus, like endometriosis and fibroids.

It won’t be easy to maintain a uterus for that long, cautions Brandacher. As far as he knows, no one has been able to maintain a liver for more than seven days. “No studies out there … have shown 30-day survival in a machine perfusion circuit,” he says.

But it’s worth the effort. The team’s main interest is learning more about how embryos implant in the uterine lining at the start of a pregnancy. They hope to be able to test the process in their outside-the-body uteruses.

They won’t be allowed to use human embryos for this, says González—that would cross an ethical boundary. Instead, they plan to use embryo-like structures made from stem cells. The structures closely resemble human embryos but are created in a lab without sperm or eggs.

Simon himself has grander ambitions.

He sees a future in which a machine like “Mother” will be able to fully gestate a human, all the way from embryo to newborn. It could offer a new path to parenthood for people who don’t have a uterus, for example, or who are not able to get pregnant for other reasons.

He appreciates that it sounds futuristic, to say the least. “I don’t know if we will end up having pregnancies inside of the uterus outside of the body, but at least we are ready to understand all the steps to do that,” he says. “You have to start somewhere.”

  •  
❌