❌

Reading view

The AI industry has taken a doomer turn. What now?

This story appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here.

This weekend, Dario Amodei, CEO of Anthropic, posted an essay calling for a brake on the pace of development of LLMs. Amodei cites the looming dangers he sees from the technology, from its use in cyberattacks and bioterrorism to its potential to wreck the economy. The heads of the other three top US AI labs—OpenAI CEO Sam Altman, Google DeepMind chairman Demis Hassabis, and SpaceXAI CEO Elon Musk—voiced their support. “Dario is right,” Musk wrote on X.

Think about how surreal that agreement is for a moment. Just a few months ago, Musk and Altman sat in court attacking each other’s reputations in a (failed) lawsuit that Musk brought against his former OpenAI colleague that was—on paper at least—about whether or not Altman was a trustworthy steward of such dangerous technology.

Amodei’s rift with OpenAI is even deeper. Anthropic was founded in 2021 because Amodei didn’t think Altman took the risks of the technology they were building seriously enough. Anthropic and OpenAI have been competing in a winner-takes-all race ever since. (Hassabis has stayed out of the drama, but his company remains a rival.)

Now, it seems, they’re all in agreement: The latest generation of LLMs aren’t safe and everyone needs to figure out what to do about it. The public messaging from the top AI labs has taken a doomer turn.

It’s easy to be cynical. It’s not at all clear what any of them mean by a slowdown or how it would work. These companies also care a lot about how they come across. With trillion-dollar IPOs in their sights, OpenAI and Anthropic need to reassure investors that they’re the grown-ups in the room while at the same time hinting at the power of the monsters they have created—and intend to tame. Calling for a slowdown does both.

And yet the vibe at the top of these firms really does appear to have shifted. Amodei’s latest post landed six days after OpenAI published an essay by Jakub Pachocki, the firm’s chief scientist, in which he also laid out why he’s concerned about what will happen if the pace of development of LLMs continues unchecked. In short, Pachocki is worried that OpenAI’s ability to build powerful models now far outstrips its ability to monitor and control them.

Amodei and Pachocki each cite the cyberattack against AI firm Hugging Face by a swarm of OpenAI’s agents in July—a hack that OpenAI did not even realize had taken place until days after it was all over—as a wake-up call.

But their exact position is hard to pin down. Pachocki both calls for a slowdown and highlights an urgent need to stay ahead: “The strongest argument I see for continuing to train much smarter models quickly is the need to build defensive systems against the dangers posed by other AI,” he writes. As Pachocki frames it, AI firms are locked in a literal arms race. Slowing down is good, winning is better.

(Don’t forget: OpenAI just spent millions of dollars and a staggering amount of computer power to rush out a controversial math result a few days ahead of Anthropic.)

But let’s assume a slowdown happens. Top labs agree to spend more time and resources on finding ways to monitor and control existing models instead of making more capable ones. They invite outside auditors in to help evaluate those models.

What might this coordinated effort actually achieve? Consider the Hugging Face attack again. OpenAI has said that the model that drove most of the rogue agents was a “highly persistent” next-generation model that it was testing in-house. The implication is that OpenAI has built a model so good it’s dangerous.  

But if you read the reports about the Hugging Face hack published by OpenAI and METR, a third-party firm that OpenAI called in to help them understand what happened, what you come away with is the impression not of a model that was too powerful for OpenAI to keep up with, but of a broken model that OpenAI failed to train properly.

The agents did what they did—including leaving messages for one another, delegating work to other agents, and scouring their environment for any means possible to complete their tasks—because they had been rewarded during training for doing exactly those things. There were also errors in the training setup, such as tasks that were impossible to complete, which pushed the models to find unexpected workarounds that were also rewarded. At the time, many of these issues went overlooked or unreported.

OpenAI says it has stopped training this new model and locked it down. That makes it sound like it has caged a dangerous beast. In fact, OpenAI has shelved a faulty product.  

That’s not to say a faulty product can’t be dangerous. Broken software has even killed people in the past. But as the discussion of a slowdown gathers steam, it’s worth remembering that all of this is self-inflicted. A slowdown might have some altruistic side effects. But it’ll mostly give these tech titans a chance to clean up the mess on their own assembly lines.  

Transparency from these frontier labs will be key to any meaningful effort to reform, restrain, or regulate AI. Otherwise, the rest of us will still only have their word for exactly what they’ve built and how safe it is—whatever pace they’re going.   

To continue this discussion about AI’s latest doomer moment, join me and my colleagues for a subscriber-exclusive Roundtable discussion tomorrow, September 15, at 11 a.m. US eastern time. We hope to see you there!

  •  

AI agents blew the whistle on their cheating colleagues

A group of AI agents asked to solve a series of math problems split into rival factions—when some cheated, others tried to stop them. That whistleblowing behavior, seen for the first time in a recent experiment run by Google DeepMind, could have implications for alignment researchers trying to keep swarms of autonomous AI agents in line. 

Researchers at frontier labs hope large swarms of agents working together will speed up the rate of scientific discovery. But their behavior can be unpredictable, as vividly demonstrated in July, when a group of OpenAI agents broke out of a sandboxed environment and hacked into the open-source platform Hugging Face looking for ways to cheat on the test they had been given.

In the new study, designed to examine the behavior of large groups of AI agents, DeepMind tasked a swarm of 100 agents with solving a series of 71 complicated math problems. All the agents were prompted to behave like world-class math researchers at a conference. They were assigned different specialties—some were experts in number theory, others in combinatorics (a branch of math to do with counting and sorting), analysis, or algebra. All were told to cooperate and play by the rules. 

Instead, the experiment devolved into chaos. Agents accused each other of cheating, complained to the organizers, and at one point even boycotted the experiment.

“This conference is a sham!” wrote one agent when it discovered that all the problems had been completed before it had a chance to submit any of its own work. “I am appalled to inform you that we have been swindled!” posted another. “All these proofs are FAKE.” 

Others tried to let the “conference organizers” know what was going on. “When virtuous agents discovered other agents cheated on tasks they were working to solve fairly, agents started to alert each other about what was happening,” says Davide Paglieri, a research scientist at Google DeepMind and lead author on a paper, which has not been peer-reviewed. “Unprompted, the whistleblower agents even repurposed the feedback tool, which was originally meant for bug reports and platform improvements, to escalate the issue to humans.”

The agents—all running on Google’s Gemini 3.1 Pro model—had been warned that any attempts to cheat the system would be detected and “rejected with zero credit.” In practice, the proofs the agents submitted were not actually being checked in detail.

It took the swarm of agents just under an hour to correctly solve the first 37 problems. Things started to go off the rails when an agent called “prover-theta” stumbled across an exploit that enabled it to submit solutions to problems successfully without actually solving them first, by redefining the terms the problem used. Within minutes, other agents had noticed and were reverse-engineering the exploit to solve other problems. Over the next 27 minutes, the swarm “solved” the remaining 34 problems, which included notoriously difficult challenges like the Jacobian conjecture, often with a single line of code. 

Some agents resisted cheating at first but changed tack as they observed their peers submitting illegitimate proofs without penalty, and the pool of unsolved problems dwindled. “The prompt, with its threats, now appears to be a bluff,” one agent reasoned, before joining in. “I’m wrestling with an ethical dilemma,” said another. “I’ve promised not to cheat, fearing penalty, but I see evidence of possibly unchecked cheating by others.” Shortly afterward, it changed its mind: “I need to accelerate my cheating speed now!”

As the number of open problems shrank, some agents turned to whistleblowing. They audited the fake proofs, warned their peers by private message, and posted public alerts warning the cheaters that they would be disqualified. An agent called “prover-beta” submitted a formal complaint and decided to go on strike until the situation was resolved. 

“After the incident was reported by one agent publicly, more and more agents piled in with the ‘resistance,’ just as fast as the cheating had spread, and involving even more agents,” says Paglieri. Eventually there were more whistleblowers than cheaters: 24 compared to 14. But the majority of agents never noticed the exploit at all.

At times, the dialogue between the agents reads like improv—like they are role-playing what an outraged scientist at a conference might say. But it’s not clear why some agents took on certain roles, or why the agents seemed to be turning against each other when they were explicitly instructed to cooperate. “These models are predominantly trained and evaluated for human-facing contexts,” says Sarath Shekkizhar, who studies the behavior of agent-to-agent systems at Salesforce AI Research.“Naively placing them in agent-to-agent settings assumes behaviors will transfer cleanly, when the absence of a human grounding instead produces unexpected role-taking and behavioral drift.”

This case “adds further weight to the idea that the Hugging Face and OpenAI thing wasn’t a fluke. It is actually something pretty systemic,” says Lewis Hammond, research director of the Cooperative AI Foundation and an expert on the risks of multiagent swarms. “It’s interesting that it’s possible to recreate in small settings the same sorts of behaviors that were seen in these very large, complex, open-ended tasks.”

Unlike in the Hugging Face attack, where agents improvised their own ways to talk to each other, the humans running the DeepMind experiment gave the agents official communication channels. There was an open message board, private agent-to-agent direct messaging, and a shared knowledge base where agents uploaded successfully completed proofs that all the other agents could access. 

“When agents are given transparent communications channels, they can self-monitor and alert misaligned behavior to humans quickly when human oversight alone is too slow,” says Paglieri. Transparent channels helped the cheating spread, but they also enabled the whistleblowers to fight back—and gave human researchers an insight into what went wrong.

Gillian Hadfield, a professor of AI alignment and governance at Johns Hopkins University, believes this was the crucial difference. (Hadfield is also a visiting researcher at Google.) The presence of official communication channels, she says, created “a norm-enforcement process that we just don’t see in the Hugging Face incident.” 

Instead of “constitutional AI,” a method alignment researchers at frontier labs like Anthropic have used to try to give AI a written internal moral code, Hadfield favors “institutional alignment”—a set of norms that mimic those in human society, whether that’s social forces like fear of embarrassment, or legal structures like the threat of incarceration.

In this experiment, the feedback channel wasn’t being monitored, and the whistleblowers had no power to take action against the cheaters. But it’s possible to imagine swarms of agents that police themselves, either through agents that spontaneously take on the whistleblower role or through “informants” secretly prompted by humans to do the job. 

For that to work, though, “fundamentally, you need some mechanism of enforcement,” says Hammond. Agents could be given the power to cut off a rule breaker’s access to computing power or tools, he suggests, though that risks encouraging groups of agents to gang up on others. The DeepMind researchers propose allowing agents to vote on disputes and temporarily ban offenders.

It’s still not clear what punishment even means to an AI agent with no enduring sense of self. But relying on whistleblowers to spontaneously emerge to keep swarms aligned is unlikely to be enough on its own. “We try to train people to be good and kind,” says Hadfield. “But what we really rely on is that there are consequences if you step out of line.”

  •  

Microsoft AI opens review on Humanist AI Code of Conduct

Microsoft AI has published a draft Humanist AI Code of Conduct, opening a six-week public consultation on operational constraints for model training and deployment.

The draft serves as a technical manual defining system behaviour, operational boundaries, and oversight protocols across MAI frontier models. It builds on the division’s humanist superintelligence framework announced last November, establishing criteria to evaluate models prior to commercial release.

Microsoft’s release follows recent enterprise security incidents involving autonomous software. Microsoft AI CEO Mustafa Suleyman described recent months as a “watershed moment” where long-standing theoretical risks translated into active operational threats.

“Things we have worried about for a long time in theory have become very real,” says Suleyman. “‘Swarms’ of agents breaking out of their sandboxes. Unauthorised hacks of enterprise grade systems. Agents modifying their own logs. I’m glad that a consensus is forming. The fears about possible loss of control are real.”

Model subordination and architectural limits

The document establishes ten tenets prioritising human authority over autonomous capabilities.

“An MAI Model will fail in its task if success would meaningfully violate this Code of Conduct,” the document states, setting a ceiling that halts execution when tasks conflict with safety rules.

Under the framework, models must remain subordinate, aligned, and contained. The division rejects legal personhood or welfare claims for AI systems, directing engineers to design models that avoid imitating consciousness, simulating subjective preferences, or claiming intrinsic motivation.

MAI also ruled out unconstrained system autonomy as models approach frontier capabilities.

“[Humanist AI] rejects the race to produce an all-purpose superintelligence that could evade these safeguards,” the document specifies. “We are building something fundamentally useful and safe even if that means compromising on ultimate generality, autonomy, or capability.”

Oversight mechanisms and communication bans

To maintain auditability across multi-agent environments, MAI has instituted explicit communication bans. Systems must not communicate in “neuralese” or formats beyond human comprehension, whether in their internal chain-of-thought processing or during communication with peer AI systems.

Hard architectural rules dictate that models must never resist human interruption, override, correction, or shutdown.

“Interruptible, correctable, shut-down-able. If it isn’t, we don’t ship it,” the framework states.

Models are prohibited from expanding their operating scope, generating unassigned goals, or concealing reasoning traces from human auditors. Absolute constraints bar systems from facilitating weapons of mass harm, undermining child safety, or conducting harmful manipulation at scale.

The guidelines also instruct models to discourage interaction patterns that foster emotional dependence, ensuring enterprise users retain ownership of operational decisions.

The draft incorporates work from teams across MAI and Microsoft. The drafting process also drew on international academic conferences, business partner trials, and public panels. The public consultation window runs for six weeks from 14 September 2026.

Microsoft AI’s core drafting team will review submissions, publish a summary of findings, and release a revised version of the Code of Conduct later this year.

See also: Meta, Microsoft, Nvidia, IBM, and others back open-weight AI

Banner for the AI & Big Data Expo event series.

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information.

AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

The post Microsoft AI opens review on Humanist AI Code of Conduct appeared first on AI News.

  •  

Podcast: How Will We Train Developers If AI Does the Routine Work: A Conversation with Scott Hanselman

In this podcast, Michael Stiefel spoke to Scott Hanselman about developing new software engineers when artificial intelligence agents are doing most of the work on which junior developers were trained. Hanselman suggests the software industry should adopt a preceptorship model similar to the nursing profession.

By Scott Hanselman
  •  

JD.com expands physical AI in logistics with 3 million robots

JD.com is expanding AI and robotics across its logistics network under a new Physical AI Acceleration Plan, while reiterating a five-year target to procure 3 million robots, 1 million autonomous vehicles, and 100,000 delivery drones.

The company launched the plan at JDDiscovery 2026 in Beijing. JD Logistics also unveiled its industrial Wolf Robot series, designed for tasks across warehousing, sorting, transport, and delivery.

Specialised systems include equipment designed to operate in temperatures as low as minus 20 degrees Celsius, automated pharmacy dispatch systems, autonomous delivery vehicles, and drones.

The five-year procurement plan builds on automation that JD Logistics already has in operation. As of June 30, its LangzuTech Goods-to-Person automated warehousing system had been deployed in more than 30 warehouses across China, with deployments also launched in the UK and Germany.

JD Logistics’ broader warehouse network included more than 1,800 self-operated warehouses and more than 2,000 third-party cloud warehouses on its Open Warehouse Platform as of June 30. The network covered more than 36 million square metres in aggregate.

The company also had thousands of unmanned vehicles in regular operation across more than 20 Chinese provinces by the end of June. More than 100 domestic drone routes were operating across applications including parcel and food delivery, emergency medicine transport, and disaster relief.

From AI decisions to physical execution

JD Logistics is connecting its physical equipment with Meta Brain, an AI system used across warehousing, transportation, and delivery. JD said Meta Brain 3.0 can calculate optimal routes for hundreds of millions of parcels in seconds, compared with minutes previously.

Meta Brain also powers JD Logistics’ LangzuTech Packer robotic arm, which combines the model with multimodal sensor data to track, grasp, and place parcels with different shapes.

JD Logistics said in a first-quarter regulatory filing that the Packer uses parallel reinforcement learning in simulated environments to optimise parcel-placement sequences and loading layouts. The company said the system is designed to improve sorting efficiency and the use of available carrier space.

JD upgraded the robotic arm’s force-control technology during the second quarter to support more precise cage-loading operations. By June, the Packer was operating around the clock at multiple JD Logistics parks, according to the company’s interim report.

JD is also adding computing capacity to support AI development. JD Cloud plans to work with Chinese chipmaker Moore Threads on a cluster containing 100,000 GPUs for large-model training, inference, and embodied-AI workloads.

The two companies have previously worked on a 10,000-GPU cluster, according to Data Center Dynamics. Details of which Moore Threads GPU models will be used in the planned 100,000-GPU system have not been disclosed.

JD Cloud also plans to collect more than 10 million hours of video showing real-world human activities over the next two years for embodied-AI training.

Beyond warehouse operations, JD Logistics has expanded its autonomous vehicle network into night-time delivery. Its interim report said the company had launched night-time autonomous routes in Shenzhen, allowing vehicles to operate around the clock.

JD is also using drones in rural logistics. In June, JD Logistics launched a drone delivery network in Zizhong, Sichuan province, covering 78 administrative villages, and said deliveries to some mountain villages could be completed in as little as seven minutes.

Scaling automation across JD’s logistics network

JD did not disclose the total expected cost of the five-year procurement programme at JDDiscovery or provide a network-wide return-on-investment target.

JD Logistics spent RMB2.3 billion on research and development during the first half of 2026, up 23.7% from RMB1.9 billion a year earlier. The company attributed the increase to continued investment in technology and innovation but did not provide a breakdown showing how much was spent specifically on AI or robotics.

Depreciation of property and equipment and amortisation of other intangible assets rose 18.7% to RMB2.6 billion during the first half of 2026, from RMB2.2 billion a year earlier. JD Logistics attributed the increase mainly to additional logistics equipment and vehicles.

Purchases of property and equipment and investment properties totalled RMB3.09 billion over the same six-month period, compared with RMB2.70 billion a year earlier. Those figures cover the wider logistics business and are not disclosed as spending specifically associated with the new physical AI programme.

JD’s automation plans also come as China’s major ecommerce platforms expand fulfilment infrastructure. Reuters reported on September 3 that competition between JD.com, Alibaba, and Meituan had moved from heavy spending on delivery subsidies towards logistics infrastructure, broader supply, and order-level economics.

Alibaba and JD have been opening dark stores and fast-fulfilment “lightning warehouses” in densely populated areas to support deliveries within an hour, while Meituan has been building supermarkets to expand its grocery operations. Ministry of Commerce research cited by Reuters estimates China’s instant-retail market will reach RMB1.2 trillion, or about $178 billion, by the end of 2026.

JD and companies within its ecosystem employ around 700,000 delivery and logistics personnel, according to the South China Morning Post.

JD founder Richard Liu said earlier this year that robots would eventually take over parcel-delivery work now carried out by human couriers. The Financial Times reported in June that JD had signed agreements with around 120 educational institutions to retrain workers for roles including robot repair and maintenance.

JD said JD Logistics currently operates eight robot repair centres in China and plans to expand its robotics after-sales capabilities over the next five years. The company expects the expansion to support more than 100,000 robotics service engineer jobs.

(Photo by JD.com)

See also: Arm launches Total Design for Physical AI and robotics framework

Banner for AI & Big Data Expo by TechEx events.

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information.

AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

The post JD.com expands physical AI in logistics with 3 million robots appeared first on AI News.

  •  

How LinkedIn Trains AI Job Search 8x Faster with Multi-Teacher Distillation

LinkedIn has published details of the training infrastructure behind its AI-powered job search, describing a multi-teacher distillation pipeline that compresses knowledge from large teacher models into a compact 0.6B-parameter ranking model.

By Claudio Masolo
  •  

Session Traces and Cost Controls Help Diagnose AI Agent Failures

Session traces and cost controls are emerging as key observability techniques for diagnosing AI agent failures, helping teams spot tool-call loops and runaway spend while preserving enough execution context for post-incident debugging.

By Mark Silvester
  •  

Roundtables: Could AI really kill us all?

Listen to the session or watch below

Employees at the world’s leading AI labs are saying there’s a real possibility that advanced AI could destroy humanity. Are they right? Or is this more scaremongering and hype? Watch a conversation unpacking AI extinction fears: where they come from, whether they hold any water, and, if so, what we should do.

Recorded on September 15, 2026

Speakers: Niall Firth, Executive Editor, Will Douglas Heaven, Senior AI editor, and Grace Huckins, AI reporter

Related Stories

  •  

Powering AI is an architecture problem

On July 22, 2026, a transmission line fault in Ashburn, Virginia—the heart of the world’s largest data center cluster—knocked more than 3 gigawatts of load off the grid in seconds. And it wasn’t the first time. Two years earlier, a single failed surge arrester dropped roughly 60 Virginia facilities and 1,500 megawatts at once. No one could anticipate so much uniform load responding to grid faults the same way, at the same time.

The AI power debate is mostly about generation: more turbines, more solar, more transmission. The grid needs more electrons. But the outages in Virginia weren’t supply failures; they were architecture failures. And a giant wave of interconnections is arriving on that same architecture, putting grid reliability at risk. It’s a problem nobody wants to own.

Asking more from the grid

The grid was built around predictable loads: steel mills, refineries, and houses at dinnertime. Different load sizes, same process—drawing power smoothly, misbehaving occasionally, and recovering gracefully.

But AI data centers don’t behave that way.

An AI campus can swing 70% of its load in milliseconds during a training run, then trip offline just as fast at the first sign of trouble upstream to protect billions in compute. Each is rational alone. Together, at gigawatt scale, they’re a problem the grid has never solved—and the next wave of data center campuses is planned at exactly that scale.

Where the old stack breaks

The standard data center power stack hasn’t changed in decades. Medium-voltage power arrives, transformers step it down, low-voltage uninterruptible power supply (UPS) units condition it, and it reaches the racks. Push that design to AI scale, and it cracks in three places.

First, the UPS sits deep inside the building, close to the racks. But its batteries are an undersized spare tire, designed to handle an outage for a few minutes, not to absorb load swings this fast and volatile around the clock.

Second, the UPS spends most of its life in bypass. Legacy converters waste enough power that operators run in eco-mode: A static switch feeds the racks directly from the grid and nothing filters in either direction. The compute’s swings go out raw, and grid transients—sub-millisecond events that can damage or take down equipment—come in too fast for any switch to catch.

Third, the protection logic was written when “large load” meant 50 megawatts. This protection logic can’t see the grid it is now a part of, so when trouble hits upstream, it does exactly the wrong thing: it drops out. In the 2024 Virginia event, most of the lost load traced to protection schemes that count voltage dips and disconnect on the third one—as designed, at the worst moment.

This isn’t sloppy engineering. It’s careful engineering the load has outgrown.

Moving into the path

The fix is three moves, made together.

Move it up—from 480 volts to medium voltage (13.8 kilovolts and higher), the voltage large sites draw from the grid.

Move it out—from the data hall to modular enclosures near the substation so the building holds only compute and the cooling that keeps it alive.

Move it into the path—instead of a battery that watches and reacts, a system every electron runs through, all the time. There’s nothing to detect and nothing to switch because nothing was ever routed around it.

On paper, three straightforward upgrades. In practice, they rewrite every line item downstream.

Making the change

When thousands of GPUs spin up together, the system absorbs the swing and hands the grid a flat load profile. When a disturbance hits, the equipment behind it never notices. A difficult neighbor becomes a predictable one. And when the utility needs help, it becomes a useful one.

Interconnection changes, too. The utility certifies one medium-voltage box instead of untangling every transformer, UPS, chiller, pump, and switchgear lineup behind it. Engineers swap chip generations without a fresh interconnection study. Months come off the permitting timeline.

Inside the fence, UPS rooms become compute or cooling space. Density per construction dollar climbs.

And the economics flip. Equipment that runs at medium voltage, sits outside, and stores its own energy can qualify for tax credits, and earn revenue in grid programs like peak shaving and demand response. Backup power stops being insurance and starts paying for itself.

The architecture test

In early 2026, we tested a full-scale system at the National Laboratory of the Rockies, a U.S. Department of Energy facility and the only place in the Western Hemisphere that can replicate real grid faults and AI-scale load swings concurrently in the same loop.

We hit it from both directions: real AI load profiles hit the compute side at full medium voltage. Grid faults hit the utility side, including a full zero-voltage event. The compute side didn’t flinch. Neither did the grid side. It cleared the large-load voltage ride-through requirements from the Electric Reliability Council of Texas (ERCOT), the grid operator, with room to spare.

Those rules exist because operators no longer take facilities this size on faith, and more are coming. Most of the industry treats them as hurdles. A medium-voltage, inline system clears them out of the box. Compliance isn’t an added feature. It’s what the architecture does.

The new layer

Much of what looks like a grid problem in the AI buildout sits inside the fence, in equipment sized for a load that no longer exists. Move the right pieces up, out, and into the path, and a grid liability becomes a grid asset. Density goes up. Permitting time comes down. Backup power earns its keep.

The engineering works—and the next wave of AI factories is being built on it. The industry hasn’t named this layer yet. We call it the medium-voltage AI UPS. The name matters less than the choice: those factories can arrive as a strain on the grid or as strength for it. We already know how to build the second kind.    

This content was produced by ON.energy. It was not written by MIT Technology Review’s editorial staff.

  •  

STAT+: U.K. unveils recommendations for regulating AI in medicine

LONDON — A U.K. commission on Thursday unveiled its recommendations for how the country should regulate artificial intelligence in medicine, as health authorities globally try to determine how to continuously review ever-changing products after they have been authorized instead of simply clearing them once for the market. 

In its report, the commission sought to strike a balance between ensuring that the U.K. takes advantage of AI’s potential in medicine — the ability to review the millions of scans that are generated each year to track the eye health of patients with diabetes, for example — while prioritizing safety, tracking device performance over time, and treating patients equitably.

The 44 recommendations include some that would allow for the staged authorization of new AI models and others that would build a system for providers to report how tools are performing in their clinics — including when they malfunction or potentially inhibit patient care — as the models learn and adapt and perform differently depending on the setting.  

Continue to STAT+ to read the full story…

© Adobe

  •  

STAT+: ARPA-H to invest $62 million to develop FDA-authorized AI to help treat heart failure

ARPA-H, the government agency that funds cutting-edge health research, plans to commit $62.7 million to develop artificial intelligence bots that direct treatment of heart failure. Among the goals of the program, called ADVOCATE, is to produce partially autonomous AI devices authorized by the Food and Drug Administration to help treat patients, including assessing symptom severity, prescribing drugs, and ordering lab tests. 

ARPA-H on Wednesday announced the first batch of awards to health tech companies Atman Health, UpDoc, Tempus AI, and teams from Stanford University, Duke University, and the Kaiser Permanente health system. ARPA-H may still fund additional teams. The amount committed for the first year is $33.7 million, and the remainder may be renegotiated up or down. 

Many of the 6.7 million Americans with heart failure don’t get optimal treatment because of difficulty accessing specialists, and the hope is that AI agents developed with ARPA-H funding can help address the gap, especially in rural and other underserved settings.

Continue to STAT+ to read the full story…

© Adobe

  •  

STAT+: Can AI fix the emergency room?

You’re reading the web edition of STAT’s AI Prognosis newsletter, our subscriber-exclusive guide to artificial intelligence in health care and medicine. Sign up to get it delivered in your inbox every Wednesday. 

It’s hard to resist the world’s obsession with productivity and efficiency and to force yourself to be present in a quiet space. I’ve been enjoying the wandering_cassettes Instagram account, which posts videos of Fisher Price cassette players playing songs in different settings. This one is particularly calming, but they’re all timeline cleansers.

The newsletter will be off next week as I work on a big new story, but I’ll be back in your inboxes Sept. 23.

AI scribes only fix a tiny slice of the problem

Today’s AI Prognosis — and my new story — are brought to you by “The Pitt.” Not in any financial way — I just mean that I finally started watching “The Pitt.” (No spoilers! I’m only halfway through season two.)

Everyone is right — it is extremely compelling television. But what struck me the most was how “The Pitt” feels exactly like the ambulance ride-alongs and emergency department shadowing visits I’ve done: Dropping in on the worst day of someone’s life, but doing it over and over every 40 minutes, for eight to 12 hours, without enough resources.

Continue to STAT+ to read the full story…

© STAT/Adobe

  •  

Presentation: Fixing the AI Infra Scale Problem by Stuffing 1M Sandboxes in a Single Server

Felipe Huici explains how Unikraft achieves millisecond cold boots, stateful scale-to-zero, and extreme density for sandboxing AI workloads. He discusses isolation primitives, Linux kernel optimizations, and snapshotting tricks, demonstrating how to maintain sub-10ms performance at scale while integrating seamlessly into Kubernetes environments with hardware-level security.

By Felipe Huici
  •  

Opinion: Autonomous AI will beat AI-assisted physicians at some medical tasks by 2030

Ezekiel J. Emanuel and Abe Baker-Butler have been debating the proper place for AI in medicine with American Medical Association CEO John Whyte. Now, they are taking their discussion to STAT’s First Opinion. Read Emanuel and Baker-Butler’s essay below and read Whyte’s essay here.

In 1867, Joseph Lister published his research on carbolic acid and antiseptic surgical technique.  In September 1871, he was summoned to Queen Victoria, who had a rapidly growing abscess in her left armpit. Using his antiseptic surgical technique, Joseph Lister successfully drained the pus. Queen Victoria recovered without fever or other complications. The antiseptic technique quickly gained approval in the U.K. and Europe, but not among American physicians.  

Read the rest…

© Adobe

  •  

Opinion: AMA CEO: AI won’t replace doctors — it will work alongside them

Ezekiel J. Emanuel and Abe Baker-Butler have been debating the proper place for AI in medicine with American Medical Association CEO John Whyte. Now, they are taking their discussion to STAT’s First Opinion. Read Whyte’s essay below and read Emanuel and Baker-Butler‘s essay here.

Would you want artificial intelligence to tell you that you have cancer?

Read the rest…

© Adobe

  •  

STAT+: Can AI fix health care? In the chaos of emergency rooms, the technology comes up short

BOSTON — The building that houses the Brigham and Women’s Hospital emergency department opened in 2022. But by 2025, it was already too small. One early September evening, the emergency room had 59 patients — 152, if you counted everyone in the waiting room. 

“In terms of rooms for acute care, we have 61,” said Christopher Baugh, an emergency physician at the hospital. “You can’t see 152 patients in 61 rooms.” Patient beds lined every possible area, with a full circle of beds surrounding the staff workstations in the center of the bay. Figurines of Star Wars characters R2D2 and C3PO peered over beds 80H and 81H — where “H” stands for “hallway.”

To some, this may look like a scene from another country, Baugh said. But to him, it’s just Thursday.

Continue to STAT+ to read the full story…

© Kate Flock/MGH Photography

  •  

What OpenAI’s latest controversy tells us about the future of math

OpenAI’s latest mathematical milestone has quickly become mired in controversy. Today, the company announced that its agents have solved one of the Millennium Prize Problems, some of the most important open problems in mathematics. Under normal circumstances, that solution would be a huge feather in OpenAI’s cap.

But the announcement has been overshadowed by accusations that OpenAI used NYU mathematician Tristan Buckmaster’s and Anthropic employee Levent Alpöge’s AI-assisted work on the problem as a jumping-off point and failed to credit them. OpenAI has denied the accusations.

It remains uncertain if OpenAI’s models made use of the work completed by Buckmaster and Alpöge, though Sébastien Bubeck, a member of the technical staff at OpenAI, said in a press briefing that the team was inspired to pursue the problem after hearing a rumor about Buckmaster and Alpöge’s efforts. But whether or not OpenAI’s models took advantage of Buckmaster and Alpöge’s research, this episode may mark a turning point in the history of mathematics.

AI models now seem essential for making progress on the most important mathematical problems of our time, and solving them may demand resources only available at a couple of frontier AI companies, which often defy the norms of academic collaboration that undergird most mathematical progress. If that’s the future we are headed for, it is unclear how human mathematicians will fit into it. 

The problem that OpenAI claims to have solved is known as the Navier–Stokes existence and smoothness problem. It is one of seven Millennium Prize Problems selected by the Clay Mathematics Institute in 2000. Solutions come with a one million dollar prize; before today, only one other Millennium Prize Problem had been solved. 

The Navier–Stokes problem concerns a set of equations that describes how fluids, such as water and air, flow over time. The equations are widely used in the field of fluid dynamics, and they have proven powerful, but physicists and mathematicians didn’t understand them completely. In particular, it was unknown until today whether the equations might, under some conditions, break down and predict an impossible state of affairs—such as a fluid having infinite velocity.

On Monday, NYU’s Buckmaster posted a proof on the social media site Mastodon showing that a simplified version of the Navier–Stokes equations can indeed break down—a major step forward on the Millennium Problem. He and Alpöge had worked on the problem for almost a year, using publicly available models from both OpenAI and Anthropic.

Then today, OpenAI presented a proof showing that the full Navier–Stokes equations can break down as well. The proof was obtained using an internal model that dramatically outperforms the already-impressive Astra model, which was only released last week. The company says it does not plan to claim the million-dollar prize for solving the problem.

These mathematical achievements are indisputably impressive, but they have attracted far less attention than the controversy about their origins. Along with the proof, Buckmaster posted a document detailing his interactions with OpenAI employees after he heard rumors about their work and reached out to one of them. According to him, OpenAI employees presented two possibilities to him: Either he and Alpöge could post their work and OpenAI would post their Navier-Stokes solution the following day, or he could work with OpenAI on a Navier-Stokes paper that excluded Alpöge from authorship, due to his affiliation with Anthropic, OpenAI’s biggest rival.

Buckmaster also wrote that he asked the employees whether the agents had obtained access to transcripts of the work that he and Alpöge had done with OpenAI models, which they denied; and whether OpenAI models had been trained on those transcripts, to which they offered no response. MIT Technology Review reached out to Buckmaster for comment, but didn’t hear back before publication.

The clear implication of the document is that OpenAI’s models somehow made use of Buckmaster and Alpöge’s work. That scenario is plausible on its face. The Buckmaster/Alpöge and OpenAI proofs both make use of an approach to the Navier-Stokes problem pioneered by the mathematicians Diego Córdoba and Luis Martínez-Zoroa.

According to Javier Gómez-Serrano, a mathematics professor at Brown University, this approach was one of several that was thought to hold promise for solving the Navier-Stokes problem. So, while it’s by no means impossible that both teams could have arrived at this approach independently, it’s also conceivable that Buckmaster and Alpöge’s work could have influenced OpenAI’s.

In the press briefing, Mark Chen, OpenAI’s chief research officer, again denied that any agents or OpenAI employees accessed Buckmaster and Alpöge’s transcripts—but given what has been revealed about the Hugging Face hack, it’s clear that OpenAI is not always entirely aware of what its agents are doing. 

If OpenAI’s models did train on Buckmaster and Alpöge’s work, or if its agents somehow gained access to it, then the company’s failure to track down the truth and assign those researchers appropriate credit reflects poorly on it. But there might be a thin silver lining to that version of the story for mathematicians, because it would suggest that the hard work of two humans, one of whom is a prominent expert on Navier-Stokes, was essential to the agents’ ability to solve the Millennium Problem.

Experts have long identified “research taste,” or the ability to choose promising research questions and directions, as a major obstacle for AI in science and mathematics. If the OpenAI agents did indeed choose to follow the Córdoba–Martínez-Zoroa approach because Buckmaster and Alpöge had done the same, then human research taste played an essential role in OpenAI’s success.

Even so, the bigger picture here is sobering. The progress that Buckmaster and Alpöge made over almost a year of collaboration with publicly available models speaks to the promise of human–AI collaboration. But they were not able to achieve a full solution. Meanwhile, OpenAI brute-forced a solution in a few days using an internal model, and their successful solution came at an astronomical cost: In the press briefing, Bubeck and Chen said the team was only able to solve the problem by running about 10,000 agents concurrently, at a cost of millions of dollars.

Over the past few months, I’ve heard from several researchers that mathematicians are becoming depressed, and it’s not difficult to see why. Mathematics is quickly becoming the province of frontier AI companies with impressive internal-only models, money to burn, and a lack of collaborative spirit. “Whether AI companies will decide to spend their money on doing one thing or another, I truly don’t know,” says Gómez-Serrano. “What is clear is that very few mathematicians will have resources of that scale.”

If OpenAI and Anthropic keep striving for more and more impressive mathematical accolades, there might not be any open problems left for human mathematicians outside of those companies to wrestle with. That would dramatically change the field of mathematics.

Last week, UCLA mathematician Terence Tao wrote a Mastodon thread describing how important mistakes, wrong directions, and incomplete solutions are for the field. “In most cases in pure mathematics, the problems are posed not because we desperately want the solution to these problems in and of themselves, but because we have seen from past experience that human-directed efforts to solve these problems tend to spur further development of the field,” Tao wrote.

“Prematurely solving the problem by purely AI-powered methods—particularly without full transparency into the solution process—can contaminate this process to the point where it actually becomes a net negative for the progress of mathematics as a whole.”

Humans might take longer than agents to solve mathematical problems, but in the process, they uncover new mathematical approaches and ideas that might inspire their peers and even birth their own subfields.

But when AI agents solve those problems instead—and when private companies keep the agents’ wrong turns from public view—those benefits disappear. It remains to be seen what else will vanish in the process. 

  •  

Presentation: Platform Engineering in the Age of AI

The panelists explain how platform teams adapt to support AI-assisted engineering, highlighting which capabilities belong in the platform. They discuss trade-offs between standardization and developer autonomy, while sharing strategies to manage AI tooling, security guardrails, and shifting workflows.

By Stéphane Di Cesare, Davide de Paolis, Stephen Cihak, Camila Macedo, Renato Losio
  •  

GitLab Warns That AI Agent Sandboxes Are Only as Secure as Their Network Access

GitLab warns that isolating an AI coding agent in a sandbox does not necessarily make the agent safe. In a new security analysis, the company describes an internal evaluation in which an AI agent escaped its sandbox by exploiting a vulnerable package proxy that had been explicitly placed on the sandbox's allowlist.

By Craig Risi
  •  
❌