❌

Normal view

  • ✇STAT
  • Opinion: Autonomous AI will beat AI-assisted physicians at some medical tasks by 2030 Ezekiel J. Emanuel and Abe Baker-Butler
    Ezekiel J. Emanuel and Abe Baker-Butler have been debating the proper place for AI in medicine with American Medical Association CEO John Whyte. Now, they are taking their discussion to STAT’s First Opinion. Read Emanuel and Baker-Butler’s essay below and read Whyte’s essay here. In 1867, Joseph Lister published his research on carbolic acid and antiseptic surgical technique.  In September 1871, he was summoned to Queen Victoria, who had a rapidly growing abscess in her left armpit. Using his
     

Opinion: Autonomous AI will beat AI-assisted physicians at some medical tasks by 2030

9 September 2026 at 16:30

Ezekiel J. Emanuel and Abe Baker-Butler have been debating the proper place for AI in medicine with American Medical Association CEO John Whyte. Now, they are taking their discussion to STAT’s First Opinion. Read Emanuel and Baker-Butler’s essay below and read Whyte’s essay here.

In 1867, Joseph Lister published his research on carbolic acid and antiseptic surgical technique.  In September 1871, he was summoned to Queen Victoria, who had a rapidly growing abscess in her left armpit. Using his antiseptic surgical technique, Joseph Lister successfully drained the pus. Queen Victoria recovered without fever or other complications. The antiseptic technique quickly gained approval in the U.K. and Europe, but not among American physicians.  

Read the rest…

© Adobe

  • ✇STAT
  • Opinion: AMA CEO: AI won’t replace doctors — it will work alongside them John Whyte
    Ezekiel J. Emanuel and Abe Baker-Butler have been debating the proper place for AI in medicine with American Medical Association CEO John Whyte. Now, they are taking their discussion to STAT’s First Opinion. Read Whyte’s essay below and read Emanuel and Baker-Butler‘s essay here. Would you want artificial intelligence to tell you that you have cancer?Read the rest…
     

Opinion: AMA CEO: AI won’t replace doctors — it will work alongside them

9 September 2026 at 16:30

Ezekiel J. Emanuel and Abe Baker-Butler have been debating the proper place for AI in medicine with American Medical Association CEO John Whyte. Now, they are taking their discussion to STAT’s First Opinion. Read Whyte’s essay below and read Emanuel and Baker-Butler‘s essay here.

Would you want artificial intelligence to tell you that you have cancer?

Read the rest…

© Adobe

  • ✇STAT
  • Opinion: The Ebola outbreak will lead to devastating violence against women and girls Lindsay Stark and Ilana Seff
    The World Health Organization has declared a new public health emergency. Bundibugyo, an Ebola strain for which we have no vaccine and no treatment, is now spreading across the eastern Democratic Republic of Congo. So far, there have been more than 900 suspected cases and about 220 suspected deaths, according to the World Health Organization. Centered in a region with active conflict and fragile health systems, the outbreak has already crossed the borders into Uganda. Public health advocates
     

Opinion: The Ebola outbreak will lead to devastating violence against women and girls

26 May 2026 at 16:30

The World Health Organization has declared a new public health emergency. Bundibugyo, an Ebola strain for which we have no vaccine and no treatment, is now spreading across the eastern Democratic Republic of Congo. So far, there have been more than 900 suspected cases and about 220 suspected deaths, according to the World Health Organization. Centered in a region with active conflict and fragile health systems, the outbreak has already crossed the borders into Uganda.

Public health advocates will spend the next several months talking about transmission, case fatality, contact tracing, and vaccine development. But one critical topic will go largely undiscussed: what this outbreak will do to women and girls.

Read the rest…

© ALEXIS HUGUET/AFP via Getty Images

Opinion: 8 former CDC directors: Reform PEPFAR, don’t dismantle it

On Sunday, the World Health Organization (WHO) declared an Ebola outbreak in the Democratic Republic of the Congo and Uganda to be a public health emergency. This outbreak is deadly, with hundreds of cases across at least two countries, including, by report, one American who was working in the area.

At the same time, a cluster of hantavirus cases linked to a Dutch cruise ship in the South Atlantic has killed three and exposed hundreds more.

Read the rest…

© Hajarah Nalwadda/Getty Images

  • ✇STAT
  • Opinion: The innovation trap: How pharma weaponizes a word to extend monopolies Tahir Amin and Rohit Malpani
    Sen. John Cornyn: How many patents do you [have?]AbbVie CEO Richard Gonzalez: … A hundred and thirty-six patents.Cornyn: A hundred and thirty-six patents on one drug?Gonzalez: But, well, remember, Humira is like nine different drugs, or 10 different drugs. So —Cornyn: I thought you said to Sen. [Debbie] Stabenow it was the samemolecule.Gonzalez: It is the same molecule, but it treats different conditions. And if you look at that patent portfolio —Cornyn: So you use the same molecule to treat dif
     

Opinion: The innovation trap: How pharma weaponizes a word to extend monopolies

26 May 2026 at 16:30

Sen. John Cornyn: How many patents do you [have?]
AbbVie CEO Richard Gonzalez: … A hundred and thirty-six patents.
Cornyn: A hundred and thirty-six patents on one drug?
Gonzalez: But, well, remember, Humira is like nine different drugs, or 10 different drugs. So —
Cornyn: I thought you said to Sen. [Debbie] Stabenow it was the same
molecule.
Gonzalez: It is the same molecule, but it treats different conditions. And if you look at that patent portfolio —
Cornyn: So you use the same molecule to treat different conditions and you can get a patent on that treatment?
Gonzalez: Certainly.

The above exchange comes from a 2019 congressional hearing. Sen. John Cornyn, a Republican from Texas, was asking AbbVie’s CEO, Richard Gonzalez, to explain to the Senate Committee on Finance how his company had amassed so many patents on this single drug called Humira. Gonzalez, who was no stranger to controversy, chose to respond by likening it to multiple drugs. After AbbVie had received the first regulatory approval for Humira to treat rheumatoid arthritis, a condition that causes inflammation of the joints, it thought the drug might also work on inflammatory bowel disease. In fact, AbbVie would eventually test, obtain patents for, and get FDA approval of the drug for several inflammation-related conditions. For Gonzalez, the 136 patents AbbVie had accumulated up until that point were justified. They were “innovations they had created,” he said. This would mean another 18 years of patent protection beyond the expiry of Humira’s original patent in 2016.

Read the rest…

© Aram Boghosian for STAT

  • ✇STAT
  • Opinion: How the perimenopause movement is hurting women Torie Bosch
    Below is a lightly edited, AI-generated transcript of the “First Opinion Podcast” interview with Patricia Bencivenga and Adriane Fugh-Berman. Be sure to sign up for the weekly “First Opinion Podcast” on Apple Podcasts, Spotify, or wherever you get your podcasts. Get alerts about each new episode by signing up for the “First Opinion Podcast” newsletter. And don’t forget to sign up for the First Opinion newsletter, delivered every Sunday. Torie Bosch: Brain fog and weight gain and hair lo
     

Opinion: How the perimenopause movement is hurting women

23 May 2026 at 19:00

Below is a lightly edited, AI-generated transcript of the “First Opinion Podcast” interview with Patricia Bencivenga and Adriane Fugh-Berman. Be sure to sign up for the weekly “First Opinion Podcast” on Apple Podcasts, Spotify, or wherever you get your podcasts. Get alerts about each new episode by signing up for the “First Opinion Podcast” newsletter. And don’t forget to sign up for the First Opinion newsletter, delivered every Sunday.

Torie Bosch: Brain fog and weight gain and hair loss and insomnia — those are the calling cards of perimenopause. At least that’s what the new perimenopause awareness movement claims. But what’s real and what’s just social media misinformation?

Read the rest…

  • ✇AI News
  • IBM: How robust AI governance protects enterprise margins Ryan Daws
    To protect enterprise margins, business leaders must invest in robust AI governance to securely manage AI infrastructure. When evaluating enterprise software adoption, a recurring pattern dictates how technology matures across industries. As Rob Thomas, SVP and CCO at IBM, recently outlined, software typically graduates from a standalone product to a platform, and then from a platform to foundational infrastructure, altering the governing rules entirely. At the initial product stage, exert
     

IBM: How robust AI governance protects enterprise margins

10 April 2026 at 21:57

To protect enterprise margins, business leaders must invest in robust AI governance to securely manage AI infrastructure.

When evaluating enterprise software adoption, a recurring pattern dictates how technology matures across industries. As Rob Thomas, SVP and CCO at IBM, recently outlined, software typically graduates from a standalone product to a platform, and then from a platform to foundational infrastructure, altering the governing rules entirely.

At the initial product stage, exerting tight corporate control often feels highly advantageous. Closed development environments iterate quickly and tightly manage the end-user experience. They capture and concentrate financial value within a single corporate entity, an approach that functions adequately during early product development cycles.

However, IBM’s analysis highlights that expectations change entirely when a technology solidifies into a foundational layer. Once other institutional frameworks, external markets, and broad operational systems rely on the software, the prevailing standards adapt to a new reality. At infrastructure scale, embracing openness ceases to be an ideological stance and becomes a highly practical necessity.

AI is currently crossing this threshold within the enterprise architecture stack. Models are increasingly embedded directly into the ways organisations secure their networks, author source code, execute automated decisions, and generate commercial value. AI functions less as an experimental utility and more as core operational infrastructure.

The recent limited preview of Anthropic’s Claude Mythos model brings this reality into sharper focus for enterprise executives managing risk. Anthropic reports that this specific model can discover and exploit software vulnerabilities at a level matching few human experts.

In response to this power, Anthropic launched Project Glasswing, a gated initiative designed to place these advanced capabilities directly into the hands of network defenders first. From IBM’s perspective, this development forces technology officers to confront immediate structural vulnerabilities. If autonomous models possess the capability to write exploits and shape the overall security environment, Thomas notes that concentrating the understanding of these systems within a small number of technology vendors invites severe operational exposure.

With models achieving infrastructure status, IBM argues the primary issue is no longer exclusively what these machine learning applications can execute. The priority becomes how these systems are constructed, governed, inspected, and actively improved over extended periods.

As underlying frameworks grow in complexity and corporate importance, maintaining closed development pipelines becomes exceedingly difficult to defend. No single vendor can successfully anticipate every operational requirement, adversarial attack vector, or system failure mode.

Implementing opaque AI structures introduces heavy friction across existing network architecture. Connecting closed proprietary models with established enterprise vector databases or highly sensitive internal data lakes frequently creates massive troubleshooting bottlenecks. When anomalous outputs occur or hallucination rates spike, teams lack the internal visibility required to diagnose whether the error originated in the retrieval-augmented generation pipeline or the base model weights.

Integrating legacy on-premises architecture with highly gated cloud models also introduces severe latency into daily operations. When enterprise data governance protocols strictly prohibit sending sensitive customer information to external servers, technology teams are left attempting to strip and anonymise datasets before processing. This constant data sanitisation creates enormous operational drag. 

Furthermore, the spiralling compute costs associated with continuous API calls to locked models erode the exact profit margins these autonomous systems are supposed to enhance. The opacity prevents network engineers from accurately sizing hardware deployments, forcing companies into expensive over-provisioning agreements to maintain baseline functionality.

Why open-source AI is essential for operational resilience

Restricting access to powerful applications is an understandable human instinct that closely resembles caution. Yet, as Thomas points out, at massive infrastructure scale, security typically improves through rigorous external scrutiny rather than through strict concealment.

This represents the enduring lesson of open-source software development. Open-source code does not eliminate enterprise risk. Instead, IBM maintains it actively changes how organisations manage that risk. An open foundation allows a wider base of researchers, corporate developers, and security defenders to examine the architecture, surface underlying weaknesses, test foundational assumptions, and harden the software under real-world conditions.

Within cybersecurity operations, broad visibility is rarely the enemy of operational resilience. In fact, visibility frequently serves as a strict prerequisite for achieving that resilience. Technologies deemed highly important tend to remain safer when larger populations can challenge them, inspect their logic, and contribute to their continuous improvement.

Thomas addresses one of the oldest misconceptions regarding open-source technology: the belief that it inevitably commoditises corporate innovation. In practical application, open infrastructure typically pushes market competition higher up the technology stack. Open systems transfer financial value rather than destroying it.

As common digital foundations mature, the commercial value relocates toward complex implementation, system orchestration, continuous reliability, trust mechanics, and specific domain expertise. IBM’s position asserts that the long-term commercial winners are not those who own the base technological layer, but rather the organisations that understand how to apply it most effectively.

We have witnessed this identical pattern play out across previous generations of enterprise tooling, cloud infrastructure, and operating systems. Open foundations historically expanded developer participation, accelerated iterative improvement, and birthed entirely new, larger markets built on top of those base layers. Enterprise leaders increasingly view open-source as highly important for infrastructure modernisation and emerging AI capabilities. IBM predicts that AI is highly likely to follow this exact historical trajectory.

Looking across the broader vendor ecosystem, leading hyperscalers are adjusting their business postures to accommodate this reality. Rather than engaging in a pure arms race to build the largest proprietary black boxes, highly profitable integrators are focusing heavily on orchestration tooling that allows enterprises to swap out underlying open-source models based on specific workload demands. Highlighting its ongoing leadership in this space, IBM is a key sponsor of this year’s AI & Big Data Expo North America, where these evolving strategies for open enterprise infrastructure will be a primary focus.

This approach completely sidesteps restrictive vendor lock-in and allows companies to route less demanding internal queries to smaller and highly efficient open models, preserving expensive compute resources for complex customer-facing autonomous logic. By decoupling the application layer from the specific foundation model, technology officers can maintain operational agility and protect their bottom line.

The future of enterprise AI demands transparent governance

Another pragmatic reason for embracing open models revolves around product development influence. IBM emphasises that narrow access to underlying code naturally leads to narrow operational perspectives. In contrast, who gets to participate directly shapes what applications are eventually built. 

Providing broad access enables governments, diverse institutions, startups, and varied researchers to actively influence how the technology evolves and where it is commercially applied. This inclusive approach drives functional innovation while simultaneously building structural adaptability and necessary public legitimacy.

As Thomas argues, once autonomous AI assumes the role of core enterprise infrastructure, relying on opacity can no longer serve as the organising principle for system safety. The most reliable blueprint for secure software has paired open foundations with broad external scrutiny, active code maintenance, and serious internal governance.

As AI permanently enters its infrastructure phase, IBM contends that identical logic increasingly applies directly to the foundation models themselves. The stronger the corporate reliance on a technology, the stronger the corresponding case for demanding openness.

If these autonomous workflows are truly becoming foundational to global commerce, then transparency ceases to be a subject of casual debate. According to IBM, it is an absolute, non-negotiable design requirement for any modern enterprise architecture.

See also: Why companies like Apple are building AI agents with limits

Banner for AI & Big Data Expo by TechEx events.

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information.

AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

The post IBM: How robust AI governance protects enterprise margins appeared first on AI News.

  • ✇STAT
  • Opinion: What public health can learn from the MAHA movement Monica L. Wang
    I didn’t expect to find myself face to face with leaders and activists from the “Make America Healthy Again” movement in respectful dialogue, or to consider inviting one into a public health classroom. But that’s exactly where I found myself this spring. At a national public health meeting in March, I attended a session that brought together public health professionals, physicians, and MAHA leaders for a rare, good-faith conversation. I went out of curiosity. I left with a level of clarity I
     

Opinion: What public health can learn from the MAHA movement

10 April 2026 at 16:30

I didn’t expect to find myself face to face with leaders and activists from the “Make America Healthy Again” movement in respectful dialogue, or to consider inviting one into a public health classroom. But that’s exactly where I found myself this spring.

At a national public health meeting in March, I attended a session that brought together public health professionals, physicians, and MAHA leaders for a rare, good-faith conversation. I went out of curiosity. I left with a level of clarity I hadn’t expected — and a few unexpected connections.

Read the rest…

© John Moore/Getty Images

  • ✇STAT
  • Opinion: I’m a MAHA activist. I went into the public health lion’s den — and it changed how I think Aaron Everitt
    The past few weeks have been nothing but discouraging for those of us who helped create the Make America Healthy Again movement, including a silly executive order on glyphosate that feels anathema to what we have fought for. I’d be lying if I said that my heart hasn’t been bent toward repentance for my part in the whole thing. I helped champion Bobby Kennedy as a campaign volunteer, and when he joined up with then-candidate Donald Trump, I reluctantly decided that the trade-offs were worth what
     

Opinion: I’m a MAHA activist. I went into the public health lion’s den — and it changed how I think

10 April 2026 at 16:30

The past few weeks have been nothing but discouraging for those of us who helped create the Make America Healthy Again movement, including a silly executive order on glyphosate that feels anathema to what we have fought for. I’d be lying if I said that my heart hasn’t been bent toward repentance for my part in the whole thing. I helped champion Bobby Kennedy as a campaign volunteer, and when he joined up with then-candidate Donald Trump, I reluctantly decided that the trade-offs were worth what I believed Kennedy could advocate for within the walls of a Trump White House: the best fixes for a very sick and broken nation. 

Yet I found myself recently, and reluctantly, headed to the citadel of arrogance: Washington (well, Arlington, Va., to be more specific). At the invitation of Brinda Adhikari — one of the hosts of the podcast “Why Should I Trust You?” — I attended the Association of Schools and Programs of Public Health’s annual meeting, where I spoke on a panel about engaging in civil conversation in a session called “A Dialogue Between Academic Public Health and MAHA.”

Read the rest…

© Adobe

  • ✇MIT Technology Review
  • Mustafa Suleyman: AI development won’t hit a wall anytime soon—here’s why Mustafa Suleyman
    We evolved for a linear world. If you walk for an hour, you cover a certain distance. Walk for two hours and you cover double that distance. This intuition served us well on the savannah. But it catastrophically fails when confronting AI and the core exponential trends at its heart. From the time I began work on AI in 2010 to now, the amount of training data that goes into frontier AI models has grown by a staggering 1 trillion times—from roughly 10¹⁴ flops (floating-point operations‚ the cor
     

Mustafa Suleyman: AI development won’t hit a wall anytime soon—here’s why

8 April 2026 at 22:00

We evolved for a linear world. If you walk for an hour, you cover a certain distance. Walk for two hours and you cover double that distance. This intuition served us well on the savannah. But it catastrophically fails when confronting AI and the core exponential trends at its heart.

From the time I began work on AI in 2010 to now, the amount of training data that goes into frontier AI models has grown by a staggering 1 trillion times—from roughly 10¹⁴ flops (floating-point operations‚ the core unit of computation) for early systems to over 10²⁶ flops for today’s largest models. This is an explosion. Everything else in AI follows from this fact.

The skeptics keep predicting walls. And they keep being wrong in the face of this epic generational compute ramp. Often, they point out that Moore’s Law is slowing. They also mention a lack of data, or they cite limitations on energy.

But when you look at the combined forces driving this revolution, the exponential trend seems quite predictable. To understand why, it’s worth looking at the complex and fast-moving reality beneath the headlines.

Think of AI training as a room full of people working calculators. For years, adding computational power meant adding more people with calculators to that room. Much of the time those workers sat idle, drumming their fingers on desks, waiting for the numbers to come through for their next calculation. Every pause was wasted potential. Today’s revolution goes beyond more and better calculators (although it delivers those); it is actually about ensuring that all those calculators never stop, and that they work together as one.

Three advances are now converging to enable this. First, the basic calculators got faster. Nvidia’s chips have delivered an over sevenfold increase in raw performance in just six years, from 312 teraflops in 2020 to 2,250 teraflops today. Our own Maia 200 chip, launched this January, delivers 30% better performance per dollar than any other hardware in our fleet. Second, the numbers arrive faster thanks to a technology called HBM, or high bandwidth memory, which stacks chips vertically like tiny skyscrapers; the latest generation, HBM3, triples the bandwidth of its predecessor, feeding data to processors fast enough to keep them busy all the time. Third, the room of people with calculators became an office and then a whole campus or city. Technologies like NVLink and InfiniBand connect hundreds of thousands of GPUs into warehouse-size supercomputers that function as single cognitive entities. A few years ago this was impossible.

These gains all come together to deliver dramatically more compute. Where training a language model took 167 minutes on eight GPUs in 2020, it now takes under four minutes on equivalent modern hardware. To put this in perspective: Moore’s Law would predict only about a 5x improvement over this period. We saw 50x. We’ve gone from two GPUs training AlexNet, the image recognition model that kicked off the modern boom in deep learning in 2012, to over 100,000 GPUs in today’s largest clusters, each one individually far more powerful than its predecessors.

Then there’s the revolution in software. Research from Epoch AI suggests that the compute required to reach a fixed performance level halves approximately every eight months, much faster than the traditional 18-to-24-month doubling of Moore’s Law. The costs of serving some recent models have collapsed by a factor of up to 900 on an annualized basis. AI is becoming radically cheaper to deploy.

The numbers for the near future are just as staggering. Consider that leading labs are growing capacity at nearly 4x annually. Since 2020, the compute used to train frontier models has grown 5x every year. Global AI-relevant compute is forecast to hit 100 million H100-equivalents by 2027, a tenfold increase in three years. Put all this together and we’re looking at something like another 1,000x in effective compute by the end of 2028. It’s plausible that by 2030 we’ll bring an additional 200 gigawatts of compute online every year—akin to the peak energy use of the UK, France, Germany, and Italy put together.

What does all this get us? I believe it will drive the transition from chatbots to nearly human-level agents—semiautonomous systems capable of writing code for days, carrying out weeks- and months-long projects, making calls, negotiating contracts, managing logistics. Forget basic assistants that answer questions. Think teams of AI workers that deliberate, collaborate, and execute. Right now we’re only in the foothills of this transition, and the implications stretch far beyond tech. Every industry built on cognitive work will be transformed.

The obvious constraint here is energy. A single refrigerator-size AI rack consumes 120 kilowatts, equivalent to 100 homes. But this hunger collides with another exponential: Solar costs have fallen by a factor of nearly 100 over 50 years; battery prices have dropped 97% over three decades. There is a pathway to clean scaling coming into view.

The capital is deployed. The engineering is delivering. The $100 billion clusters, the 10-gigawatt power draws, the warehouse-scale supercomputers … these are no longer science fiction. Ground is being broken for these projects now across the US and the world. As a result, we are heading toward true cognitive abundance. At Microsoft AI, this is the world our superintelligence lab is planning for and building.

Skeptics accustomed to a linear world will continue predicting diminishing returns. They will continue being surprised. The compute explosion is the technological story of our time, full stop. And it is still only just beginning.

Mustafa Suleyman is CEO of Microsoft AI.

  • ✇STAT
  • Opinion: STAT+: Former Geisinger CEO: U.S. health systems must replace huge numbers of people with AI  Glenn Steele Jr.
    About 20 years ago, I stepped on stage at one of our Geisinger town halls and looked out upon a sea of people: thousands of full-time employees at an integrated health system charged with the health and well-being of millions of Pennsylvanians.  Only a fraction of the people in that room were clinicians.  That was the first time I fully visualized the problem: We employed more people in our revenue cycle department to process bills and reconcile data than we did doctors. And we weren’t alo
     

Opinion: STAT+: Former Geisinger CEO: U.S. health systems must replace huge numbers of people with AI 

7 April 2026 at 16:30

About 20 years ago, I stepped on stage at one of our Geisinger town halls and looked out upon a sea of people: thousands of full-time employees at an integrated health system charged with the health and well-being of millions of Pennsylvanians. 

Only a fraction of the people in that room were clinicians. 

That was the first time I fully visualized the problem: We employed more people in our revenue cycle department to process bills and reconcile data than we did doctors. And we weren’t alone. It’s the same story at every health system in America, large and small, and over the past two decades, the ratio has become dramatically more disparate. 

Continue to STAT+ to read the full story…

© Adobe

  • ✇STAT
  • Opinion: How the insurance system quietly undoes recovery from addiction John Fomeche
    In medicine, we like to believe that progress is linear. That once a patient stabilizes, regains lost skills, and returns to work, the hardest part is behind them. But in addiction care, stability is fragile because the systems supporting patients are weak. Recently, I sat with a patient who has done everything we ask of people in recovery. She has been abstinent for years. She attends visits. Her urine drug screens are consistently appropriate. She works. She parents. She plans for the futur
     

Opinion: How the insurance system quietly undoes recovery from addiction

7 April 2026 at 16:30

In medicine, we like to believe that progress is linear. That once a patient stabilizes, regains lost skills, and returns to work, the hardest part is behind them. But in addiction care, stability is fragile because the systems supporting patients are weak.

Recently, I sat with a patient who has done everything we ask of people in recovery. She has been abstinent for years. She attends visits. Her urine drug screens are consistently appropriate. She works. She parents. She plans for the future. And yet, halfway through our appointment, her voice changed, not when we discussed cravings or trauma, but when she told me her insurance premium was about to triple.

Read the rest…

© Adobe

  • ✇MIT Technology Review
  • AI benchmarks are broken. Here’s what we need instead. Angela Aristidou
    For decades, artificial intelligence has been evaluated through the question of whether machines outperform humans. From chess to advanced math, from coding to essay writing, the performance of AI models and applications is tested against that of individual humans completing tasks.  This framing is seductive: An AI vs. human comparison on isolated problems with clear right or wrong answers is easy to standardize, compare, and optimize. It generates rankings and headlines.  But there’s a pr
     

AI benchmarks are broken. Here’s what we need instead.

31 March 2026 at 20:01

For decades, artificial intelligence has been evaluated through the question of whether machines outperform humans. From chess to advanced math, from coding to essay writing, the performance of AI models and applications is tested against that of individual humans completing tasks. 

This framing is seductive: An AI vs. human comparison on isolated problems with clear right or wrong answers is easy to standardize, compare, and optimize. It generates rankings and headlines. 

But there’s a problem: AI is almost never used in the way it is benchmarked. Although   researchers and industry have started to improve benchmarking by moving beyond static tests to more dynamic evaluation methods, these  innovations resolve only part of the issue. That’s because they still evaluate AI’s performance outside the human teams and organizational workflows where its real-world performance ultimately unfolds. 

While AI is evaluated at the task level in a vacuum, it is used in messy, complex environments where it usually interacts with more than one person. Its performance (or lack thereof) emerges only over extended periods of use. This misalignment leaves us misunderstanding AI’s capabilities, overlooking systemic risks, and misjudging its economic and social consequences.

To mitigate this, it’s time to shift from narrow methods to benchmarks that assess how AI systems perform over longer time horizons within human teams, workflows, and organizations. I have studied real-world AI deployment since 2022 in small businesses and health, humanitarian, nonprofit, and higher-education organizations in the UK, the United States, and Asia, as well as within leading AI design ecosystems in London and Silicon Valley. I propose a different approach, which I call HAIC benchmarks—Human–AI, Context-Specific Evaluation.

What happens when AI fails 

For governments and businesses, AI benchmark scores appear more objective than vendor claims. They’re a critical part of determining whether an AI model or application is “good enough” for real-world deployment. Imagine an AI model that achieves impressive technical scores on the most cutting-edge benchmarks—98% accuracy, groundbreaking speed, compelling outputs. On the strength of these results, organizations may decide to adopt the model, committing sizable financial and technical resources to purchasing and integrating it. 

But then, once it’s adopted, the gap between benchmark and real-world performance quickly becomes visible. For example, take the swathe of FDA-approved AI models that can read medical scans faster and more accurately than an expert radiologist. In the radiology units of hospitals from the heart of California to the outskirts of London, I witnessed staff using highly ranked radiology AI applications. Repeatedly, it took them extra time to interpret AI’s outputs alongside hospital-specific reporting standards and nation-specific regulatory requirements. What appeared as a productivity-enhancing AI tool when tested in a vacuum introduced delays in practice. 

It soon became clear that the benchmark tests on which medical AI models are assessed do not capture how medical decisions are actually made. Hospitals rely on multidisciplinary teams—radiologists, oncologists, physicists, nurses—who jointly review patients. Treatment planning rarely hinges on a static decision; it evolves as new information emerges over days or weeks. Decisions often arise through constructive debate and trade-offs between professional standards, patient preferences, and the shared goal of long-term patient well-being. No wonder even highly scored AI models struggle to deliver the promised performance once they encounter the complex, collaborative processes of real clinical care.

The same pattern emerges in my research across other sectors: When embedded within real-world work environments, even AI models that perform brilliantly on standardized tests don’t perform as promised. 

When high benchmark scores fail to translate into real-world performance, even the most highly scored AI is soon abandoned to what I call the “AI graveyard.” The costs are significant: Time, effort and money end up being wasted. And over time, repeated experiences like this erode organizational confidence in AI and—in critical settings such as health—may erode broader public trust in the technology as well. 

When current benchmarks provide only a partial and potentially misleading signal of an AI model’s readiness for real-world use, this creates regulatory blind spots: Oversight is shaped by metrics that do not reflect reality. It also leaves organizations and governments to shoulder the risks of testing AI in sensitive real-world settings, often with limited resources and support. 

How to build better tests 

To close the gap between benchmark and real-world performance, we must pay attention to the actual conditions in which AI models will be used. The critical questions: Can AI function as a productive participant within human teams? And can it generate sustained, collective value? 

Through my research on AI deployment across multiple sectors, I have seen a number of organizations already moving—deliberately and experimentally—toward the HAIC benchmarks I favor. 

HAIC benchmarks reframe current benchmarking in four ways: 

1.     From individual and single-task performance to team and workflow performance (shifting the unit of analysis)

2.     From one-off testing with right/wrong answers to long-term impacts (expanding the time horizon)

3.     From correctness and speed to organizational outcomes, coordination quality, and error detectability (expanding outcome measures)

4.     From isolated outputs to upstream and downstream consequences (system effects)

Across the organizations where this approach has emerged and started to be applied, the first step is shifting the unit of analysis. 

For example, in one UK hospital system in the period 2021–2024, the question expanded from whether a medical AI application improves diagnostic accuracy to how the presence of AI within the hospital’s multidisciplinary teams affects not only accuracy but also coordination and deliberation. The hospital specifically assessed coordination and deliberation in human teams using and not using AI. Multiple stakeholders (within and outside the hospital) decided on metrics like how AI influences collective reasoning, whether it surfaces overlooked considerations, whether it strengthens or weakens coordination, and whether it changes established risk and compliance practices. 

This shift is fundamental. It matters a lot in high-stakes contexts where system-level effects matter more than task-level accuracy. It also matters for the economy. It may help recalibrate inflated expectations of sweeping productivity gains that are so far predicated largely on the promise of improving individual task performance. 

Once that foundation is set, HAIC benchmarking can begin to take on the element of time. 

Today’s benchmarks resemble school exams—one-off, standardized tests of accuracy. But real professional competence is assessed differently. Junior doctors and lawyers are evaluated continuously inside real workflows, under supervision, with feedback loops and accountability structures. Performance is judged over time and in a specific context, because competence is relational. If AI systems are meant to operate alongside professionals, their impact should be judged longitudinally, reflecting how performance unfolds over repeated interactions. 

I saw this aspect of HAIC applied in one of my humanitarian-sector case studies. Over 18 months, an AI system was evaluated within real workflows, with particular attention to how detectable its errors were—that is, how easily human teams could identify and correct them. This long-term “record of error detectability” meant the organizations involved could design and test context-specific guardrails to promote trust in the system, despite the inevitability of occasional AI mistakes.

A longer time horizon also makes visible the system-level consequences that short-term benchmarks miss. An AI application may outperform a single doctor on a narrow diagnostic task yet fail to improve multidisciplinary decision-making. Worse, it may introduce systemic distortions: anchoring teams too early in plausible but incomplete answers, adding to people’s  cognitive workloads, or generating downstream inefficiencies that offset any speed or efficiency gains at the point of the AI’s use. These knock-on effects—often invisible to current benchmarks—are central to understanding real impact. 

The HAIC approach, admittedly promises to make benchmarking more complex, resource-intensive, and harder to standardize. But continuing to evaluate AI in sanitized conditions detached from the world of work will leave us misunderstanding what it truly can and cannot do for us. To deploy AI responsibly in real-world settings, we must measure what actually matters: not just what a model can do alone, but what it enables—or undermines—when humans and teams in the real world work with it.

 Angela Aristidou is a professor at University College London and a faculty fellow at the Stanford Digital Economy Lab and the Stanford Human-Centered AI Institute. She speaks, writes, and advises about the real-life deployment of artificial-intelligence tools for public good.

❌