❌

Normal view

The diagnostic accuracy of wearable digital technology in detecting fertility window and menstrual cycles: a systematic review and Bayesian network meta-analysis

npj Digital Medicine, Published online: 24 January 2026; doi:10.1038/s41746-025-02320-8

The diagnostic accuracy of wearable digital technology in detecting fertility window and menstrual cycles: a systematic review and Bayesian network meta-analysis
  • ✇MIT Technology Review
  • The Download: chatbots for health, and US fights over AI regulation Charlotte Jee
    This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. “Dr. Google” had its issues. Can ChatGPT Health do better?   For the past two decades, there’s been a clear first step for anyone who starts experiencing new medical symptoms: Look them up online. The practice was so common that it gained the pejorative moniker “Dr. Google.” But times are changing, and many medical-information seekers are now using L
     

The Download: chatbots for health, and US fights over AI regulation

23 January 2026 at 21:07

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology.


“Dr. Google” had its issues. Can ChatGPT Health do better?  

For the past two decades, there’s been a clear first step for anyone who starts experiencing new medical symptoms: Look them up online. The practice was so common that it gained the pejorative moniker “Dr. Google.” But times are changing, and many medical-information seekers are now using LLMs. According to OpenAI, 230 million people ask ChatGPT health-related queries each week.  

That’s the context around the launch of OpenAI’s new ChatGPT Health product, which debuted earlier this month. The big question is: can the obvious risks of using AI for health-related queries be mitigated enough for them to be a net benefit? Read the full story. 

—Grace Huckins

America’s coming war over AI regulation  

In the final weeks of 2025, the battle over regulating artificial intelligence in the US reached boiling point. On December 11, after Congress failed twice to pass a law banning state AI laws, President Donald Trump signed a sweeping executive order seeking to handcuff states from regulating the booming industry.  

Instead, he vowed to work with Congress to establish a “minimally burdensome” national AI policy. The move marked a victory for tech titans, who have been marshaling multimillion-dollar war chests to oppose AI regulations, arguing that a patchwork of state laws would stifle innovation.

In 2026, the battleground will shift to the courts. While some states might back down from passing AI laws, others will charge ahead. Read our story about what’s on the horizon. 

—Michelle Kim

This story is from MIT Technology Review’s What’s Next series of stories that look across industries, trends, and technologies to give you a first look at the future. You can read the rest of them here.  

Measles is surging in the US. Wastewater tracking could help.

This week marked a rather unpleasant anniversary: It’s a year since Texas reported a case of measles—the start of a significant outbreak that ended up spreading across multiple states. Since the start of January 2025, there have been over 2,500 confirmed cases of measles in the US. Three people have died. 

As vaccination rates drop and outbreaks continue, scientists have been experimenting with new ways to quickly identify new cases and prevent the disease from spreading. And they are starting to see some success with wastewater surveillance. Read the full story.

—Jessica Hamzelou 

This story is from The Checkup, our weekly newsletter giving you the inside track on all things health and biotech. Sign up to receive it in your inbox every Thursday.

The must-reads

I’ve combed the internet to find you today’s most fun/important/scary/fascinating stories about technology.

1 The US is dismantling itself
A foreign enemy could not invent a better chain of events to wreck its standing in the world. (Wired $)  
+ We need to talk about whether Donald Trump might be losing it.  (New Yorker $)

2 Big Tech is taking on more debt to fund its AI aspirations
And the bubble just keeps growing. (WP $)
+ Forget unicorns. 2026 is shaping up to be the year of the “hectocorn.” (The Guardian)
+ Everyone in tech agrees we’re in a bubble. They just can’t agree on what happens when it pops. (MIT Technology Review)

3 DOGE accessed even more personal data than we thought 
Even now, the Trump administration still can’t say how much data is at risk, or what it was used for. (NPR)

4 TikTok has finalized a deal to create a new US entity 
Ending years of uncertainty about its fate in America. (CNN)
+ Why China is the big winner out of all of this. (FT $)

5 The US is now officially out of the World Health Organization 
And it’s leaving behind nearly $300 million in bills unpaid. (Ars Technica) 
+ The US withdrawal from the WHO will hurt us all. (MIT Technology Review)

6 AI-powered disinformation swarms pose a threat to democracy
A would-be autocrat could use them to persuade populations to accept cancelled elections or overturn results. (The Guardian)
+ The era of AI persuasion in elections is about to begin. (MIT Technology Review)

7 We’re about to start seeing more robots everywhere
But exactly what they’ll look like remains up for debate. (Vox $)
+ Chinese companies are starting to dominate entire sectors of AI and robotics. (MIT Technology Review)

8 Some people seem to be especially vulnerable to loneliness
If you’re ‘other-directed’, you could particularly benefit from less screentime. (New Scientist $)

9 This academic lost two years of work with a single click
TL;DR: Don’t rely on ChatGPT to store your data. (Nature)

10 How animals develop a sense of direction 🦇🧭
Their ‘internal compass’ seems to be informed by landmarks that help them form a mental map. (Quanta $)

Quote of the day

“The rate at which AI is progressing, I think we have AI that is smarter than any human this year, and no later than next year.”

—Elon Musk simply cannot resist the urge to make wild predictions at Davos, Wired reports. 

One more thing

ADAM DETOUR

Africa fights rising hunger by looking to foods of the past

After falling steadily for decades, the prevalence of global hunger is now on the rise—nowhere more so than in sub-Saharan Africa. 

Africa’s indigenous crops are often more nutritious and better suited to the hot and dry conditions that are becoming more prevalent, yet many have been neglected by science, which means they tend to be more vulnerable to diseases and pests and yield well below their theoretical potential.

Now the question is whether researchers, governments, and farmers can work together in a way that gets these crops onto plates and provides Africans from all walks of life with the energy and nutrition that they need to thrive, whatever climate change throws their way. Read the full story.

—Jonathan W. Rosen

We can still have nice things

A place for comfort, fun and distraction to brighten up your day. (Got any ideas? Drop me a line or skeet ’em at me.)

+ The only thing I fancy dry this January is a martini. Here’s how to make one.
+ If you absolutely adore the Bic crystal pen, you might want this lamp. 
+ Cozy up with a nice long book this winter. ($)
+ Want to eat healthier? Slow down and tune out food ‘noise’. ($)

Principles to guide clinical AI readiness and move from benchmarks to real-world evaluation

23 January 2026 at 08:00

Nature Medicine, Published online: 23 January 2026; doi:10.1038/s41591-025-04198-1

We propose straightforward principles to foster an evaluation-forward operating system that can transform the adoption of clinical artificial intelligence from a leap of faith into a stepwise, trust-building process.

The Responsibility Vacuum: Organizational Failure in Scaled Agent Systems

arXiv:2601.15059v1 Announce Type: new Abstract: Modern CI/CD pipelines integrating agent-generated code exhibit a structural failure in responsibility attribution. Decisions are executed through formally correct approval processes, yet no entity possesses both the authority to approve those decisions and the epistemic capacity to meaningfully understand their basis. We define this condition as responsibility vacuum: a state in which decisions occur, but responsibility cannot be attributed because authority and verification capacity do not coincide. We show that this is not a process deviation or technical defect, but a structural property of deployments where decision generation throughput exceeds bounded human verification capacity. We identify a scaling limit under standard deployment assumptions, including parallel agent generation, CI-based validation, and individualized human approval gates. Beyond a throughput threshold, verification ceases to function as a decision criterion and is replaced by ritualized approval based on proxy signals. Personalized responsibility becomes structurally unattainable in this regime. We further characterize a CI amplification dynamic, whereby increasing automated validation coverage raises proxy signal density without restoring human capacity. Under fixed time and attention constraints, this accelerates cognitive offloading in the broad sense and widens the gap between formal approval and epistemic understanding. Additional automation therefore amplifies, rather than mitigates, the responsibility vacuum. We conclude that unless organizations explicitly redesign decision boundaries or reassign responsibility away from individual decisions toward batch- or system-level ownership, responsibility vacuum remains an invisible but persistent failure mode in scaled agent deployments.

Towards Execution-Grounded Automated AI Research

arXiv:2601.14525v1 Announce Type: cross Abstract: Automated AI research holds great potential to accelerate scientific discovery. However, current LLMs often generate plausible-looking but ineffective ideas. Execution grounding may help, but it is unclear whether automated execution is feasible and whether LLMs can learn from the execution feedback. To investigate these, we first build an automated executor to implement ideas and launch large-scale parallel GPU experiments to verify their effectiveness. We then convert two realistic research problems - LLM pre-training and post-training - into execution environments and demonstrate that our automated executor can implement a large fraction of the ideas sampled from frontier LLMs. We analyze two methods to learn from the execution feedback: evolutionary search and reinforcement learning. Execution-guided evolutionary search is sample-efficient: it finds a method that significantly outperforms the GRPO baseline (69.4% vs 48.0%) on post-training, and finds a pre-training recipe that outperforms the nanoGPT baseline (19.7 minutes vs 35.9 minutes) on pre-training, all within just ten search epochs. Frontier LLMs often generate meaningful algorithmic ideas during search, but they tend to saturate early and only occasionally exhibit scaling trends. Reinforcement learning from execution reward, on the other hand, suffers from mode collapse. It successfully improves the average reward of the ideator model but not the upper-bound, due to models converging on simple ideas. We thoroughly analyze the executed ideas and training dynamics to facilitate future efforts towards execution-grounded automated AI research.

Automated Rubrics for Reliable Evaluation of Medical Dialogue Systems

arXiv:2601.15161v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used for clinical decision support, where hallucinations and unsafe suggestions may pose direct risks to patient safety. These risks are particularly challenging as they often manifest as subtle clinical errors that evade detection by generic metrics, while expert-authored fine-grained rubrics remain costly to construct and difficult to scale. In this paper, we propose a retrieval-augmented multi-agent framework designed to automate the generation of instance-specific evaluation rubrics. Our approach grounds evaluation in authoritative medical evidence by decomposing retrieved content into atomic facts and synthesizing them with user interaction constraints to form verifiable, fine-grained evaluation criteria. Evaluated on HealthBench, our framework achieves a Clinical Intent Alignment (CIA) score of 60.12%, a statistically significant improvement over the GPT-4o baseline (55.16%). In discriminative tests, our rubrics yield a mean score delta ($\mu_{\Delta} = 8.658$) and an AUROC of 0.977, nearly doubling the quality separation achieved by GPT-4o baseline (4.972). Beyond evaluation, our rubrics effectively guide response refinement, improving quality by 9.2% (from 59.0% to 68.2%). This provides a scalable and transparent foundation for both evaluating and improving medical LLMs. The code is available at https://anonymous.4open.science/r/Automated-Rubric-Generation-AF3C/.

Towards AI Transparency and Accountability: A Global Framework for Exchanging Information on AI Systems

arXiv:2307.13658v3 Announce Type: replace-cross Abstract: We propose that future AI transparency and accountability regulations are based on an open global standard for exchanging information about AI systems, which allows co-existence of potentially conflicting local regulations. Then, we discuss key components of a lightweight and effective AI transparency and/or accountability regulation. To prevent overregulation, the proposed approach encourages collaboration between regulators and industry to create a scalable and cost-efficient mutually beneficial solution. This includes using automated assessments and benchmarks with results transparently communicated through AI cards in an open AI register to facilitate meaningful public comparisons of competing AI systems. Such AI cards should report standardized measures tailored to the specific high-risk applications of AI systems and could be used for conformity assessments under AI transparency and accountability policies such as the European Union's AI Act.

PPGFlowECG: Latent Rectified Flow with Cross-Modal Encoding for PPG-Guided ECG Generation and Cardiovascular Disease Detection

arXiv:2509.19774v2 Announce Type: replace-cross Abstract: Electrocardiography (ECG) is the clinical gold standard for cardiovascular disease (CVD) assessment, yet continuous monitoring is constrained by the need for dedicated hardware and trained personnel. Photoplethysmography (PPG) is ubiquitous in wearable devices and readily scalable, but it lacks electrophysiological specificity, limiting diagnostic reliability. While generative methods aim to translate PPG into clinically useful ECG signals, existing approaches are limited by the misalignment of physiological semantics in generative models and the complexity of modeling in high-dimensional signals. To address these limitations, we propose PPGFlowECG, a two-stage framework that aligns PPG and ECG in a shared latent space using the CardioAlign Encoder and then synthesizes ECGs with latent rectified flow. We further provide a formal analysis of this coupling, showing that the CardioAlign Encoder is necessary to guarantee stable and semantically consistent ECG synthesis under our formulation. Extensive experiments on four datasets demonstrate improved synthesis fidelity and downstream diagnostic utility. These results indicate that PPGFlowECG supports scalable, wearable-first CVD screening when standard ECG acquisition is unavailable.

Large language models improve transferability of electronic health record-based predictions across countries and coding systems

npj Digital Medicine, Published online: 22 January 2026; doi:10.1038/s41746-026-02363-5

Large language models improve transferability of electronic health record-based predictions across countries and coding systems

Multimodal AI generates virtual population for tumor microenvironment modeling

GigaTIME leverages multimodal AI to generate virtual multiplex immunofluorescence (mIF) profiles from standard H&E slides, enabling comprehensive tumor immune microenvironment modeling across a large (>14,000) and diverse patient population. This virtual approach unlocks new opportunities for large-scale clinical discoveries that were previously hindered by the scarcity of mIF data.

Research progress in diagnosis and treatment of pancreatic cancer with mismatch repair and microsatellite instability

Clin Transl Oncol. 2026 Jan 21. doi: 10.1007/s12094-025-04214-3. Online ahead of print.

ABSTRACT

Pancreatic cancer (PC), predominantly pancreatic ductal adenocarcinoma, remains one of the most lethal malignancies, largely due to late diagnosis and intrinsic resistance to conventional therapies. In recent years, mismatch repair deficiency (dMMR) and microsatellite instability-high (MSI-H) have emerged as clinically actionable biomarkers in a small but distinct subset of PC, accounting for approximately 1-2% of cases. These tumors display unique molecular characteristics, including a high prevalence of wild-type KRAS and TP53, elevated tumor mutational burden, and recurrent kinase fusions, which together confer enhanced immunogenicity and increased sensitivity to immune checkpoint inhibitors (ICIs). In addition to their therapeutic relevance, dMMR/MSI-H status has important diagnostic implications for the identification of Lynch syndrome-associated pancreatic cancers, informing genetic counseling and familial risk assessment. This review summarizes current understanding of the molecular basis of mismatch repair deficiency and microsatellite instability in PC, evaluates available diagnostic approaches such as immunohistochemistry, polymerase chain reaction, and next-generation sequencing, and discusses the prognostic and predictive significance of dMMR/MSI-H status. Emerging clinical evidence supporting the use of ICIs in selected patients across neoadjuvant, adjuvant, and advanced disease settings is also reviewed, along with challenges related to assay discordance, tumor heterogeneity, and immunotherapy resistance. Finally, future directions are highlighted, emphasizing the need for standardized testing algorithms, integration of multi-omics and spatial profiling technologies, and prospective clinical studies to optimize precision treatment strategies for this rare but clinically meaningful subtype of pancreatic cancer.

PMID:41563663 | DOI:10.1007/s12094-025-04214-3

United multi-omics and machine learning refine regulatory T cell-defined hepatocellular carcinoma subtypes

iScience. 2025 Dec 3;29(1):114328. doi: 10.1016/j.isci.2025.114328. eCollection 2026 Jan 16.

ABSTRACT

Hepatocellular carcinoma (HCC) is highly heterogeneous and aggressive, and the absence of precision individual treatment regimen enables repeated immune escape. Exploiting regulatory T cell (Treg) marker genes as key classifiers, we used 10 clustering algorithms to integrate the multi-omics HCC patient data and combined them with 10 machine learning (ML) algorithms to delineate molecular subtypes predictive of prognosis and immune response. We identified two cancer subtypes (CSs) that are associated with prognosis, with the second subtype (CS2) showing the most favorable prognostic outcomes. Subsequently, 9 key genes were screened for HCC model scoring, stratifying patients into low-risk (good prognosis, responsive to immunotherapy) and high-risk (poor outcome, not responsive to immunotherapy) groups. The high-risk group may be effective against the mTOR inhibitor AZD8055. Comprehensive multi-omics data and multiple ML algorithms offer key insights into HCC occurrence and evolution, with model scores guiding patient prognosis and treatment clinically.

PMID:41561382 | PMC:PMC12814435 | DOI:10.1016/j.isci.2025.114328

Responsible AI for General-Purpose Systems: Overview, Challenges, and A Path Forward

arXiv:2601.13122v1 Announce Type: new Abstract: Modern general-purpose AI systems made using large language and vision models, are capable of performing a range of tasks like writing text articles, generating and debugging codes, querying databases, and translating from one language to another, which has made them quite popular across industries. However, there are risks like hallucinations, toxicity, and stereotypes in their output that make them untrustworthy. We review various risks and vulnerabilities of modern general-purpose AI along eight widely accepted responsible AI (RAI) principles (fairness, privacy, explainability, robustness, safety, truthfulness, governance, and sustainability) and compare how they are non-existent or less severe and easily mitigable in traditional task-specific counterparts. We argue that this is due to the non-deterministically high Degree of Freedom in output (DoFo) of general-purpose AI (unlike the deterministically constant or low DoFo of traditional task-specific AI systems), and there is a need to rethink our approach to RAI for general-purpose AI. Following this, we derive C2V2 (Control, Consistency, Value, Veracity) desiderata to meet the RAI requirements for future general-purpose AI systems, and discuss how recent efforts in AI alignment, retrieval-augmented generation, reasoning enhancements, etc. fare along one or more of the desiderata. We believe that the goal of developing responsible general-purpose AI can be achieved by formally modeling application- or domain-dependent RAI requirements along C2V2 dimensions, and taking a system design approach to suitably combine various techniques to meet the desiderata.

DeepEvidence: Empowering Biomedical Discovery with Deep Knowledge Graph Research

arXiv:2601.11560v1 Announce Type: cross Abstract: Biomedical knowledge graphs (KGs) encode vast, heterogeneous information spanning literature, genes, pathways, drugs, diseases, and clinical trials, but leveraging them collectively for scientific discovery remains difficult. Their structural differences, continual evolution, and limited cross-resource alignment require substantial manual integration, limiting the depth and scale of knowledge exploration. We introduce DeepEvidence, an AI-agent framework designed to perform Deep Research across various heterogeneous biomedical KGs. Unlike generic Deep Research systems that rely primarily on internet-scale text, DeepEvidence incorporates specialized knowledge-graph tooling and coordinated exploration strategies to systematically bridge heterogeneous resources. At its core is an orchestrator that directs two complementary agents: Breadth-First ReSearch (BFRS) for broad, multi-graph entity search, and Depth-First ReSearch (DFRS) for multi-hop, evidence-focused reasoning. An internal, incrementally built evidence graph provides a structured record of retrieved entities, relations, and supporting evidence. To operate at scale, DeepEvidence includes unified interfaces for querying diverse biomedical APIs and an execution sandbox that enables programmatic data retrieval, extraction, and analysis. Across established deep-reasoning benchmarks and four key stages of the biomedical discovery lifecycle: drug discovery, pre-clinical experimentation, clinical trial development, and evidence-based medicine, DeepEvidence demonstrates substantial gains in systematic exploration and evidence synthesis. These results highlight the potential of knowledge-graph-driven Deep Research to accelerate biomedical discovery.

Knowing When to Abstain: Medical LLMs Under Clinical Uncertainty

arXiv:2601.12471v1 Announce Type: cross Abstract: Current evaluation of large language models (LLMs) overwhelmingly prioritizes accuracy; however, in real-world and safety-critical applications, the ability to abstain when uncertain is equally vital for trustworthy deployment. We introduce MedAbstain, a unified benchmark and evaluation protocol for abstention in medical multiple-choice question answering (MCQA) -- a discrete-choice setting that generalizes to agentic action selection -- integrating conformal prediction, adversarial question perturbations, and explicit abstention options. Our systematic evaluation of both open- and closed-source LLMs reveals that even state-of-the-art, high-accuracy models often fail to abstain with uncertain. Notably, providing explicit abstention options consistently increases model uncertainty and safer abstention, far more than input perturbations, while scaling model size or advanced prompting brings little improvement. These findings highlight the central role of abstention mechanisms for trustworthy LLM deployment and offer practical guidance for improving safety in high-stakes applications.

A Cloud-based Multi-Agentic Workflow for Science

arXiv:2601.12607v1 Announce Type: cross Abstract: As Large Language Models (LLMs) become ubiquitous across various scientific domains, their lack of ability to perform complex tasks like running simulations or to make complex decisions limits their utility. LLM-based agents bridge this gap due to their ability to call external resources and tools and thus are now rapidly gaining popularity. However, coming up with a workflow that can balance the models, cloud providers, and external resources is very challenging, making implementing an agentic system more of a hindrance than a help. In this work, we present a domain-agnostic, model-independent workflow for an agentic framework that can act as a scientific assistant while being run entirely on cloud. Built with a supervisor agent marshaling an array of agents with individual capabilities, our framework brings together straightforward tasks like literature review and data analysis with more complex ones like simulation runs. We describe the framework here in full, including a proof-of-concept system we built to accelerate the study of Catalysts, which is highly important in the field of Chemistry and Material Science. We report the cost to operate and use this framework, including the breakdown of the cost by services use. We also evaluate our system on a custom-curated synthetic benchmark and a popular Chemistry benchmark, and also perform expert validation of the system. The results show that our system is able to route the task to the correct agent 90% of the time and successfully complete the assigned task 97.5% of the time for the synthetic tasks and 91% of the time for real-world tasks, while still achieving better or comparable accuracy to most frontier models, showing that this is a viable framework for other scientific domains to replicate.

SciHorizon-GENE: Benchmarking LLM for Life Sciences Inference from Gene Knowledge to Functional Understanding

arXiv:2601.12805v1 Announce Type: cross Abstract: Large language models (LLMs) have shown growing promise in biomedical research, particularly for knowledge-driven interpretation tasks. However, their ability to reliably reason from gene-level knowledge to functional understanding, However, their ability to reliably reason from gene-level knowledge to functional understanding, a core requirement for knowledge-enhanced cell atlas interpretation, remains largely underexplored. To address this gap, we introduce SciHorizon-GENE, a large-scale gene-centric benchmark constructed from authoritative biological databases. The benchmark integrates curated knowledge for over 190K human genes and comprises more than 540K questions covering diverse gene-to-function reasoning scenarios relevant to cell type annotation, functional interpretation, and mechanism-oriented analysis. Motivated by behavioral patterns observed in preliminary examinations, SciHorizon-GENE evaluates LLMs along four biologically critical perspectives: research attention sensitivity, hallucination tendency, answer completeness, and literature influence, explicitly targeting failure modes that limit the safe adoption of LLMs in biological interpretation pipelines. We systematically evaluate a wide range of state-of-the-art general-purpose and biomedical LLMs, revealing substantial heterogeneity in gene-level reasoning capabilities and persistent challenges in generating faithful, complete, and literature-grounded functional interpretations. Our benchmark establishes a systematic foundation for analyzing LLM behavior at the gene scale and offers insights for model selection and development, with direct relevance to knowledge-enhanced biological interpretation.

SciCoQA: Quality Assurance for Scientific Paper--Code Alignment

arXiv:2601.12910v1 Announce Type: cross Abstract: We present SciCoQA, a dataset for detecting discrepancies between scientific publications and their codebases to ensure faithful implementations. We construct SciCoQA from GitHub issues and reproducibility papers, and to scale our dataset, we propose a synthetic data generation method for constructing paper-code discrepancies. We analyze the paper-code discrepancies in detail and propose discrepancy types and categories to better understand the occurring mismatches. In total, our dataset consists of 611 paper-code discrepancies (81 real, 530 synthetic), spanning diverse computational science disciplines, including AI, Physics, Quantitative Biology, and others. Our evaluation of 21 LLMs highlights the difficulty of SciCoQA, particularly for instances involving omitted paper details, long-context inputs, and data outside the models' pre-training corpus. The best performing model in our evaluation, GPT-5, can only detect 45.7\% of real-world paper-code discrepancies.

AI-generated data contamination erodes pathological variability and diagnostic reliability

arXiv:2601.12946v1 Announce Type: cross Abstract: Generative artificial intelligence (AI) is rapidly populating medical records with synthetic content, creating a feedback loop where future models are increasingly at risk of training on uncurated AI-generated data. However, the clinical consequences of this AI-generated data contamination remain unexplored. Here, we show that in the absence of mandatory human verification, this self-referential cycle drives a rapid erosion of pathological variability and diagnostic reliability. By analysing more than 800,000 synthetic data points across clinical text generation, vision-language reporting, and medical image synthesis, we find that models progressively converge toward generic phenotypes regardless of the model architecture. Specifically, rare but critical findings, including pneumothorax and effusions, vanish from the synthetic content generated by AI models, while demographic representations skew heavily toward middle-aged male phenotypes. Crucially, this degradation is masked by false diagnostic confidence; models continue to issue reassuring reports while failing to detect life-threatening pathology, with false reassurance rates tripling to 40%. Blinded physician evaluation confirms that this decoupling of confidence and accuracy renders AI-generated documentation clinically useless after just two generations. We systematically evaluate three mitigation strategies, finding that while synthetic volume scaling fails to prevent collapse, mixing real data with quality-aware filtering effectively preserves diversity. Ultimately, our results suggest that without policy-mandated human oversight, the deployment of generative AI threatens to degrade the very healthcare data ecosystems it relies upon.
❌