Normal view
-
cs.AI, q-bio.NC updates on arXiv.org
-
High-Fidelity Longitudinal Patient Simulation Using Real-World Data
arXiv:2601.17310v1 Announce Type: new Abstract: Simulation is a powerful tool for exploring uncertainty. Its potential in clinical medicine is transformative and includes personalized treatment planning and virtual clinical trials. However, simulating patient trajectories is challenging because of complex biological and sociocultural influences. Here, we show that real-world clinical records can be leveraged to empirically model patient timelines. We developed a generative simulator model that
-
cs.AI, q-bio.NC updates on arXiv.org
-
Unheard in the Digital Age: Rethinking AI Bias and Speech Diversity
arXiv:2601.18641v1 Announce Type: cross Abstract: Speech remains one of the most visible yet overlooked vectors of inclusion and exclusion in contemporary society. While fluency is often equated with credibility and competence, individuals with atypical speech patterns are routinely marginalized. Given the current state of the debate, this article focuses on the structural biases that shape perceptions of atypical speech and are now being encoded into artificial intelligence. Automated speech r
Unheard in the Digital Age: Rethinking AI Bias and Speech Diversity
-
cs.AI, q-bio.NC updates on arXiv.org
-
MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications
arXiv:2409.07314v2 Announce Type: replace-cross Abstract: While Large Language Models (LLMs) achieve superhuman performance on standardized medical licensing exams, these static benchmarks have become saturated and increasingly disconnected from the functional requirements of clinical workflows. To bridge the gap between theoretical capability and verified utility, we introduce MEDIC, a comprehensive evaluation framework establishing leading indicators across various clinical dimensions. Beyond
MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications
-
cs.AI, q-bio.NC updates on arXiv.org
-
Pretrain Value, Not Reward: Decoupled Value Policy Optimization
arXiv:2502.16944v2 Announce Type: replace-cross Abstract: In this paper, we explore how directly pretraining a value model simplifies and stabilizes reinforcement learning from human feedback (RLHF). In reinforcement learning, value estimation is the key to policy optimization, distinct from reward supervision. The value function predicts the \emph{return-to-go} of a partial answer, that is, how promising the partial answer is if it were continued to completion. In RLHF, however, the standard p
Pretrain Value, Not Reward: Decoupled Value Policy Optimization
-
cs.AI, q-bio.NC updates on arXiv.org
-
uPVC-Net: A Universal Premature Ventricular Contraction Detection Deep Learning Algorithm
arXiv:2506.11238v2 Announce Type: replace-cross Abstract: Introduction: Premature Ventricular Contractions (PVCs) are common cardiac arrhythmias originating from the ventricles. Accurate detection remains challenging due to variability in electrocardiogram (ECG) waveforms caused by differences in lead placement, recording conditions, and population demographics. Methods: We developed uPVC-Net, a universal deep learning model to detect PVCs from any single-lead ECG recordings. The model is devel
uPVC-Net: A Universal Premature Ventricular Contraction Detection Deep Learning Algorithm
-
cs.AI, q-bio.NC updates on arXiv.org
-
On the Fundamental Limits of LLMs at Scale
arXiv:2511.12869v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have benefited enormously from scaling, yet these gains are bounded by five fundamental limitations: (1) hallucination, (2) context compression, (3) reasoning degradation, (4) retrieval fragility, and (5) multimodal misalignment. While existing surveys describe these phenomena empirically, they lack a rigorous theoretical synthesis connecting them to the foundational limits of computation, information, and le
On the Fundamental Limits of LLMs at Scale
-
cs.AI, q-bio.NC updates on arXiv.org
-
Empowering LLMs for Structure-Based Drug Design via Exploration-Augmented Latent Inference
arXiv:2601.15333v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) possess strong representation and reasoning capabilities, but their application to structure-based drug design (SBDD) is limited by insufficient understanding of protein structures and unpredictable molecular generation. To address these challenges, we propose Exploration-Augmented Latent Inference for LLMs (ELILLM), a framework that reinterprets the LLM generation process as an encoding, latent space explora
Empowering LLMs for Structure-Based Drug Design via Exploration-Augmented Latent Inference
-
npj Digital Medicine
-
Multimodal digital biopsy for preoperative prediction of occult peritoneal metastasis in gastric cancer
npj Digital Medicine, Published online: 26 January 2026; doi:10.1038/s41746-025-02268-9Multimodal digital biopsy for preoperative prediction of occult peritoneal metastasis in gastric cancer
Multimodal digital biopsy for preoperative prediction of occult peritoneal metastasis in gastric cancer
npj Digital Medicine, Published online: 26 January 2026; doi:10.1038/s41746-025-02268-9
Multimodal digital biopsy for preoperative prediction of occult peritoneal metastasis in gastric cancer-
Journal of Medical Internet Research
-
Feasibility, Acceptability, and Perspectives Regarding the Use of Activity Tracking Wearable Devices Among Home Health Aides: Mixed Methods Study
Background: Home health aides and attendants (HHAs) provide in-home care to the growing population of older adults who want to age in place. Despite their vital role in patient care, HHAs are an underserved and vulnerable population of health care professionals who often experience poor health themselves. Activity tracking devices offer a promising way to improve HHAs’ health-related awareness and promote health behavior change, particularly regarding physical activity and sleep quality, 2 areas
Feasibility, Acceptability, and Perspectives Regarding the Use of Activity Tracking Wearable Devices Among Home Health Aides: Mixed Methods Study
-
InfoQ

-
Article: Virtual Panel - AI in the Trenches: How Developers Are Rewriting the Software Process
This virtual panel brings together engineers, architects, and technical leaders to explore how AI is changing the landscape of software development. Practitioners share their insights on successes and failures when AI is incorporated into daily workflows, emphasizing the significance of context, validation, and cultural adaptation in making AI a sustainable element of modern engineering practices. By Arthur Casals, Mariia Bulycheva, May Walter, Phil Calçado, Andreas Kollegger
Article: Virtual Panel - AI in the Trenches: How Developers Are Rewriting the Software Process
This virtual panel brings together engineers, architects, and technical leaders to explore how AI is changing the landscape of software development. Practitioners share their insights on successes and failures when AI is incorporated into daily workflows, emphasizing the significance of context, validation, and cultural adaptation in making AI a sustainable element of modern engineering practices.
By Arthur Casals, Mariia Bulycheva, May Walter, Phil Calçado, Andreas Kollegger-
cs.AI, q-bio.NC updates on arXiv.org
-
PyHealth 2.0: A Comprehensive Open-Source Toolkit for Accessible and Reproducible Clinical Deep Learning
arXiv:2601.16414v1 Announce Type: cross Abstract: Difficulty replicating baselines, high computational costs, and required domain expertise create persistent barriers to clinical AI research. To address these challenges, we introduce PyHealth 2.0, an enhanced clinical deep learning toolkit that enables predictive modeling in as few as 7 lines of code. PyHealth 2.0 offers three key contributions: (1) a comprehensive toolkit addressing reproducibility and compatibility challenges by unifying 15+
PyHealth 2.0: A Comprehensive Open-Source Toolkit for Accessible and Reproducible Clinical Deep Learning
-
cs.AI, q-bio.NC updates on arXiv.org
-
DeepEra: A Deep Evidence Reranking Agent for Scientific Retrieval-Augmented Generated Question Answering
arXiv:2601.16478v1 Announce Type: cross Abstract: With the rapid growth of scientific literature, scientific question answering (SciQA) has become increasingly critical for exploring and utilizing scientific knowledge. Retrieval-Augmented Generation (RAG) enhances LLMs by incorporating knowledge from external sources, thereby providing credible evidence for scientific question answering. But existing retrieval and reranking methods remain vulnerable to passages that are semantically similar but
DeepEra: A Deep Evidence Reranking Agent for Scientific Retrieval-Augmented Generated Question Answering
-
cs.AI, q-bio.NC updates on arXiv.org
-
Evaluating Large Vision-language Models for Surgical Tool Detection
arXiv:2601.16895v1 Announce Type: cross Abstract: Surgery is a highly complex process, and artificial intelligence has emerged as a transformative force in supporting surgical guidance and decision-making. However, the unimodal nature of most current AI systems limits their ability to achieve a holistic understanding of surgical workflows. This highlights the need for general-purpose surgical AI systems capable of comprehensively modeling the interrelated components of surgical scenes. Recent a
Evaluating Large Vision-language Models for Surgical Tool Detection
-
cs.AI, q-bio.NC updates on arXiv.org
-
Advances in Artificial Intelligence: A Review for the Creative Industries
arXiv:2501.02725v5 Announce Type: replace Abstract: Artificial intelligence (AI) has undergone transformative advances since 2022, particularly through generative AI, large language models (LLMs), and diffusion models, fundamentally reshaping the creative industries. However, existing reviews have not comprehensively addressed these recent breakthroughs and their integrated impact across the creative production pipeline. This paper addresses this gap by providing a systematic review of AI techn
Advances in Artificial Intelligence: A Review for the Creative Industries
-
cs.AI, q-bio.NC updates on arXiv.org
-
EmbedAgent: Benchmarking Large Language Models in Embedded System Development
arXiv:2506.11003v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have shown promise in various tasks, yet few benchmarks assess their capabilities in embedded system development. In this paper, we introduce EmbedAgent, a paradigm designed to simulate real-world roles in embedded system development, such as Embedded System Programmer, Architect, and Integrator. This paradigm enables LLMs to be tested in tasks that bridge the gap between digital and physical systems, allowin
EmbedAgent: Benchmarking Large Language Models in Embedded System Development
-
cs.AI, q-bio.NC updates on arXiv.org
-
CRADLE Bench: A Clinician-Annotated Benchmark for Multi-Faceted Mental Health Crisis and Safety Risk Detection
arXiv:2510.23845v2 Announce Type: replace-cross Abstract: Detecting mental health crisis situations such as suicide ideation, rape, domestic violence, child abuse, and sexual harassment is a critical yet underexplored challenge for language models. When such situations arise during user--model interactions, models must reliably flag them, as failure to do so can have serious consequences. In this work, we introduce CRADLE BENCH, a benchmark for multi-faceted crisis detection. Unlike previous ef
CRADLE Bench: A Clinician-Annotated Benchmark for Multi-Faceted Mental Health Crisis and Safety Risk Detection
-
npj Digital Medicine
-
Deep learning for malignancy and tumor origin prediction using cytology or histopathology whole slide images
npj Digital Medicine, Published online: 24 January 2026; doi:10.1038/s41746-026-02359-1Deep learning for malignancy and tumor origin prediction using cytology or histopathology whole slide images
Deep learning for malignancy and tumor origin prediction using cytology or histopathology whole slide images
npj Digital Medicine, Published online: 24 January 2026; doi:10.1038/s41746-026-02359-1
Deep learning for malignancy and tumor origin prediction using cytology or histopathology whole slide images-
npj Digital Medicine
-
The diagnostic accuracy of wearable digital technology in detecting fertility window and menstrual cycles: a systematic review and Bayesian network meta-analysis
npj Digital Medicine, Published online: 24 January 2026; doi:10.1038/s41746-025-02320-8The diagnostic accuracy of wearable digital technology in detecting fertility window and menstrual cycles: a systematic review and Bayesian network meta-analysis
The diagnostic accuracy of wearable digital technology in detecting fertility window and menstrual cycles: a systematic review and Bayesian network meta-analysis
npj Digital Medicine, Published online: 24 January 2026; doi:10.1038/s41746-025-02320-8
The diagnostic accuracy of wearable digital technology in detecting fertility window and menstrual cycles: a systematic review and Bayesian network meta-analysis-
MIT Technology Review

-
The Download: chatbots for health, and US fights over AI regulation
This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. “Dr. Google” had its issues. Can ChatGPT Health do better? For the past two decades, there’s been a clear first step for anyone who starts experiencing new medical symptoms: Look them up online. The practice was so common that it gained the pejorative moniker “Dr. Google.” But times are changing, and many medical-information seekers are now using L
The Download: chatbots for health, and US fights over AI regulation
This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology.
“Dr. Google” had its issues. Can ChatGPT Health do better?
For the past two decades, there’s been a clear first step for anyone who starts experiencing new medical symptoms: Look them up online. The practice was so common that it gained the pejorative moniker “Dr. Google.” But times are changing, and many medical-information seekers are now using LLMs. According to OpenAI, 230 million people ask ChatGPT health-related queries each week.
That’s the context around the launch of OpenAI’s new ChatGPT Health product, which debuted earlier this month. The big question is: can the obvious risks of using AI for health-related queries be mitigated enough for them to be a net benefit? Read the full story.
—Grace Huckins
America’s coming war over AI regulation
In the final weeks of 2025, the battle over regulating artificial intelligence in the US reached boiling point. On December 11, after Congress failed twice to pass a law banning state AI laws, President Donald Trump signed a sweeping executive order seeking to handcuff states from regulating the booming industry.
Instead, he vowed to work with Congress to establish a “minimally burdensome” national AI policy. The move marked a victory for tech titans, who have been marshaling multimillion-dollar war chests to oppose AI regulations, arguing that a patchwork of state laws would stifle innovation.
In 2026, the battleground will shift to the courts. While some states might back down from passing AI laws, others will charge ahead. Read our story about what’s on the horizon.
—Michelle Kim
This story is from MIT Technology Review’s What’s Next series of stories that look across industries, trends, and technologies to give you a first look at the future. You can read the rest of them here.
Measles is surging in the US. Wastewater tracking could help.
This week marked a rather unpleasant anniversary: It’s a year since Texas reported a case of measles—the start of a significant outbreak that ended up spreading across multiple states. Since the start of January 2025, there have been over 2,500 confirmed cases of measles in the US. Three people have died.
As vaccination rates drop and outbreaks continue, scientists have been experimenting with new ways to quickly identify new cases and prevent the disease from spreading. And they are starting to see some success with wastewater surveillance. Read the full story.
—Jessica Hamzelou
This story is from The Checkup, our weekly newsletter giving you the inside track on all things health and biotech. Sign up to receive it in your inbox every Thursday.
The must-reads
I’ve combed the internet to find you today’s most fun/important/scary/fascinating stories about technology.
1 The US is dismantling itself
A foreign enemy could not invent a better chain of events to wreck its standing in the world. (Wired $)
+ We need to talk about whether Donald Trump might be losing it. (New Yorker $)
2 Big Tech is taking on more debt to fund its AI aspirations
And the bubble just keeps growing. (WP $)
+ Forget unicorns. 2026 is shaping up to be the year of the “hectocorn.” (The Guardian)
+ Everyone in tech agrees we’re in a bubble. They just can’t agree on what happens when it pops. (MIT Technology Review)
3 DOGE accessed even more personal data than we thought
Even now, the Trump administration still can’t say how much data is at risk, or what it was used for. (NPR)
4 TikTok has finalized a deal to create a new US entity
Ending years of uncertainty about its fate in America. (CNN)
+ Why China is the big winner out of all of this. (FT $)
5 The US is now officially out of the World Health Organization
And it’s leaving behind nearly $300 million in bills unpaid. (Ars Technica)
+ The US withdrawal from the WHO will hurt us all. (MIT Technology Review)
6 AI-powered disinformation swarms pose a threat to democracy
A would-be autocrat could use them to persuade populations to accept cancelled elections or overturn results. (The Guardian)
+ The era of AI persuasion in elections is about to begin. (MIT Technology Review)
7 We’re about to start seeing more robots everywhere
But exactly what they’ll look like remains up for debate. (Vox $)
+ Chinese companies are starting to dominate entire sectors of AI and robotics. (MIT Technology Review)
8 Some people seem to be especially vulnerable to loneliness
If you’re ‘other-directed’, you could particularly benefit from less screentime. (New Scientist $)
9 This academic lost two years of work with a single click
TL;DR: Don’t rely on ChatGPT to store your data. (Nature)
10 How animals develop a sense of direction ![]()
![]()
Their ‘internal compass’ seems to be informed by landmarks that help them form a mental map. (Quanta $)
Quote of the day
“The rate at which AI is progressing, I think we have AI that is smarter than any human this year, and no later than next year.”
—Elon Musk simply cannot resist the urge to make wild predictions at Davos, Wired reports.
One more thing

Africa fights rising hunger by looking to foods of the past
After falling steadily for decades, the prevalence of global hunger is now on the rise—nowhere more so than in sub-Saharan Africa.
Africa’s indigenous crops are often more nutritious and better suited to the hot and dry conditions that are becoming more prevalent, yet many have been neglected by science, which means they tend to be more vulnerable to diseases and pests and yield well below their theoretical potential.
Now the question is whether researchers, governments, and farmers can work together in a way that gets these crops onto plates and provides Africans from all walks of life with the energy and nutrition that they need to thrive, whatever climate change throws their way. Read the full story.
—Jonathan W. Rosen
We can still have nice things
A place for comfort, fun and distraction to brighten up your day. (Got any ideas? Drop me a line or skeet ’em at me.)
+ The only thing I fancy dry this January is a martini. Here’s how to make one.
+ If you absolutely adore the Bic crystal pen, you might want this lamp.
+ Cozy up with a nice long book this winter. ($)
+ Want to eat healthier? Slow down and tune out food ‘noise’. ($)
-
Nature Medicine
-
Principles to guide clinical AI readiness and move from benchmarks to real-world evaluation
Nature Medicine, Published online: 23 January 2026; doi:10.1038/s41591-025-04198-1We propose straightforward principles to foster an evaluation-forward operating system that can transform the adoption of clinical artificial intelligence from a leap of faith into a stepwise, trust-building process.
Principles to guide clinical AI readiness and move from benchmarks to real-world evaluation
Nature Medicine, Published online: 23 January 2026; doi:10.1038/s41591-025-04198-1
We propose straightforward principles to foster an evaluation-forward operating system that can transform the adoption of clinical artificial intelligence from a leap of faith into a stepwise, trust-building process.