❌

Normal view

  • ✇Latest Science News -- ScienceDaily
  • A simple blood test mismatch linked to kidney failure and death
    A major global study suggests that a hidden mismatch between two common blood tests could quietly signal serious trouble ahead. When results from creatinine and cystatin C—two markers used to assess kidney health—don’t line up, the risk of kidney failure, heart disease, and even death appears to rise sharply. Researchers found that this gap is especially common among hospitalized and older patients, and that relying on just one test may miss early warning signs.
     

A simple blood test mismatch linked to kidney failure and death

22 January 2026 at 01:19
A major global study suggests that a hidden mismatch between two common blood tests could quietly signal serious trouble ahead. When results from creatinine and cystatin C—two markers used to assess kidney health—don’t line up, the risk of kidney failure, heart disease, and even death appears to rise sharply. Researchers found that this gap is especially common among hospitalized and older patients, and that relying on just one test may miss early warning signs.
  • ✇STAT
  • STAT+: OpenEvidence raises $250 million, doubling its valuation Mario Aguilar
    OpenEvidence, maker of a popular chatbot that helps doctors search clinical evidence, on Wednesday announced $250 million in new funding. The new round led by Thrive Capital and DST Global values OpenEvidence at $12 billion, and the company has announced $735 million in funding in the last 12 months. OpenEvidence is free to use by any clinician with a national provider identifier number. The company’s primary business model is advertising shown to clinicians.  Founded in 2022, OpenEvidence
     

STAT+: OpenEvidence raises $250 million, doubling its valuation

21 January 2026 at 23:38

OpenEvidence, maker of a popular chatbot that helps doctors search clinical evidence, on Wednesday announced $250 million in new funding.

The new round led by Thrive Capital and DST Global values OpenEvidence at $12 billion, and the company has announced $735 million in funding in the last 12 months. OpenEvidence is free to use by any clinician with a national provider identifier number. The company’s primary business model is advertising shown to clinicians. 

Founded in 2022, OpenEvidence is one of the most prominent and best-funded companies from a wave of health artificial intelligence companies that emerged since the widespread availability of large language models. Reflecting on the eye-popping fundraising for health AI, a Silicon Valley Bank report released earlier in January raised an eyebrow at the ability of companies like OpenEvidence to deliver on their stratospheric valuations with advertising and software-as-a-service business models. “It won’t be a surprise to see them tap the value of the data they’re already collecting” to offer more services to pharma or other customers, the authors wrote.

Continue to STAT+ to read the full story…

© Adobe

Research progress in diagnosis and treatment of pancreatic cancer with mismatch repair and microsatellite instability

Clin Transl Oncol. 2026 Jan 21. doi: 10.1007/s12094-025-04214-3. Online ahead of print.

ABSTRACT

Pancreatic cancer (PC), predominantly pancreatic ductal adenocarcinoma, remains one of the most lethal malignancies, largely due to late diagnosis and intrinsic resistance to conventional therapies. In recent years, mismatch repair deficiency (dMMR) and microsatellite instability-high (MSI-H) have emerged as clinically actionable biomarkers in a small but distinct subset of PC, accounting for approximately 1-2% of cases. These tumors display unique molecular characteristics, including a high prevalence of wild-type KRAS and TP53, elevated tumor mutational burden, and recurrent kinase fusions, which together confer enhanced immunogenicity and increased sensitivity to immune checkpoint inhibitors (ICIs). In addition to their therapeutic relevance, dMMR/MSI-H status has important diagnostic implications for the identification of Lynch syndrome-associated pancreatic cancers, informing genetic counseling and familial risk assessment. This review summarizes current understanding of the molecular basis of mismatch repair deficiency and microsatellite instability in PC, evaluates available diagnostic approaches such as immunohistochemistry, polymerase chain reaction, and next-generation sequencing, and discusses the prognostic and predictive significance of dMMR/MSI-H status. Emerging clinical evidence supporting the use of ICIs in selected patients across neoadjuvant, adjuvant, and advanced disease settings is also reviewed, along with challenges related to assay discordance, tumor heterogeneity, and immunotherapy resistance. Finally, future directions are highlighted, emphasizing the need for standardized testing algorithms, integration of multi-omics and spatial profiling technologies, and prospective clinical studies to optimize precision treatment strategies for this rare but clinically meaningful subtype of pancreatic cancer.

PMID:41563663 | DOI:10.1007/s12094-025-04214-3

  • ✇MIT Technology Review
  • Everyone wants AI sovereignty. No one can truly have it. Cathy Li
    Governments plan to pour $1.3 trillion into AI infrastructure by 2030 to invest in “sovereign AI,” with the premise being that countries should be in control of their own AI capabilities. The funds include financing for domestic data centers, locally trained models, independent supply chains, and national talent pipelines. This is a response to real shocks: covid-era supply chain breakdowns, rising geopolitical tensions, and the war in Ukraine.   But the pursuit of absolute autonomy is runnin
     

Everyone wants AI sovereignty. No one can truly have it.

21 January 2026 at 22:00

Governments plan to pour $1.3 trillion into AI infrastructure by 2030 to invest in “sovereign AI,” with the premise being that countries should be in control of their own AI capabilities. The funds include financing for domestic data centers, locally trained models, independent supply chains, and national talent pipelines. This is a response to real shocks: covid-era supply chain breakdowns, rising geopolitical tensions, and the war in Ukraine.  

But the pursuit of absolute autonomy is running into reality. AI supply chains are irreducibly global: Chips are designed in the US and manufactured in East Asia; models are trained on data sets drawn from multiple countries; applications are deployed across dozens of jurisdictions.  

If sovereignty is to remain meaningful, it must shift from a defensive model of self-reliance to a vision that emphasizes the concept of orchestration, balancing national autonomy with strategic partnership. 

Why infrastructure-first strategies hit walls 

A November survey by Accenture found that 62% of European organizations are now seeking sovereign AI solutions, driven primarily by geopolitical anxiety rather than technical necessity. That figure rises to 80% in Denmark and 72% in Germany. The European Union has appointed its first Commissioner for Tech Sovereignty. 

This year, $475 billion is flowing into AI data centers globally. In the United States, AI data centers accounted for roughly one-fifth of GDP growth in the second quarter of 2025. But the obstacle for other nations hoping to follow suit isn’t just money. It’s energy and physics. Global data center capacity is projected to hit 130 gigawatts by 2030, and for every $1 billion spent on these facilities, $125 million is needed for electricity networks. More than $750 billion in planned investment is already facing grid delays. 

And it’s also talent. Researchers and entrepreneurs are mobile, drawn to ecosystems with access to capital, competitive wages, and rapid innovation cycles. Infrastructure alone won’t attract or retain world-class talent.  

What works: An orchestrated sovereignty

What nations need isn’t sovereignty through isolation but through specialization and orchestration. This means choosing which capabilities you build, which you pursue through partnership, and where you can genuinely lead in shaping the global AI landscape. 

The most successful AI strategies don’t try to replicate Silicon Valley; they identify specific advantages and build partnerships around them. 

Singapore offers a model. Rather than seeking to duplicate massive infrastructure, it invested in governance frameworks, digital-identity platforms, and applications of AI in logistics and finance, areas where it can realistically compete. 

Israel shows a different path. Its strength lies in a dense network of startups and military-adjacent research institutions delivering outsize influence despite the country’s small size. 

South Korea is instructive too. While it has national champions like Samsung and Naver, these firms still partner with Microsoft and Nvidia on infrastructure. That’s deliberate collaboration reflecting strategic oversight, not dependence.  

Even China, despite its scale and ambition, cannot secure full-stack autonomy. Its reliance on global research networks and on foreign lithography equipment, such as extreme ultraviolet systems needed to manufacture advanced chips and GPU architectures, shows the limits of techno-nationalism. 

The pattern is clear: Nations that specialize and partner strategically can outperform those trying to do everything alone. 

Three ways to align ambition with reality 

1.  Measure added value, not inputs.  

Sovereignty isn’t how many petaflops you own. It’s how many lives you improve and how fast the economy grows. Real sovereignty is the ability to innovate in support of national priorities such as productivity, resilience, and sustainability while maintaining freedom to shape governance and standards.  

Nations should track the use of AI in health care and monitor how the technology’s adoption correlates with manufacturing productivity, patent citations, and international research collaborations. The goal is to ensure that AI ecosystems generate inclusive and lasting economic and social value.  

2. Cultivate a strong AI innovation ecosystem. 

Build infrastructure, but also build the ecosystem around it: research institutions, technical education, entrepreneurship support, and public-private talent development. Infrastructure without skilled talent and vibrant networks cannot deliver a lasting competitive advantage.   

3. Build global partnerships.  

Strategic partnerships enable nations to pool resources, lower infrastructure costs, and access complementary expertise. Singapore’s work with global cloud providers and the EU’s collaborative research programs show how nations advance capabilities faster through partnership than through isolation. Rather than competing to set dominant standards, nations should collaborate on interoperable frameworks for transparency, safety, and accountability.  

What’s at stake 

Overinvesting in independence fragments markets and slows cross-border innovation, which is the foundation of AI progress. When strategies focus too narrowly on control, they sacrifice the agility needed to compete. 

The cost of getting this wrong isn’t just wasted capital—it’s a decade of falling behind. Nations that double down on infrastructure-first strategies risk ending up with expensive data centers running yesterday’s models, while competitors that choose strategic partnerships iterate faster, attract better talent, and shape the standards that matter. 

The winners will be those who define sovereignty not as separation, but as participation plus leadership—choosing who they depend on, where they build, and which global rules they shape. Strategic interdependence may feel less satisfying than independence, but it’s real, it is achievable, and it will separate the leaders from the followers over the next decade. 

The age of intelligent systems demands intelligent strategies—ones that measure success not by infrastructure owned, but by problems solved. Nations that embrace this shift won’t just participate in the AI economy; they’ll shape it. That’s sovereignty worth pursuing. 

Cathy Li is head of the Centre for AI Excellence at the World Economic Forum.

Communication Strategies to Promote Patient Engagement in Telemedicine: Systematic Review

Background: The rapid growth of telemedicine offers convenience, flexibility, and accessibility for patients to have health care services worldwide. To succeed in telemedicine, health care practitioners and telemedicine tools must engage patients through effective communication. However, a research gap exists in understanding the communication strategies used in telemedicine and how they effectively engage patients. Objective: This study aims to identify communication strategies influencing patient engagement in telemedicine with provider-patient interactions, as well as how included studies evaluate patient engagement through a systematic review. Methods: We searched the literature comprehensively using 6 databases, Web of Science, PubMed, Scopus, MEDLINE, CINAHL, and Embase, from inception to October 2025. We included empirical, English-language studies that examined communication strategies affecting patient engagement in telemedicine with provider-patient interactions. Studies lacking actual patients or provider-patient interactions in telemedicine were excluded. We used content analysis to identify texts that were related to Theme 1: the communication strategies affecting patient engagement, and Theme 2: evaluation of patient engagement. Coded texts were analyzed to develop subthemes and themes of identified communication strategies. Methods for evaluating patient engagement were summarized. A narrative synthesis was conducted because of heterogeneity across study design and outcomes. We used the Mixed Methods Appraisal Tool to assess the quality of research included in this study. Results: This study systematically reviewed 34 peer-reviewed articles, revealing 3 overarching themes of effective communication strategies that enhance patient engagement: interpersonal communication strategies, with 6 subthemes (building relationships, supportive attitude, interactive dialogic loop, nonverbal communication, professionalism and accuracy, and tailored communication); team-level communication strategies, with 3 subthemes (training and preparation, teamwork and care coordination, and cultural and linguistic sensitivity); and system-level communication strategies, with 3 subthemes (usefulness of information, ease of use, and data privacy and security). We also found that included studies predominantly used qualitative research methods, such as semistructured interviews and focus groups, to collect patient engagement data. Conclusions: This review provides an innovative synthesis of communication strategies that promote patient engagement in telemedicine by integrating interpersonal (micro), team (meso), and system-level (macro) perspectives. Unlike previous reviews that focused on single aspects or levels of communication, this study offers a holistic framework that advances theoretical understanding of how multilevel communication strategies collectively shape patient engagement. Practically, the findings offer actionable guidance for health care professionals, telemedicine developers, and policymakers seeking to enhance the quality and sustainability of telemedicine services. In real-world settings, the identified strategies can inform professional training, platform design, and policy development to support patient-centered digital care. This review is the first to systematically bring together communication strategies for patient engagement in telemedicine across all 3 levels. Future research should build on this framework by developing and validating quantitative measures of patient engagement and examining the relationships between communication strategies and telemedicine outcomes.

United multi-omics and machine learning refine regulatory T cell-defined hepatocellular carcinoma subtypes

iScience. 2025 Dec 3;29(1):114328. doi: 10.1016/j.isci.2025.114328. eCollection 2026 Jan 16.

ABSTRACT

Hepatocellular carcinoma (HCC) is highly heterogeneous and aggressive, and the absence of precision individual treatment regimen enables repeated immune escape. Exploiting regulatory T cell (Treg) marker genes as key classifiers, we used 10 clustering algorithms to integrate the multi-omics HCC patient data and combined them with 10 machine learning (ML) algorithms to delineate molecular subtypes predictive of prognosis and immune response. We identified two cancer subtypes (CSs) that are associated with prognosis, with the second subtype (CS2) showing the most favorable prognostic outcomes. Subsequently, 9 key genes were screened for HCC model scoring, stratifying patients into low-risk (good prognosis, responsive to immunotherapy) and high-risk (poor outcome, not responsive to immunotherapy) groups. The high-risk group may be effective against the mTOR inhibitor AZD8055. Comprehensive multi-omics data and multiple ML algorithms offer key insights into HCC occurrence and evolution, with model scores guiding patient prognosis and treatment clinically.

PMID:41561382 | PMC:PMC12814435 | DOI:10.1016/j.isci.2025.114328

Responsible AI for General-Purpose Systems: Overview, Challenges, and A Path Forward

arXiv:2601.13122v1 Announce Type: new Abstract: Modern general-purpose AI systems made using large language and vision models, are capable of performing a range of tasks like writing text articles, generating and debugging codes, querying databases, and translating from one language to another, which has made them quite popular across industries. However, there are risks like hallucinations, toxicity, and stereotypes in their output that make them untrustworthy. We review various risks and vulnerabilities of modern general-purpose AI along eight widely accepted responsible AI (RAI) principles (fairness, privacy, explainability, robustness, safety, truthfulness, governance, and sustainability) and compare how they are non-existent or less severe and easily mitigable in traditional task-specific counterparts. We argue that this is due to the non-deterministically high Degree of Freedom in output (DoFo) of general-purpose AI (unlike the deterministically constant or low DoFo of traditional task-specific AI systems), and there is a need to rethink our approach to RAI for general-purpose AI. Following this, we derive C2V2 (Control, Consistency, Value, Veracity) desiderata to meet the RAI requirements for future general-purpose AI systems, and discuss how recent efforts in AI alignment, retrieval-augmented generation, reasoning enhancements, etc. fare along one or more of the desiderata. We believe that the goal of developing responsible general-purpose AI can be achieved by formally modeling application- or domain-dependent RAI requirements along C2V2 dimensions, and taking a system design approach to suitably combine various techniques to meet the desiderata.

Virtual Urbanism: An AI-Driven Framework for Quantifying Urban Identity. A Tokyo-Based Pilot Study Using Diffusion-Generated Synthetic Environments

arXiv:2601.13846v1 Announce Type: new Abstract: This paper introduces Virtual Urbanism (VU), a multimodal AI-driven analytical framework for quantifying urban identity through the medium of synthetic urban replicas. The framework aims to advance computationally tractable urban identity metrics. To demonstrate feasibility, the pilot study Virtual Urbanism and Tokyo Microcosms is presented. A pipeline integrating Stable Diffusion and LoRA models was used to produce synthetic replicas of nine Tokyo areas rendered as dynamic synthetic urban sequences, excluding existing orientation markers to elicit core identity-forming elements. Human-evaluation experiments (I) assessed perceptual legitimacy of replicas; (II) quantified area-level identity; (III) derived core identity-forming elements. Results showed a mean identification accuracy of ~81%, confirming the validity of the replicas. Urban Identity Level (UIL) metric enabled assessment of identity levels across areas, while semantic analysis revealed culturally embedded typologies as core identity-forming elements, positioning VU as a viable framework for AI-augmented urban analysis, outlining a path toward automated, multi-parameter identity metrics.

Medication counseling with large language models: balancing flexibility and rigidity

arXiv:2601.11544v1 Announce Type: cross Abstract: The introduction of large language models (LLMs) has greatly enhanced the capabilities of software agents. Instead of relying on rule-based interactions, agents can now interact in flexible ways akin to humans. However, this flexibility quickly becomes a problem in fields where errors can be disastrous, such as in a pharmacy context, but the opposite also holds true; a system that is too inflexible will also lead to errors, as it can become too rigid to handle situations that are not accounted for. Work using LLMs in a pharmacy context have adopted a wide scope, accounting for many different medications in brief interactions -- our strategy is the opposite: focus on a more narrow and long task. This not only enables a greater understanding of the task at hand, but also provides insight into what challenges are present in an interaction of longer nature. The main challenge, however, remains the same for a narrow and wide system: it needs to strike a balance between adherence to conversational requirements and flexibility. In an effort to strike such a balance, we present a prototype system meant to provide medication counseling while juggling these two extremes. We also cover our design in constructing such a system, with a focus on methods aiming to fulfill conversation requirements, reduce hallucinations and promote high-quality responses. The methods used have the potential to increase the determinism of the system, while simultaneously not removing the dynamic conversational abilities granted by the usage of LLMs. However, a great deal of work remains ahead, and the development of this kind of system needs to involve continuous testing and a human-in-the-loop. It should also be evaluated outside of commonly used benchmarks for LLMs, as these do not adequately capture the complexities of this kind of conversational system.

DeepEvidence: Empowering Biomedical Discovery with Deep Knowledge Graph Research

arXiv:2601.11560v1 Announce Type: cross Abstract: Biomedical knowledge graphs (KGs) encode vast, heterogeneous information spanning literature, genes, pathways, drugs, diseases, and clinical trials, but leveraging them collectively for scientific discovery remains difficult. Their structural differences, continual evolution, and limited cross-resource alignment require substantial manual integration, limiting the depth and scale of knowledge exploration. We introduce DeepEvidence, an AI-agent framework designed to perform Deep Research across various heterogeneous biomedical KGs. Unlike generic Deep Research systems that rely primarily on internet-scale text, DeepEvidence incorporates specialized knowledge-graph tooling and coordinated exploration strategies to systematically bridge heterogeneous resources. At its core is an orchestrator that directs two complementary agents: Breadth-First ReSearch (BFRS) for broad, multi-graph entity search, and Depth-First ReSearch (DFRS) for multi-hop, evidence-focused reasoning. An internal, incrementally built evidence graph provides a structured record of retrieved entities, relations, and supporting evidence. To operate at scale, DeepEvidence includes unified interfaces for querying diverse biomedical APIs and an execution sandbox that enables programmatic data retrieval, extraction, and analysis. Across established deep-reasoning benchmarks and four key stages of the biomedical discovery lifecycle: drug discovery, pre-clinical experimentation, clinical trial development, and evidence-based medicine, DeepEvidence demonstrates substantial gains in systematic exploration and evidence synthesis. These results highlight the potential of knowledge-graph-driven Deep Research to accelerate biomedical discovery.

Measuring Stability Beyond Accuracy in Small Open-Source Medical Large Language Models for Pediatric Endocrinology

arXiv:2601.11567v1 Announce Type: cross Abstract: Small open-source medical large language models (LLMs) offer promising opportunities for low-resource deployment and broader accessibility. However, their evaluation is often limited to accuracy on medical multiple choice question (MCQ) benchmarks, and lacks evaluation of consistency, robustness, or reasoning behavior. We use MCQ coupled to human evaluation and clinical review to assess six small open-source medical LLMs (HuatuoGPT-o1 (Chen 2024), Diabetica-7B, Diabetica-o1 (Wei 2024), Meditron3-8B (Sallinen2025), MedFound-7B (Liu 2025), and ClinicaGPT-base-zh (Wang 2023)) in pediatric endocrinology. In deterministic settings, we examine the effect of prompt variation on models' output and self-assessment bias. In stochastic settings, we evaluate output variability and investigate the relationship between consistency and correctness. HuatuoGPT-o1-8B achieved the highest performance. The results show that high consistency across the model response is not an indicator of correctness, although HuatuoGPT-o1-8B showed the highest consistency rate. When tasked with selecting correct reasoning, both HuatuoGPT-o1-8B and Diabetica-o1 exhibit self-assessment bias and dependency on the order of the candidate explanations. Expert review of incorrect reasoning rationales identified a mix of clinically acceptable responses and clinical oversight. We further show that system-level perturbations, such as differences in CUDA builds, can yield statistically significant shifts in model output despite stable accuracy. This work demonstrates that small, semantically negligible prompt perturbations lead to divergent outputs, raising concerns about reproducibility of LLM-based evaluations and highlights the output variability under different stochastic regimes, emphasizing the need of a broader diagnostic framework to understand potential pitfalls in real-world clinical decision support scenarios.

Knowing When to Abstain: Medical LLMs Under Clinical Uncertainty

arXiv:2601.12471v1 Announce Type: cross Abstract: Current evaluation of large language models (LLMs) overwhelmingly prioritizes accuracy; however, in real-world and safety-critical applications, the ability to abstain when uncertain is equally vital for trustworthy deployment. We introduce MedAbstain, a unified benchmark and evaluation protocol for abstention in medical multiple-choice question answering (MCQA) -- a discrete-choice setting that generalizes to agentic action selection -- integrating conformal prediction, adversarial question perturbations, and explicit abstention options. Our systematic evaluation of both open- and closed-source LLMs reveals that even state-of-the-art, high-accuracy models often fail to abstain with uncertain. Notably, providing explicit abstention options consistently increases model uncertainty and safer abstention, far more than input perturbations, while scaling model size or advanced prompting brings little improvement. These findings highlight the central role of abstention mechanisms for trustworthy LLM deployment and offer practical guidance for improving safety in high-stakes applications.

A Cloud-based Multi-Agentic Workflow for Science

arXiv:2601.12607v1 Announce Type: cross Abstract: As Large Language Models (LLMs) become ubiquitous across various scientific domains, their lack of ability to perform complex tasks like running simulations or to make complex decisions limits their utility. LLM-based agents bridge this gap due to their ability to call external resources and tools and thus are now rapidly gaining popularity. However, coming up with a workflow that can balance the models, cloud providers, and external resources is very challenging, making implementing an agentic system more of a hindrance than a help. In this work, we present a domain-agnostic, model-independent workflow for an agentic framework that can act as a scientific assistant while being run entirely on cloud. Built with a supervisor agent marshaling an array of agents with individual capabilities, our framework brings together straightforward tasks like literature review and data analysis with more complex ones like simulation runs. We describe the framework here in full, including a proof-of-concept system we built to accelerate the study of Catalysts, which is highly important in the field of Chemistry and Material Science. We report the cost to operate and use this framework, including the breakdown of the cost by services use. We also evaluate our system on a custom-curated synthetic benchmark and a popular Chemistry benchmark, and also perform expert validation of the system. The results show that our system is able to route the task to the correct agent 90% of the time and successfully complete the assigned task 97.5% of the time for the synthetic tasks and 91% of the time for real-world tasks, while still achieving better or comparable accuracy to most frontier models, showing that this is a viable framework for other scientific domains to replicate.

SciHorizon-GENE: Benchmarking LLM for Life Sciences Inference from Gene Knowledge to Functional Understanding

arXiv:2601.12805v1 Announce Type: cross Abstract: Large language models (LLMs) have shown growing promise in biomedical research, particularly for knowledge-driven interpretation tasks. However, their ability to reliably reason from gene-level knowledge to functional understanding, However, their ability to reliably reason from gene-level knowledge to functional understanding, a core requirement for knowledge-enhanced cell atlas interpretation, remains largely underexplored. To address this gap, we introduce SciHorizon-GENE, a large-scale gene-centric benchmark constructed from authoritative biological databases. The benchmark integrates curated knowledge for over 190K human genes and comprises more than 540K questions covering diverse gene-to-function reasoning scenarios relevant to cell type annotation, functional interpretation, and mechanism-oriented analysis. Motivated by behavioral patterns observed in preliminary examinations, SciHorizon-GENE evaluates LLMs along four biologically critical perspectives: research attention sensitivity, hallucination tendency, answer completeness, and literature influence, explicitly targeting failure modes that limit the safe adoption of LLMs in biological interpretation pipelines. We systematically evaluate a wide range of state-of-the-art general-purpose and biomedical LLMs, revealing substantial heterogeneity in gene-level reasoning capabilities and persistent challenges in generating faithful, complete, and literature-grounded functional interpretations. Our benchmark establishes a systematic foundation for analyzing LLM behavior at the gene scale and offers insights for model selection and development, with direct relevance to knowledge-enhanced biological interpretation.

SciCoQA: Quality Assurance for Scientific Paper--Code Alignment

arXiv:2601.12910v1 Announce Type: cross Abstract: We present SciCoQA, a dataset for detecting discrepancies between scientific publications and their codebases to ensure faithful implementations. We construct SciCoQA from GitHub issues and reproducibility papers, and to scale our dataset, we propose a synthetic data generation method for constructing paper-code discrepancies. We analyze the paper-code discrepancies in detail and propose discrepancy types and categories to better understand the occurring mismatches. In total, our dataset consists of 611 paper-code discrepancies (81 real, 530 synthetic), spanning diverse computational science disciplines, including AI, Physics, Quantitative Biology, and others. Our evaluation of 21 LLMs highlights the difficulty of SciCoQA, particularly for instances involving omitted paper details, long-context inputs, and data outside the models' pre-training corpus. The best performing model in our evaluation, GPT-5, can only detect 45.7\% of real-world paper-code discrepancies.

AI-generated data contamination erodes pathological variability and diagnostic reliability

arXiv:2601.12946v1 Announce Type: cross Abstract: Generative artificial intelligence (AI) is rapidly populating medical records with synthetic content, creating a feedback loop where future models are increasingly at risk of training on uncurated AI-generated data. However, the clinical consequences of this AI-generated data contamination remain unexplored. Here, we show that in the absence of mandatory human verification, this self-referential cycle drives a rapid erosion of pathological variability and diagnostic reliability. By analysing more than 800,000 synthetic data points across clinical text generation, vision-language reporting, and medical image synthesis, we find that models progressively converge toward generic phenotypes regardless of the model architecture. Specifically, rare but critical findings, including pneumothorax and effusions, vanish from the synthetic content generated by AI models, while demographic representations skew heavily toward middle-aged male phenotypes. Crucially, this degradation is masked by false diagnostic confidence; models continue to issue reassuring reports while failing to detect life-threatening pathology, with false reassurance rates tripling to 40%. Blinded physician evaluation confirms that this decoupling of confidence and accuracy renders AI-generated documentation clinically useless after just two generations. We systematically evaluate three mitigation strategies, finding that while synthetic volume scaling fails to prevent collapse, mixing real data with quality-aware filtering effectively preserves diversity. Ultimately, our results suggest that without policy-mandated human oversight, the deployment of generative AI threatens to degrade the very healthcare data ecosystems it relies upon.

Multi-objective fluorescent molecule design with a data-physics dual-driven generative framework

arXiv:2601.13564v1 Announce Type: cross Abstract: Designing fluorescent small molecules with tailored optical and physicochemical properties requires navigating vast, underexplored chemical space while satisfying multiple objectives and constraints. Conventional generate-score-screen approaches become impractical under such realistic design specifications, owing to their low search efficiency, unreliable generalizability of machine-learning prediction, and the prohibitive cost of quantum chemical calculation. Here we present LUMOS, a data-and-physics driven framework for inverse design of fluorescent molecules. LUMOS couples generator and predictor within a shared latent representation, enabling direct specification-to-molecule design and efficient exploration. Moreover, LUMOS combines neural networks with a fast time-dependent density functional theory (TD-DFT) calculation workflow to build a suite of complementary predictors spanning different trade-offs in speed, accuracy, and generalizability, enabling reliable property prediction across diverse scenarios. Finally, LUMOS employs a property-guided diffusion model integrated with multi-objective evolutionary algorithms, enabling de novo design and molecular optimization under multiple objectives and constraints. Across comprehensive benchmarks, LUMOS consistently outperforms baseline models in terms of accuracy, generalizability and physical plausibility for fluorescence property prediction, and demonstrates superior performance in multi-objective scaffold- and fragment-level molecular optimization. Further validation using TD-DFT and molecular dynamics (MD) simulations demonstrates that LUMOS can generate valid fluorophores that meet various target specifications. Overall, these results establish LUMOS as a data-physics dual-driven framework for general fluorophore inverse design.

Neural Organ Transplantation (NOT): Checkpoint-Based Modular Adaptation for Transformer Models

arXiv:2601.13580v1 Announce Type: cross Abstract: We introduce Neural Organ Transplantation (NOT), a modular adaptation framework that enables trained transformer layers to function as reusable transferable checkpoints for domain adaptation. Unlike conventional fine-tuning approaches that tightly couple trained parameters to specific model instances and training data, NOT extracts contiguous layer subsets ("donor organs") from pre-trained models, trains them independently on domain-specific data, and saves them as standalone checkpoint files that can be transplanted into compatible recipient models without access to the original training data. Through experiments on three decoder-only transformer architectures spanning 124M to 20B parameters (GPT-2, TinyLlama, and GPT-OSS), we demonstrate that donor transplantation substantially outperforms existing adaptation methods, achieving an order-of-magnitude improvement in perplexity over LoRA while training significantly faster. The method exhibits position dependence, with early insertion positions yielding optimal results. Cross-domain transfer at billion-parameter scale reveals unexpected regularization benefits. These findings demonstrate that transformer middle layers can support efficient modular transfer for decoder-only architectures, enabling privacy-preserving expertise sharing through checkpoint distribution. We note that this approach is currently limited to decoder-only models; preliminary experiments on encoder-based architectures show reduced effectiveness.
❌