❌

Normal view

Efficient and accurate search in petabase-scale sequence repositories

Nature, Published online: 08 October 2025; doi:10.1038/s41586-025-09603-w

MetaGraph enables scalable indexing of large sets of DNA, RNA or protein sequences using annotated de Bruijn graphs.

Pathobiology and Genetics

Pneumologie. 2025 Oct;79(10):701-711. doi: 10.1055/a-2625-4648. Epub 2025 Oct 6.

ABSTRACT

Genetics and pathobiology were addressed at the 7th World Symposium on Pulmonary Hypertension in Task Forces 2 and 3. The Genetics Task Force also focused on precision medicine approaches, and the Pathobiology working group concentrated heavily on new omics technologies. Therefore, the following not only summarises the current state of knowledge on genetics, genetic testing methods, and molecular pathophysiological changes, but also places it in context and critically discusses it. In addition, the importance of national and international biobanks and cohorts, as well as the active involvement of patients and families, is emphasized.

PMID:41052524 | DOI:10.1055/a-2625-4648

The Potential of AI in Nursing Care: Multicenter Evaluation in Fall Risk Assessment

Background: With 28%-35% of individuals aged 65 years and older experiencing incidents of falling, falls are the second leading cause of unintentional injury–related deaths globally. Limited availability of clinical staff often impedes the timely detection and prevention of potential falls. Advances in artificial intelligence (AI) could complement existing fall risk assessment and help better allocate nursing care resources. Yet, many studies are based on small datasets from a single institution, which can restrict the generalizability of the model, and do not investigate important aspects in AI model development, such as fairness across demographic groups. Objective: This study aimed to provide a comprehensive empirical evaluation of the potential of AI in nursing care, focusing on the case of fall risk prediction. To account for demographic and contextual differences in fall incidences, we analyze data from a university and a geriatric hospital in Germany. To the best of our knowledge, these are the largest fall risk prediction datasets to date with heterogeneous data distributions. We focus on 3 key objectives. First, does AI help in improving fall risk prediction? Second, how can AI models be trained safely across different hospitals? Finally, are these models fair? Methods: This study used 2 datasets for fall risk prediction: one from a university hospital with 931,726 participants, 10,442 of whom experienced falls, and another from a geriatric hospital with 12,773 participants, 1728 of whom have fallen. State-of-the-art AI models were trained with 3 approaches, including 2 decentralized learning paradigms. First, separate models were trained on data from each hospital; second, models were retrained on the respective other dataset; and federated learning (FL) was applied to both datasets. The performance of these models was compared with the rule-based systems as implemented in clinical practice for fall risk prediction. Additional analyses were conducted to test for model fairness. Results: Our findings demonstrate that AI models consistently outperform rule-based systems across all experimental setups, with the area under the receiver operating characteristic curve of 0.735 (90% CI 0.727-0.744) for the geriatric hospital, and 0.926 (90% CI 0.924-0.928) for the university hospital. FL did not improve the fall risk prediction in this setting. Our fairness analysis ruled out disparities in model performance between different sex groups, but we found fairness infringements across age groups. Conclusions: This study demonstrates that AI models consistently outperform traditional rule-based systems across heterogeneous datasets in predicting fall risk. However, it also reveals the challenges related to demographic shifts and label distribution imbalances, which limited the FL models’ ability to generalize. While the fairness analysis indicated fair results across sex subgroups, age-related disparities emerged. Addressing data imbalances and ensuring broader representation across demographic groups will be crucial for developing more fair and generalizable models.

Evaluating Large Language Models and Retrieval-Augmented Generation Enhancement for Delivering Guideline-Adherent Nutrition Information for Cardiovascular Disease Prevention: Cross-Sectional Study

Background: Cardiovascular disease (CVD) remains the leading cause of death worldwide, yet many web-based sources on cardiovascular (CV) health are inaccessible. Large language models (LLMs) are increasingly used for health-related inquiries and offer an opportunity to produce accessible and scalable CV health information. However, because these models are trained on heterogeneous data, including unverified user-generated content, the quality and reliability of food and nutrition information on CVD prevention remain uncertain. Recent studies have examined LLM use in various health care applications, but their effectiveness for providing nutrition information remains understudied. Although retrieval-augmented generation (RAG) frameworks have been shown to enhance LLM consistency and accuracy, their use in delivering nutrition information for CVD prevention requires further evaluation. Objective: To evaluate the effectiveness of off-the-shelf and RAG-enhanced LLMs in delivering guideline-adherent nutrition information for CVD prevention, we assessed 3 off-the-shelf models (ChatGPT-4o, Perplexity, and Llama 3-70B) and a Llama 3-70B+RAG model. Methods: We curated 30 nutrition questions that comprehensively addressed CVD prevention. These were approved by a registered dietitian providing preventive cardiology services at an academic medical center and were posed 3 times to each model. We developed a 15,074-word knowledge bank incorporating the American Heart Association’s 2021 dietary guidelines and related website content to enhance Meta’s Llama 3-70B model using RAG. The model received this and a few-shot prompt as context, included citations in a Context Source section, and used vector similarity to align responses with guideline content, with the temperature parameter set to 0.5 to enhance consistency. Model responses were evaluated by 3 expert reviewers against benchmark CV guidelines for appropriateness, reliability, readability, harm, and guideline adherence. Mean scores were compared using ANOVA, with statistical significance set at P<.05. interrater agreement was measured using the cohen coefficient and readability estimated flesch-kincaid score. results: llama model scored higher than perplexity gpt-4o models on reliability appropriateness guideline adherence showed no harm.>70%; P<.001 indicated high reviewer agreement. conclusions: the llama model outperformed off-the-shelf models across all measures with no evidence of harm although responses were less readable due to technical language. scored lower on and produced some harmful responses. these findings highlight limitations demonstrate that rag system integration can enhance llm performance in delivering evidence-based dietary information.>

The Role of Data in Public Health and Health Innovation: Perspectives on Social Determinants of Health, Community-Based Data Approaches, and AI

Public health is undergoing profound transformation driven by data from the global health sector and related fields. To address systemic health disparities, scholars and practitioners are increasingly applying a data equity lens, an approach that has become even more urgent as the United States faces the erosion of public health data infrastructure. This paper summarizes insights from an April 2024 convening by the Yale School of Public Health—The Role of Data in Public Health Equity and Innovation—with intersectoral stakeholders from academia, government (local, state, and federal), healthcare, and private industry. The convening included keynote presentations and roundtables regarding the depiction of social determinants of health (SDOH) in data; effects of artificial intelligence (AI) on health data equity; and community-based models for data, providing a framework for cross-cutting discussions. Through a narrative synthesis, themes were identified and synthesized from systematically gathered information from presentations and roundtables. This process led to a set of actionable, cross-cutting recommendations to guide inclusive and impactful data practices for policymakers, public health professionals, and health innovators across diverse contexts: (1) Enable big data and interoperability connecting SDOH and health outcomes; (2) Include diverse, non-technical voices in AI and health discussions; (3) Fund research on data equity and AI in health sciences; (4) Modernize Health Insurance Portability and Accountability Act (HIPAA) with new guidelines for AI and big data; and (5) Research and conceptual frameworks are needed to elucidate interconnections between data equity and health equity.

Generative artificial intelligence in medicine

Nature Medicine, Published online: 06 October 2025; doi:10.1038/s41591-025-03983-2

This Review summarizes recent technical advancements in generative AI, outlines how new models might improve healthcare and discusses validation approaches—using lessons from recent successes and failures in the field.

Single-cell and multi-omics analysis identifies TRIM9 as a key ubiquitination regulator in pancreatic cancer

Front Immunol. 2025 Sep 19;16:1631708. doi: 10.3389/fimmu.2025.1631708. eCollection 2025.

ABSTRACT

This study investigates the role of ubiquitination-related genes in pancreatic cancer (PC) using single-cell RNA sequencing (scRNA-seq), spatial transcriptomics, and multi-omics approaches. scRNA-seq data (GSE155698) from PC samples identified 12 cell types, with endothelial cells exhibiting high ubiquitination scores (High_ubiquitin-Endo) and enriched interactions with fibroblasts/macrophages via WNT, NOTCH, and integrin pathways. Spatial transcriptomics (GSE235315) validated cell-type localization. Mendelian randomization (SMR) analysis prioritized TRIM9 as a PC-protective gene, downregulated in tumors and correlated with better survival. WGCNA revealed TRIM9-co-expressed modules linked to prognosis. A machine learning-based prognostic model (CoxBoost+RSF) integrating seven genes (TSPAN6, TSC1, RNF167, PBXIP1, LRRC49, KATNAL2, IGF2BP2) stratified patients into high/low-risk groups with distinct survival, mutation burdens, and immune infiltration. TRIM9 overexpression suppressed PC cell proliferation/migration in vitro, while knockdown enhanced malignancy. Mechanistically, TRIM9 promoted K11-linked ubiquitination and proteasomal degradation of HNRNPU, dependent on its RING domain. In vivo, TRIM9 overexpression reduced tumor growth, rescued by HNRNPU co-expression. Integrated analyses highlight TRIM9 as a tumor suppressor and prognostic biomarker, mediated via ubiquitination-dependent regulation of HNRNPU stability. This work provides insights into ubiquitination-driven PC pathogenesis and therapeutic targeting.

PMID:41050689 | PMC:PMC12491318 | DOI:10.3389/fimmu.2025.1631708

A computational medicine framework integrating multi-omics, systems biology, and artificial neural networks for Alzheimer's disease therapeutic discovery

Acta Pharm Sin B. 2025 Sep;15(9):4411-4426. doi: 10.1016/j.apsb.2025.07.018. Epub 2025 Jul 16.

ABSTRACT

The translation of genetic findings from genome-wide association studies into actionable therapeutics persists as a critical challenge in Alzheimer's disease (AD) research. Here, we present PI4AD, a computational medicine framework that integrates multi-omics data, systems biology, and artificial neural networks for therapeutic discovery. This framework leverages multi-omic and network evidence to deliver three core functionalities: clinical target prioritisation; self-organising prioritisation map construction, distinguishing AD-specific targets from those linked to neuropsychiatric disorders; and pathway crosstalk-informed therapeutic discovery. PI4AD successfully recovers clinically validated targets like APP and ESR1, confirming its prioritisation efficacy. Its artificial neural network component identifies disease-specific molecular signatures, while pathway crosstalk analysis reveals critical nodal genes (e.g., HRAS and MAPK1), drug repurposing candidates, and clinically relevant network modules. By validating targets, elucidating disease-specific therapeutic potentials, and exploring crosstalk mechanisms, PI4AD bridges genetic insights with pathway-level biology, establishing a systems genetics foundation for rational therapeutic development. Importantly, its emphasis on Ras-centred pathways-implicated in synaptic dysfunction and neuroinflammation-provides a strategy to disrupt AD progression, complementing conventional amyloid/tau-focused paradigms, with the future potential to redefine treatment strategies in conjunction with mRNA therapeutics and thereby advance translational medicine in neurodegeneration.

PMID:41049755 | PMC:PMC12491700 | DOI:10.1016/j.apsb.2025.07.018

❌