❌

Normal view

Promoting Responsible DeepSeek Deployment in Health Care: Scoping Review Comparing Grey and White Literature

Background: The rapid deployment of DeepSeek, an open-source large language model has sparked concerns of its impact on patient outcomes and safety. However, little is known about how DeepSeek is used and regulated in these facilities. Objective: This study aimed to 1) systematically review the characteristics of deployed DeepSeek in the top 100 hospitals in China; and 2) compare performances and risks from hospital disclosure with research evidence. Methods: We performed a scoping review of gray and white literature, collecting data from the top 100 Chinese hospitals. We extracted basic characteristics of DeepSeek, its aim, evaluation approach, performance, risk and hospital regulation. A coding framework was developedcovering LLMs application scenario, evaluation dimension and source of risk. Results: We identified a total of 58 DeepSeek models in 48 out of the top 100 Chinese hospitals as well as 27 studies. We observed deployed DeepSeek mainly intended to assist clinical decision making, such as patient diagnosis and treatment recommendation. However, only 36.2% hospital-deployed models clearly indicated a pre-deployment assessment, 22.4% presented assessment results, and 8.6% identified potential risks and countermeasures. We found poor transparency in hospital reporting, with none presenting evaluation details. Hospitals were likely to report DeepSeek’s higher performance and fewer risks. Conclusions: The irresponsible deployment of DeepSeek in Chinese leading hospitals poses potential risks to patient outcomes and safety. We highlight the urgent need that existing regulations should be expanded to the downstream developers and users and hospitals need to perform a more rigorous validation and transparent reporting.

PathMind: A Retrieve-Prioritize-Reason Framework for Knowledge Graph Reasoning with Large Language Models

arXiv:2511.14256v1 Announce Type: new Abstract: Knowledge graph reasoning (KGR) is the task of inferring new knowledge by performing logical deductions on knowledge graphs. Recently, large language models (LLMs) have demonstrated remarkable performance in complex reasoning tasks. Despite promising success, current LLM-based KGR methods still face two critical limitations. First, existing methods often extract reasoning paths indiscriminately, without assessing their different importance, which may introduce irrelevant noise that misleads LLMs. Second, while many methods leverage LLMs to dynamically explore potential reasoning paths, they require high retrieval demands and frequent LLM calls. To address these limitations, we propose PathMind, a novel framework designed to enhance faithful and interpretable reasoning by selectively guiding LLMs with important reasoning paths. Specifically, PathMind follows a "Retrieve-Prioritize-Reason" paradigm. First, it retrieves a query subgraph from KG through the retrieval module. Next, it introduces a path prioritization mechanism that identifies important reasoning paths using a semantic-aware path priority function, which simultaneously considers the accumulative cost and the estimated future cost for reaching the target. Finally, PathMind generates accurate and logically consistent responses via a dual-phase training strategy, including task-specific instruction tuning and path-wise preference alignment. Extensive experiments on benchmark datasets demonstrate that PathMind consistently outperforms competitive baselines, particularly on complex reasoning tasks with fewer input tokens, by identifying essential reasoning paths.

Large language models driven neural architecture search for universal and lightweight disease diagnosis on histopathology slide images

18 November 2025 at 08:00

npj Digital Medicine, Published online: 18 November 2025; doi:10.1038/s41746-025-02042-x

Large language models driven neural architecture search for universal and lightweight disease diagnosis on histopathology slide images

End to End AI System for Surgical Gesture Sequence Recognition and Clinical Outcome Prediction

arXiv:2511.11899v1 Announce Type: new Abstract: Fine-grained analysis of intraoperative behavior and its impact on patient outcomes remain a longstanding challenge. We present Frame-to-Outcome (F2O), an end-to-end system that translates tissue dissection videos into gesture sequences and uncovers patterns associated with postoperative outcomes. Leveraging transformer-based spatial and temporal modeling and frame-wise classification, F2O robustly detects consecutive short (~2 seconds) gestures in the nerve-sparing step of robot-assisted radical prostatectomy (AUC: 0.80 frame-level; 0.81 video-level). F2O-derived features (gesture frequency, duration, and transitions) predicted postoperative outcomes with accuracy comparable to human annotations (0.79 vs. 0.75; overlapping 95% CI). Across 25 shared features, effect size directions were concordant with small differences (~ 0.07), and strong correlation (r = 0.96, p

LLM4AD: Large Language Models for Autonomous Driving - Concept, Review, Benchmark, Experiments, and Future Trends

arXiv:2410.15281v4 Announce Type: replace-cross Abstract: With the broader adoption and highly successful development of Large Language Models (LLMs), there has been growing interest and demand for applying LLMs to autonomous driving technology. Driven by their natural language understanding and reasoning capabilities, LLMs have the potential to enhance various aspects of autonomous driving systems, from perception and scene understanding to interactive decision-making. In this paper, we first introduce the novel concept of designing Large Language Models for Autonomous Driving (LLM4AD), followed by a review of existing LLM4AD studies. Then, we propose a comprehensive benchmark for evaluating the instruction-following and reasoning abilities of LLM4AD systems, which includes LaMPilot-Bench, CARLA Leaderboard 1.0 Benchmark in simulation and NuPlanQA for multi-view visual question answering. Furthermore, we conduct extensive real-world experiments on autonomous vehicle platforms, examining both on-cloud and on-edge LLM deployment for personalized decision-making and motion control. Next, we explore the future trends of integrating language diffusion models into autonomous driving, exemplified by the proposed ViLaD (Vision-Language Diffusion) framework. Finally, we discuss the main challenges of LLM4AD, including latency, deployment, security and privacy, safety, trust and transparency, and personalization.

Chain-of-Scrutiny: Detecting Backdoor Attacks for Large Language Models

arXiv:2406.05948v4 Announce Type: replace-cross Abstract: Large Language Models (LLMs), especially those accessed via APIs, have demonstrated impressive capabilities across various domains. However, users without technical expertise often turn to (untrustworthy) third-party services, such as prompt engineering, to enhance their LLM experience, creating vulnerabilities to adversarial threats like backdoor attacks. Backdoor-compromised LLMs generate malicious outputs to users when inputs contain specific "triggers" set by attackers. Traditional defense strategies, originally designed for small-scale models, are impractical for API-accessible LLMs due to limited model access, high computational costs, and data requirements. To address these limitations, we propose Chain-of-Scrutiny (CoS) which leverages LLMs' unique reasoning abilities to mitigate backdoor attacks. It guides the LLM to generate reasoning steps for a given input and scrutinizes for consistency with the final output -- any inconsistencies indicating a potential attack. It is well-suited for the popular API-only LLM deployments, enabling detection at minimal cost and with little data. User-friendly and driven by natural language, it allows non-experts to perform the defense independently while maintaining transparency. We validate the effectiveness of CoS through extensive experiments on various tasks and LLMs, with results showing greater benefits for more powerful LLMs.

Interplay between gut microbial communities and metabolites modulates pan-cancer immunotherapy responses

Cell Metab. 2025 Jan 28:S1550-4131(24)00495-9. doi: 10.1016/j.cmet.2024.12.013. Online ahead of print.

ABSTRACT

Immune checkpoint blockade (ICB) therapy has revolutionized cancer treatment but remains effective in only a subset of patients. Emerging evidence suggests that the gut microbiome and its metabolites critically influence ICB efficacy. In this study, we performed a multi-omics analysis of fecal microbiomes and metabolomes from 165 patients undergoing anti-programmed cell death protein 1 (PD-1)/programmed death ligand 1 (PD-L1) therapy, identifying microbial and metabolic entities associated with treatment response. Integration of data from four public metagenomic datasets (n = 568) uncovered cross-cohort microbial and metabolic signatures, validated in an independent cohort (n = 138). An integrated predictive model incorporating these features demonstrated robust performance. Notably, we characterized five response-associated enterotypes, each linked to specific bacterial taxa and metabolites. Among these, the metabolite phenylacetylglutamine (PAGln) was negatively correlated with response and shown to attenuate anti-PD-1 efficacy in vivo. This study sheds light on the interplay among the gut microbiome, the gut metabolome, and immunotherapy response, identifying potential biomarkers to improve treatment outcomes.

PMID:39909032 | DOI:10.1016/j.cmet.2024.12.013

Global trends and risk factors in gastric cancer: a comprehensive analysis of the Global Burden of Disease Study 2021 and multi-omics data

Int J Med Sci. 2025 Jan 1;22(2):341-356. doi: 10.7150/ijms.104437. eCollection 2025.

ABSTRACT

Background: Gastric cancer (GC) remains a significant global health challenge. This study aimed to comprehensively analyze GC epidemiology and risk factors to inform prevention and intervention strategies. Methods: We analyzed the Global Burden of Disease Study 2021 data, conducted 16 different machine learning (ML) models of NHANES data, performed Mendelian randomization (MR) studies on disease phenotypes, dietary preferences, microbiome, blood-based markers, and integrated differential gene expression and expression quantitative trait loci (eQTL) data from multiple cohorts to identify factors associated with GC risk. Results: Global age-standardized disability-adjusted life year rates (ASDR) for GC declined from 886.24 to 358.42 per 100,000 population between 1990 and 2030, with significant regional disparities. Despite this decline, total disability-adjusted life years show a concerning upward trend from 2015, rising from approximately 22.9 million to a projected 24.3 million by 2030. The slope index of inequality shifted from 87 in 1990 to -184 in 2021, indicating a reversal in GC burden distribution, with higher ASDR now associated with lower socio-demographic index countries. The ML models analysis identified higher levels of clinical characteristics such as phosphorus, calcium, eosinophils percent, and triglycerides, as well as lower levels of iron and monocyte percent, may be associated with an increased risk of GC. MR analyses revealed causal associations between GC risk and disease phenotypes such as Helicobacter pylori infection, chronic gastritis, obesity, depression, and dietary preferences such as dairy and processed meats. Gut microbiome analysis showed associations with microbiome such as Phascolarctobacterium and Ruminococcaceae species. Blood-based markers analysis identified protective and risk effects for cortisol, glutamate, nicotinamide, Natural Killer %lymphocyte, CD4-CD8- T cell Absolute Count, Phosphatidylcholine (16:0_18:1), and Interleukin-1-alpha. Integrated genomic analysis identified 10 genes significantly associated with GC risk, with strong evidence for colocalization in genes such as CCR6 and PILRB. Conclusions: This systematic analysis reveals complex global trends in GC burden and identifies novel clinical, disease phenotypes, dietary preferences, microbial, blood-based, and genetic risk factors. These findings provide potential targets for improved risk stratification, prevention, and intervention strategies to reduce the global burden of GC.

PMID:39781526 | PMC:PMC11704698 | DOI:10.7150/ijms.104437

Histone demethylase KDM5D upregulation drives sex differences in colon cancer

Nature, Published online: 21 June 2023; doi:10.1038/s41586-023-06254-7

A murine colorectal cancer (CRC) model shows that mutant KRAS-STAT4-mediated upregulation of Y chromosome KDM5D contributes to the sex differences in KRAS-mutant CRC, providing an actionable therapeutic strategy for metastasis risk reduction for men afflicted with KRAS-mutant CRC.
❌