❌

Normal view

Cognitive bias in LLM reasoning compromises interpretation of clinical oncology notes

arXiv:2511.20680v1 Announce Type: cross Abstract: Despite high performance on clinical benchmarks, large language models may reach correct conclusions through faulty reasoning, a failure mode with safety implications for oncology decision support that is not captured by accuracy-based evaluation. In this two-cohort retrospective study, we developed a hierarchical taxonomy of reasoning errors from GPT-4 chain-of-thought responses to real oncology notes and tested its clinical relevance. Using breast and pancreatic cancer notes from the CORAL dataset, we annotated 600 reasoning traces to define a three-tier taxonomy mapping computational failures to cognitive bias frameworks. We validated the taxonomy on 822 responses from prostate cancer consult notes spanning localized through metastatic disease, simulating extraction, analysis, and clinical recommendation tasks. Reasoning errors occurred in 23 percent of interpretations and dominated overall errors, with confirmation bias and anchoring bias most common. Reasoning failures were associated with guideline-discordant and potentially harmful recommendations, particularly in advanced disease management. Automated evaluators using state-of-the-art language models detected error presence but could not reliably classify subtypes. These findings show that large language models may provide fluent but clinically unsafe recommendations when reasoning is flawed. The taxonomy provides a generalizable framework for evaluating and improving reasoning fidelity before clinical deployment.

How Do Companies Manage the Environmental Sustainability of AI? An Interview Study About Green AI Efforts and Regulations

arXiv:2505.07317v2 Announce Type: replace-cross Abstract: With the ever-growing adoption of artificial intelligence (AI), AI-based software and its negative impact on the environment are no longer negligible, and studying and mitigating this impact has become a critical area of research. However, it is currently unclear which role environmental sustainability plays during AI adoption in industry and how AI regulations influence Green AI practices and decision-making in industry. We therefore aim to investigate the Green AI perception and management of industry practitioners. To this end, we conducted a total of 11 interviews with participants from 10 different organizations that adopted AI-based software. The interviews explored three main themes: AI adoption, current efforts in mitigating the negative environmental impact of AI, and the influence of the EU AI Act and the Corporate Sustainability Reporting Directive (CSRD). Our findings indicate that 9 of 11 participants prioritized business efficiency during AI adoption, with minimal consideration of environmental sustainability. Monitoring and mitigation of AI's environmental impact were very limited. Only one participant monitored negative environmental effects. Regarding applied mitigation practices, six participants reported no actions, with the others sporadically mentioning techniques like prompt engineering, relying on smaller models, or not overusing AI. Awareness and compliance with the EU AI Act are low, with only one participant reporting on its influence, while the CSRD drove sustainability reporting efforts primarily in larger companies. All in all, our findings reflect a lack of urgency and priority for sustainable AI among these companies. We suggest that current regulations are not very effective, which has implications for policymakers. Additionally, there is a need to raise industry awareness, but also to provide user-friendly techniques and tools for Green AI practices.

scGALA advances graph link prediction-based cell alignment for comprehensive data integration and harmonization

Nat Commun. 2025 Nov 26. doi: 10.1038/s41467-025-66644-5. Online ahead of print.

ABSTRACT

Single-cell technologies have transformed our understanding of cellular heterogeneity through multimodal data acquisition. However, robust cell alignment remains a major challenge for data integration and harmonization, including batch correction, label transfer, and multi-omics integration. Many existing methods constrain alignment based on rigid feature-wise distance metrics, limiting their ability to capture accurate cell correspondence across diverse cell populations and conditions. We introduce scGALA, a graph-based learning framework that redefines cell alignment by combining graph attention networks with a score-driven, task-independent optimization strategy. scGALA constructs enriched graphs of cell-cell relationships by integrating gene expression profiles with auxiliary information, such as spatial coordinates, and iteratively refines alignment via self-supervised graph link prediction, where a deep neural network is trained to identify and reinforce high-confidence correspondences across datasets. In extensive benchmarks, scGALA identifies over 25 percent more high-confidence alignments without compromising accuracy. By improving the core step of cell alignment, scGALA serves as a versatile enhancer for a wide range of single-cell data integration tasks.

PMID:41298467 | DOI:10.1038/s41467-025-66644-5

Clinician-Directed Large Language Model Software Generation for Therapeutic Interventions in Physical Rehabilitation

arXiv:2511.18274v1 Announce Type: cross Abstract: Digital health interventions are increasingly used in physical and occupational therapy to deliver home exercise programs via sensor equipped devices such as smartphones, enabling remote monitoring of adherence and performance. However, digital interventions are typically programmed as software before clinical encounters as libraries of parametrized exercise modules targeting broad patient populations. At the point of care, clinicians can only select modules and adjust a narrow set of parameters like repetitions, so patient specific needs that emerge during encounters, such as distinct movement limitations, and home environments, are rarely reflected in the software. We evaluated a digital intervention paradigm that uses large language models (LLMs) to translate clinicians' exercise prescriptions into intervention software. In a prospective single arm feasibility study with 20 licensed physical and occupational therapists and a standardized patient, clinicians created 40 individualized upper extremity programs (398 instructions) that were automatically translated into executable software. Our results show a 45% increase in the proportion of personalized prescriptions that can be implemented as software compared with a template based benchmark, with unanimous consensus among therapists on ease of use. The LLM generated software correctly delivered 99.78% (397/398) of instructions as prescribed and monitored performance with 88.4% (352/398) accuracy, with 90% (18/20) of therapists judged it safe to interact with patients, and 75% (15/20) expressed willingness to adopt it. To our knowledge, this is the first prospective evaluation of clinician directed intervention software generation with LLMs in healthcare, demonstrating feasibility and motivating larger trials to assess clinical effectiveness and safety in real patient populations.

AI-Assisted Cardiovascular Risk Assessment by General Practitioners in Resource-Constrained Indonesian Settings Using a Conceptual Prototype: Randomized Controlled Study

Background: Preventive strategies integrated with digital health and artificial intelligence (AI), have significant potential to mitigate the global burden of atherosclerotic cardiovascular disease (ASCVD). AI-enabled clinical decision support (CDS) systems increasingly provide patient-specific insights beyond traditional risk factors. Despite these advances, their capacity to enhance clinical decision-making in resource-constrained settings remains largely unexplored. Objective: We conducted a randomised controlled study to assess the effect of AI-based CDS on 10-year ASCVD risk assessment and management in primary prevention. Methods: In a three-way within-subject randomised design, doctors completed nine clinical vignettes representative of primary care presentations in a resource-constrained outpatient setting. For each vignette, participants assessed 10-year ASCVD risk and made management decisions using either a conceptual prototype of AI-based CDS, automated CDS, or no decision support. The conceptual prototype represented contemporary risk calculators based on traditional machine learning models (e.g., random forest, neural networks, logistic regression) that incorporate additional predictors alongside traditional risk factors. Primary outcomes were correct risk assessment and patient management (prescription of aspirin, statins, and anti-hypertensives; referral for advanced examinations). Decision-making time and perceptions about AI utility were also measured. Results: 102 doctors from all seven geographical regions of Indonesia participated. Most participants were 26–35 years (83%), 56% male, with a median of six years of clinical experience (IQR=4.75). AI-based CDS improved risk assessment by 27% (2(2, n=102) = 48.875, P<.001 compared to unassisted or one additional correct risk classification for every patients where doctors use ai needed treat nnt="3.7;" ci the prescription of statins also improved by n="102)=" p in pairwise comparisons assisted with ai-based cds correctly assessed significantly more cases adjusted and prescribed appropriate statin often medium effect size r=".35)" control. ai-assisted required less time marginal means sec vs f however improvements aspirin anti-hypertensives did not reach statistical significance. no improvement was observed referral decisions. participants generally viewed positively agreeing strongly that they would follow its recommendations indicating it if given access. believed could enhance efficiency assessment particularly high-volume primary care settings while noting need verify against clinical guidelines each patient. conclusions: coupled reduced decision-making highlight potential utility ascvd resource-constrained efficient healthcare resources is crucial. further research ascertain whether this online study translate real-world low-resource settings.>

Decoding the cholesterol-apoptosis axis in HCC: a machine learning-based multi-omics integration and single-cell transcriptomic analysis

25 November 2025 at 19:00

Discov Oncol. 2025 Nov 25;16(1):2162. doi: 10.1007/s12672-025-04010-z.

ABSTRACT

Liver hepatocellular carcinoma (LIHC), a predominant form of primary hepatic malignancy, demonstrates a progressively escalating global incidence, imposing substantial health and economic burdens on patients and society. Early diagnosis remains challenging, often resulting in late-stage detection, which limits the efficacy of current therapeutic strategies. This study systematically examines the transcriptional signatures of apoptosis-associated and cholesterol metabolic pathways in LIHC, providing insights into its underlying mechanisms and identifying potential prognostic markers. We employed multi-omics and machine learning to evaluate gene expression variations and construct a prognostic risk scoring model. This study identified apoptosis- and cholesterol metabolism-related differentially expressed genes (ACMRDEGs). Importantly, LASSO regression analysis identified six hub genes (EPHX2, FABP5, SQLE, ADH4, HMGCS2, and CYP7A1) as critical prognostic biomarkers, demonstrating significant correlation with overall survival (OS). Furthermore, immune cell infiltration analysis indicated significant differences in 12 immune cell types within LIHC microenvironment, underscoring the immune system's involvement in disease progression. cholesterol and alcohol metabolism pathways were significantly enriched among hub gene modules, as quantified by multiple gene enrichment analyses. Single-cell analysis identified six major cell types, providing a deeper understanding of the cellular heterogeneity within LIHC. In summarize, this study presents the first integrated apoptosis-cholesterol metabolic pathway-based six-gene prognostic model for LIHC, validated for robustness across multiple cohorts, which may facilitate personalized therapeutic strategies and refined risk assessment in clinical practice.

PMID:41288805 | PMC:PMC12647489 | DOI:10.1007/s12672-025-04010-z

Decoding the cholesterol-apoptosis axis in HCC: a machine learning-based multi-omics integration and single-cell transcriptomic analysis

Discov Oncol. 2025 Nov 25;16(1):2162. doi: 10.1007/s12672-025-04010-z.

ABSTRACT

Liver hepatocellular carcinoma (LIHC), a predominant form of primary hepatic malignancy, demonstrates a progressively escalating global incidence, imposing substantial health and economic burdens on patients and society. Early diagnosis remains challenging, often resulting in late-stage detection, which limits the efficacy of current therapeutic strategies. This study systematically examines the transcriptional signatures of apoptosis-associated and cholesterol metabolic pathways in LIHC, providing insights into its underlying mechanisms and identifying potential prognostic markers. We employed multi-omics and machine learning to evaluate gene expression variations and construct a prognostic risk scoring model. This study identified apoptosis- and cholesterol metabolism-related differentially expressed genes (ACMRDEGs). Importantly, LASSO regression analysis identified six hub genes (EPHX2, FABP5, SQLE, ADH4, HMGCS2, and CYP7A1) as critical prognostic biomarkers, demonstrating significant correlation with overall survival (OS). Furthermore, immune cell infiltration analysis indicated significant differences in 12 immune cell types within LIHC microenvironment, underscoring the immune system's involvement in disease progression. cholesterol and alcohol metabolism pathways were significantly enriched among hub gene modules, as quantified by multiple gene enrichment analyses. Single-cell analysis identified six major cell types, providing a deeper understanding of the cellular heterogeneity within LIHC. In summarize, this study presents the first integrated apoptosis-cholesterol metabolic pathway-based six-gene prognostic model for LIHC, validated for robustness across multiple cohorts, which may facilitate personalized therapeutic strategies and refined risk assessment in clinical practice.

PMID:41288805 | DOI:10.1007/s12672-025-04010-z

When Alignment Fails: Multimodal Adversarial Attacks on Vision-Language-Action Models

arXiv:2511.16203v2 Announce Type: replace-cross Abstract: Vision-Language-Action models (VLAs) have recently demonstrated remarkable progress in embodied environments, enabling robots to perceive, reason, and act through unified multimodal understanding. Despite their impressive capabilities, the adversarial robustness of these systems remains largely unexplored, especially under realistic multimodal and black-box conditions. Existing studies mainly focus on single-modality perturbations and overlook the cross-modal misalignment that fundamentally affects embodied reasoning and decision-making. In this paper, we introduce VLA-Fool, a comprehensive study of multimodal adversarial robustness in embodied VLA models under both white-box and black-box settings. VLA-Fool unifies three levels of multimodal adversarial attacks: (1) textual perturbations through gradient-based and prompt-based manipulations, (2) visual perturbations via patch and noise distortions, and (3) cross-modal misalignment attacks that intentionally disrupt the semantic correspondence between perception and instruction. We further incorporate a VLA-aware semantic space into linguistic prompts, developing the first automatically crafted and semantically guided prompting framework. Experiments on the LIBERO benchmark using a fine-tuned OpenVLA model reveal that even minor multimodal perturbations can cause significant behavioral deviations, demonstrating the fragility of embodied multimodal alignment.

ConCISE: A Reference-Free Conciseness Evaluation Metric for LLM-Generated Answers

arXiv:2511.16846v1 Announce Type: cross Abstract: Large language models (LLMs) frequently generate responses that are lengthy and verbose, filled with redundant or unnecessary details. This diminishes clarity and user satisfaction, and it increases costs for model developers, especially with well-known proprietary models that charge based on the number of output tokens. In this paper, we introduce a novel reference-free metric for evaluating the conciseness of responses generated by LLMs. Our method quantifies non-essential content without relying on gold standard references and calculates the average of three calculations: i) a compression ratio between the original response and an LLM abstractive summary; ii) a compression ratio between the original response and an LLM extractive summary; and iii) wordremoval compression, where an LLM removes as many non-essential words as possible from the response while preserving its meaning, with the number of tokens removed indicating the conciseness score. Experimental results demonstrate that our proposed metric identifies redundancy in LLM outputs, offering a practical tool for automated evaluation of response brevity in conversational AI systems without the need for ground truth human annotations.

SMILE: A Composite Lexical-Semantic Metric for Question-Answering Evaluation

arXiv:2511.17432v1 Announce Type: cross Abstract: Traditional evaluation metrics for textual and visual question answering, like ROUGE, METEOR, and Exact Match (EM), focus heavily on n-gram based lexical similarity, often missing the deeper semantic understanding needed for accurate assessment. While measures like BERTScore and MoverScore leverage contextual embeddings to address this limitation, they lack flexibility in balancing sentence-level and keyword-level semantics and ignore lexical similarity, which remains important. Large Language Model (LLM) based evaluators, though powerful, come with drawbacks like high costs, bias, inconsistency, and hallucinations. To address these issues, we introduce SMILE: Semantic Metric Integrating Lexical Exactness, a novel approach that combines sentence-level semantic understanding with keyword-level semantic understanding and easy keyword matching. This composite method balances lexical precision and semantic relevance, offering a comprehensive evaluation. Extensive benchmarks across text, image, and video QA tasks show SMILE is highly correlated with human judgments and computationally lightweight, bridging the gap between lexical and semantic evaluation.

Artificial Intelligence Index Report 2025

arXiv:2504.07139v3 Announce Type: replace Abstract: Welcome to the eighth edition of the AI Index report. The 2025 Index is our most comprehensive to date and arrives at an important moment, as AI's influence across society, the economy, and global governance continues to intensify. New in this year's report are in-depth analyses of the evolving landscape of AI hardware, novel estimates of inference costs, and new analyses of AI publication and patenting trends. We also introduce fresh data on corporate adoption of responsible AI practices, along with expanded coverage of AI's growing role in science and medicine. Since its founding in 2017 as an offshoot of the One Hundred Year Study of Artificial Intelligence, the AI Index has been committed to equipping policymakers, journalists, executives, researchers, and the public with accurate, rigorously validated, and globally sourced data. Our mission has always been to help these stakeholders make better-informed decisions about the development and deployment of AI. In a world where AI is discussed everywhere - from boardrooms to kitchen tables - this mission has never been more essential. The AI Index continues to lead in tracking and interpreting the most critical trends shaping the field - from the shifting geopolitical landscape and the rapid evolution of underlying technologies, to AI's expanding role in business, policymaking, and public life. Longitudinal tracking remains at the heart of our mission. In a domain advancing at breakneck speed, the Index provides essential context - helping us understand where AI stands today, how it got here, and where it may be headed next. Recognized globally as one of the most authoritative resources on artificial intelligence, the AI Index has been cited in major media outlets such as The New York Times, Bloomberg, and The Guardian; referenced in hundreds of academic papers; and used by policymakers and government agencies around the world.

Multimodal analysis of whole slide images in colorectal cancer

npj Digital Medicine, Published online: 24 November 2025; doi:10.1038/s41746-025-02095-y

Multimodal analysis of whole slide images in colorectal cancer

Health care Experiences of Educated Young Adults With Blindness in the Digital Age: Qualitative Study

Background: The rapid advancement of digital health technologies (DHTs) offers substantial potential for improving healthcare access, yet it simultaneously risks exacerbating existing inequities for marginalized populations. Previous research on the digital divide has often treated individuals with blindness as a homogenous group, primarily focusing on barriers related to digital access and skills. However, less is known about the nuanced experiences of specific subgroups, such as educated and digitally literate young adults. This study focuses on this demographic to understand how their advanced digital capabilities interact with systemic and infrastructural barriers in healthcare. Objective: This qualitative study aimed to explore the lived healthcare experiences of educated young adults with blindness in China, specifically identifying how DHTs simultaneously contribute to their empowerment and exclusion. Methods: Eligible participants were educated young adults with blindness in China (aged 18-30 years, Mandarin speakers, smartphone users, and holding or pursuing higher education). A total of 12 semi-structured interviews were conducted in Mandarin during September 2024. All interviews were audio-recorded and transcribed verbatim. An inductive thematic analysis was employed to interpret the data and identify key themes. Results: Participants’ experiences highlighted an “empowered but excluded” dynamic. Seven key themes emerged, categorized into empowerment and exclusion. Empowerment themes included: (1) digital platforms empowering self-management and healthcare access, where DHTs enabled independent appointment booking and access to comprehensive health information; and (2) digital platforms empowering for finding medical visit companions, facilitating the discovery of companions for physical and emotional support. Exclusion themes comprised: (3) inaccessible online appointment systems, due to non-inclusive designs; (4) inaccessible healthcare environments and information formats, stemming from non-accessible self-service machines and written materials; (5) lack of provider competencies in respecting patient autonomy, as providers often assumed digital incompetence; (6) data privacy and security concerns, heightened by increased digitalization and reliance on assistive tools; and (7) challenges related to the quality and consistency of online companion support, highlighting the limitations of platform-based assistance. Conclusions: Our findings reveal an “empowered but excluded” dynamic: the potential for digital empowerment and enhanced independence is often curtailed by systematic barriers. Addressing this necessitates a multifaceted approach: enhancing technological accessibility through robust standards adherence and inclusive co-design processes; improving healthcare provider competencies in patient-centered care via targeted training; and empowering educated young blind adults by building their capacity for self-determination to achieve equitable healthcare access.

Benchmark on Drug Target Interaction Modeling from a Drug Structure Perspective

arXiv:2407.04055v2 Announce Type: replace-cross Abstract: The prediction modeling of drug-target interactions is crucial to drug discovery and design, which has seen rapid advancements owing to deep learning technologies. Recently developed methods, such as those based on graph neural networks (GNNs) and Transformers, demonstrate exceptional performance across various datasets by effectively extracting structural information. However, the benchmarking of these novel methods often varies significantly in terms of hyperparameter settings and datasets, which limits algorithmic progress. In view of these, we conducted a comprehensive survey and benchmark for drug-target interaction modeling from a structural perspective via integrating tens of explicit (i.e., GNN-based) and implicit (i.e., Transformer-based) structure learning algorithms. We conducted a macroscopical comparison between these two classes of encoding strategies as well as the different featurization techniques that inform molecules' chemical and physical properties. We then carry out the microscopical comparison between all the integrated models across the six datasets via comprehensively benchmarking their effectiveness and efficiency. To ensure fairness, we investigate model performance under individually optimized configuration. Remarkably, the summarized insights from the benchmark studies lead to the design of model combos. We demonstrate that our combos can achieve new state-of-the-art performance on various datasets associated with cost-effective memory and computation.

A Workflow for Full Traceability of AI Decisions

arXiv:2511.11275v2 Announce Type: replace Abstract: An ever increasing number of high-stake decisions are made or assisted by automated systems employing brittle artificial intelligence technology. There is a substantial risk that some of these decision induce harm to people, by infringing their well-being or their fundamental human rights. The state-of-the-art in AI systems makes little effort with respect to appropriate documentation of the decision process. This obstructs the ability to trace what went into a decision, which in turn is a prerequisite to any attempt of reconstructing a responsibility chain. Specifically, such traceability is linked to a documentation that will stand up in court when determining the cause of some AI-based decision that inadvertently or intentionally violates the law. This paper takes a radical, yet practical, approach to this problem, by enforcing the documentation of each and every component that goes into the training or inference of an automated decision. As such, it presents the first running workflow supporting the generation of tamper-proof, verifiable and exhaustive traces of AI decisions. In doing so, we expand the DBOM concept into an effective running workflow leveraging confidential computing technology. We demonstrate the inner workings of the workflow in the development of an app to tell poisonous and edible mushrooms apart, meant as a playful example of high-stake decision support.

The dual immunomodulatory role of B cells in tumorigenesis: mechanisms, microenvironment crosstalk, and therapeutic implications

Front Immunol. 2025 Oct 30;16:1649812. doi: 10.3389/fimmu.2025.1649812. eCollection 2025.

ABSTRACT

B lymphocytes exhibit a multifaceted and context-dependent role in tumor biology, acting as both promoters and suppressors of malignancy through dynamic interactions within the tumor microenvironment (TME). This review synthesizes current evidence on the dual functions of B cells in tumor immunity, highlighting their capacity to orchestrate antitumor responses via antigen presentation, antibody-dependent cytotoxicity, and tertiary lymphoid structure (TLS)-mediated T cell activation, while paradoxically driving immunosuppression through regulatory B cells (Bregs), pro-angiogenic signaling, and immune checkpoint modulation. Key mechanisms include TLS formation, which enhances cytotoxic T cell priming and correlates with improved immunotherapy outcomes, and Breg-mediated secretion of IL-10/TGF-β, which fosters T cell exhaustion and myeloid-derived suppressor cell recruitment. Tumor-type specificity is evident: TLS-rich malignancies like melanoma and Non-Small Cell Lung Cancer (NSCLC) show B cell-driven immune activation, whereas pancreatic and hepatocellular carcinomas demonstrate B cell functional plasticity influenced by metabolic and epigenetic reprogramming. Therapeutically, B cell-targeted strategies-including CD20 antibodies, CAR-T cells, and B cell epitope vaccines-demonstrate efficacy in hematologic and solid tumors, yet face challenges due to subset heterogeneity and sex-specific response disparities. Emerging approaches combine immune checkpoint inhibitors (ICBs) with TLS-inducing agents or exploit B cell-derived biomarkers for personalized therapy. Future directions emphasize deciphering B cell metabolic-niche crosstalk, optimizing combinatorial regimens, and leveraging spatial multiomics to resolve functional heterogeneity. By bridging mechanistic insights with clinical translation, this work underscores B cells as pivotal regulators of tumor immunity and advocates for precision strategies to harness their antitumor potential while mitigating pro-tumor plasticity.

PMID:41246318 | PMC:PMC12611826 | DOI:10.3389/fimmu.2025.1649812

❌