❌

Normal view

A Multi-faceted Analysis of Cognitive Abilities: Evaluating Prompt Methods with Large Language Models on the CONSORT Checklist

arXiv:2510.19139v1 Announce Type: new Abstract: Despite the rapid expansion of Large Language Models (LLMs) in healthcare, the ability of these systems to assess clinical trial reporting according to CONSORT standards remains unclear, particularly with respect to their cognitive and reasoning strategies. This study applies a behavioral and metacognitive analytic approach with expert-validated data, systematically comparing two representative LLMs under three prompt conditions. Clear differences emerged in how the models approached various CONSORT items, and prompt types, including shifts in reasoning style, explicit uncertainty, and alternative interpretations shaped response patterns. Our results highlight the current limitations of these systems in clinical compliance automation and underscore the importance of understanding their cognitive adaptations and strategic behavior in developing more explainable and reliable medical AI.

MSC-Bench: A Rigorous Benchmark for Multi-Server Tool Orchestration

arXiv:2510.19423v1 Announce Type: new Abstract: We introduce MSC-Bench, a large-scale benchmark for evaluating multi-hop, end-to-end tool orchestration by LLM agents in a hierarchical Model-Context Protocol (MCP) ecosystem. Existing benchmarks often evaluate tools in isolation, ignoring challenges such as functional overlap and cross-server orchestration, leading to overly optimistic assessments. MSC-Bench addresses these gaps by constructing ground truth through 'equal function sets', allowing objective metrics such as F1 score and reducing the dependency on LLM-as-a-judge evaluation. Organized as a five-level curriculum, it systematically tests agent capabilities from single-tool orchestration to complex cross-server planning, and robustness to out-of-scope requests. Experiments reveal that rigid hierarchies can hinder performance without co-designed strategies, and even state-of-the-art agents exhibit systemic weaknesses in robustness. MSC-Bench provides a diagnostic framework to expose these limitations and guide the development of more capable and efficient tool-using agents. The benchmark and resources are publicly available at https://github.com/snooow1029/MSC_Bench.

KnowMol: Advancing Molecular Large Language Models with Multi-Level Chemical Knowledge

arXiv:2510.19484v1 Announce Type: cross Abstract: The molecular large language models have garnered widespread attention due to their promising potential on molecular applications. However, current molecular large language models face significant limitations in understanding molecules due to inadequate textual descriptions and suboptimal molecular representation strategies during pretraining. To address these challenges, we introduce KnowMol-100K, a large-scale dataset with 100K fine-grained molecular annotations across multiple levels, bridging the gap between molecules and textual descriptions. Additionally, we propose chemically-informative molecular representation, effectively addressing limitations in existing molecular representation strategies. Building upon these innovations, we develop KnowMol, a state-of-the-art multi-modal molecular large language model. Extensive experiments demonstrate that KnowMol achieves superior performance across molecular understanding and generation tasks. GitHub: https://github.com/yzf-code/KnowMol Huggingface: https://hf.co/datasets/yzf1102/KnowMol-100K

Insights into the Unknown: Federated Data Diversity Analysis on Molecular Data

arXiv:2510.19535v1 Announce Type: cross Abstract: AI methods are increasingly shaping pharmaceutical drug discovery. However, their translation to industrial applications remains limited due to their reliance on public datasets, lacking scale and diversity of proprietary pharmaceutical data. Federated learning (FL) offers a promising approach to integrate private data into privacy-preserving, collaborative model training across data silos. This federated data access complicates important data-centric tasks such as estimating dataset diversity, performing informed data splits, and understanding the structure of the combined chemical space. To address this gap, we investigate how well federated clustering methods can disentangle and represent distributed molecular data. We benchmark three approaches, Federated kMeans (Fed-kMeans), Federated Principal Component Analysis combined with Fed-kMeans (Fed-PCA+Fed-kMeans), and Federated Locality-Sensitive Hashing (Fed-LSH), against their centralized counterparts on eight diverse molecular datasets. Our evaluation utilizes both, standard mathematical and a chemistry-informed evaluation metrics, SF-ICF, that we introduce in this work. The large-scale benchmarking combined with an in-depth explainability analysis shows the importance of incorporating domain knowledge through chemistry-informed metrics, and on-client explainability analyses for federated diversity analysis on molecular data.

Integrating Transparent Models, LLMs, and Practitioner-in-the-Loop: A Case of Nonprofit Program Evaluation

arXiv:2510.19799v1 Announce Type: cross Abstract: Public and nonprofit organizations often hesitate to adopt AI tools because most models are opaque even though standard approaches typically analyze aggregate patterns rather than offering actionable, case-level guidance. This study tests a practitioner-in-the-loop workflow that pairs transparent decision-tree models with large language models (LLMs) to improve predictive accuracy, interpretability, and the generation of practical insights. Using data from an ongoing college-success program, we build interpretable decision trees to surface key predictors. We then provide each tree's structure to an LLM, enabling it to reproduce case-level predictions grounded in the transparent models. Practitioners participate throughout feature engineering, model design, explanation review, and usability assessment, ensuring that field expertise informs the analysis at every stage. Results show that integrating transparent models, LLMs, and practitioner input yields accurate, trustworthy, and actionable case-level evaluations, offering a viable pathway for responsible AI adoption in the public and nonprofit sectors.

RoboGPT-R1: Enhancing Robot Planning with Reinforcement Learning

arXiv:2510.14828v2 Announce Type: replace Abstract: Improving the reasoning capabilities of embodied agents is crucial for robots to complete complex human instructions in long-view manipulation tasks successfully. Despite the success of large language models and vision language models based on Supervised Fine-Tuning (SFT) in planning tasks, they continue facing challenges in performing long-horizon manipulation tasks in complex real-world environments, owing to their restricted common sense and reasoning capabilities. Considering that aligning general-purpose vision language models to robotic planning tasks via supervised fine-tuning suffers from poor generalization and insufficient physical understanding, we propose RoboGPT-R1, a two-stage fine-tuning framework for embodied planning. In this framework, supervised training acquires foundational knowledge through expert sequences, followed by RL to address the model's shortcomings in visual-spatial understanding and reasoning. To achieve physical understanding and action sequence consistency in multi-step reasoning tasks, we design a rule-based reward function that simultaneously considers long-horizon performance and action constraint in the environment. The reasoning model, trained on Qwen2.5-VL-3B, significantly outperforms the larger-scale model, GPT-4o-mini, by 21.33% and surpasses other work trained on Qwen2.5-VL-7B by 20.33% on the EmbodiedBench benchmark.

The Right to Be Remembered: Preserving Maximally Truthful Digital Memory in the Age of AI

arXiv:2510.16206v2 Announce Type: replace Abstract: Since the rapid expansion of large language models (LLMs), people have begun to rely on them for information retrieval. While traditional search engines display ranked lists of sources shaped by search engine optimization (SEO), advertising, and personalization, LLMs typically provide a synthesized response that feels singular and authoritative. While both approaches carry risks of bias and omission, LLMs may amplify the effect by collapsing multiple perspectives into one answer, reducing users ability or inclination to compare alternatives. This concentrates power over information in a few LLM vendors whose systems effectively shape what is remembered and what is overlooked. As a result, certain narratives, individuals or groups, may be disproportionately suppressed, while others are disproportionately elevated. Over time, this creates a new threat: the gradual erasure of those with limited digital presence, and the amplification of those already prominent, reshaping collective memory. To address these concerns, this paper presents a concept of the Right To Be Remembered (RTBR) which encompasses minimizing the risk of AI-driven information omission, embracing the right of fair treatment, while ensuring that the generated content would be maximally truthful.

ScholaWrite: A Dataset of End-to-End Scholarly Writing Process

arXiv:2502.02904v4 Announce Type: replace-cross Abstract: Writing is a cognitively demanding activity that requires constant decision-making, heavy reliance on working memory, and frequent shifts between tasks of different goals. To build writing assistants that truly align with writers' cognition, we must capture and decode the complete thought process behind how writers transform ideas into final texts. We present ScholaWrite, the first dataset of end-to-end scholarly writing, tracing the multi-month journey from initial drafts to final manuscripts. We contribute three key advances: (1) a Chrome extension that unobtrusively records keystrokes on Overleaf, enabling the collection of realistic, in-situ writing data; (2) a novel corpus of full scholarly manuscripts, enriched with fine-grained annotations of cognitive writing intentions. The dataset includes \LaTeX-based edits from five computer science preprints, capturing nearly 62K text changes over four months; and (3) analyses and insights into the micro-dynamics of scholarly writing, highlighting gaps between human writing processes and the current capabilities of large language models (LLMs) in providing meaningful assistance. ScholaWrite underscores the value of capturing end-to-end writing data to develop future writing assistants that support, not replace, the cognitive work of scientists.

TinySQL: A Progressive Text-to-SQL Dataset for Mechanistic Interpretability Research

arXiv:2503.12730v5 Announce Type: replace-cross Abstract: Mechanistic interpretability research faces a gap between analyzing simple circuits in toy tasks and discovering features in large models. To bridge this gap, we propose text-to-SQL generation as an ideal task to study, as it combines the formal structure of toy tasks with real-world complexity. We introduce TinySQL, a synthetic dataset, progressing from basic to advanced SQL operations, and train models ranging from 33M to 1B parameters to establish a comprehensive testbed for interpretability. We apply multiple complementary interpretability techniques, including Edge Attribution Patching and Sparse Autoencoders, to identify minimal circuits and components supporting SQL generation. We compare circuits for different SQL subskills, evaluating their minimality, reliability, and identifiability. Finally, we conduct a layerwise logit lens analysis to reveal how models compose SQL queries across layers: from intent recognition to schema resolution to structured generation. Our work provides a robust framework for probing and comparing interpretability methods in a structured, progressively complex setting.

LongCodeBench: Evaluating Coding LLMs at 1M Context Windows

arXiv:2505.07897v3 Announce Type: replace-cross Abstract: Context lengths for models have grown rapidly, from thousands to millions of tokens in just a few years. The extreme context sizes of modern long-context models have made it difficult to construct realistic long-context benchmarks -- not only due to the cost of collecting million-context tasks but also in identifying realistic scenarios that require significant contexts. We identify code comprehension and repair as a natural testbed and challenge task for long-context models and introduce LongCodeBench (LCB), a benchmark to test LLM coding abilities in long-context scenarios. Our benchmark tests both the comprehension and repair capabilities of LCLMs in realistic and important settings by drawing from real-world GitHub issues and constructing QA (LongCodeQA) and bug fixing (LongSWE-Bench) tasks. We carefully stratify the complexity of our benchmark, enabling us to evaluate models across different scales -- ranging from Qwen2.5 14B Instruct to Google's flagship Gemini model. We find that long-context remains a weakness for all models, with performance drops such as from 29% to 3% for Claude 3.5 Sonnet, or from 70.2% to 40% for Qwen2.5. The LCB dataset is available publicly at https://huggingface.co/datasets/Steefano/LCB and the codebase to replicate the work on this paper at https://github.com/Zteefano/long-code-bench.

With Limited Data for Multimodal Alignment, Let the STRUCTURE Guide You

arXiv:2506.16895v2 Announce Type: replace-cross Abstract: Multimodal models have demonstrated powerful capabilities in complex tasks requiring multimodal alignment, including zero-shot classification and cross-modal retrieval. However, existing models typically rely on millions of paired multimodal samples, which are prohibitively expensive or infeasible to obtain in many domains. In this work, we explore the feasibility of building multimodal models with limited amount of paired data by aligning pretrained unimodal foundation models. We show that high-quality alignment is possible with as few as tens of thousands of paired samples$\unicode{x2013}$less than $1\%$ of the data typically used in the field. To achieve this, we introduce STRUCTURE, an effective regularization technique that preserves the neighborhood geometry of the latent space of unimodal encoders. Additionally, we show that aligning last layers is often suboptimal and demonstrate the benefits of aligning the layers with the highest representational similarity across modalities. These two components can be readily incorporated into existing alignment methods, yielding substantial gains across 24 zero-shot image classification and retrieval benchmarks, with average relative improvement of $51.6\%$ in classification and $91.8\%$ in retrieval tasks. Our results highlight the effectiveness and broad applicability of our framework for limited-sample multimodal learning and offer a promising path forward for resource-constrained domains.

ACT: Agentic Classification Tree

arXiv:2509.26433v2 Announce Type: replace-cross Abstract: When used in high-stakes settings, AI systems are expected to produce decisions that are transparent, interpretable, and auditable, a requirement increasingly expected by regulations. Decision trees such as CART provide clear and verifiable rules, but they are restricted to structured tabular data and cannot operate directly on unstructured inputs such as text. In practice, large language models (LLMs) are widely used for such data, yet prompting strategies such as chain-of-thought or prompt optimization still rely on free-form reasoning, limiting their ability to ensure trustworthy behaviors. We present the Agentic Classification Tree (ACT), which extends decision-tree methodology to unstructured inputs by formulating each split as a natural-language question, refined through impurity-based evaluation and LLM feedback via TextGrad. Experiments on text benchmarks show that ACT matches or surpasses prompting-based baselines while producing transparent and interpretable decision paths.

Landscape of T-cell exhaustion heterogeneity and HBV integration in virus-related HCC revealed by whole-exome, transcriptome, and single-cell sequencing

JHEP Rep. 2025 Jul 10;7(11):101518. doi: 10.1016/j.jhepr.2025.101518. eCollection 2025 Nov.

ABSTRACT

BACKGROUND & AIMS: To enhance our understanding of the tumor immune microenvironment (TIME) in hepatocellular carcinoma (HCC), we investigated the heterogeneity of T-cell exhaustion and its association with HBV integrations and direct oncogenic potential in HCC.

METHODS: We conducted a multi-omics analysis, including single-cell RNA sequencing, whole-exome sequencing, whole-transcriptome sequencing, and next-generation sequencing (NGS)-based HBV integration analysis, in eight patients with virus-related HCC. For validation, bulk RNA sequencing and NGS-based HBV integration analysis were performed in an independent cohort (n = 106).

RESULTS: Based on the expression scores of exhaustion markers in effector CD8+ T cells, patients were classified into high (n = 2) and low (n = 6) exhaustion groups (p <0.001). The high-exhaustion group exhibited higher clonal expansion (Gini index: 0.83 vs. 0.48, p = 0.006) and sharing of CD8+ T effector memory and cycling T cells with elevated exhaustion markers. This group also showed increased clonal expansion of CD4+ regulatory T cells and follicular helper T cells (p <0.001) with higher PDCD1 expression. In addition, the high-exhaustion group had higher TP53 mutation rates and signature scores for proliferation subtypes compared with the low-exhaustion group, who predominantly harbored TERT mutations. Moreover, the high-exhaustion group demonstrated more pronounced HBV integrations with elevated intrahepatic covalently closed circular DNA (cccDNA) and pregenomic (pg)RNA levels. Similarly, in the validation cohort, the high-exhaustion group (n = 28) demonstrated stronger proliferation subtype signatures (p <0.001), along with higher HBV integrations, S-fusion transcripts, and an increased intrahepatic viral reservoir (cccDNA/pgRNA) (p <0.05) compared with the low-exhaustion group (n = 78).

CONCLUSIONS: Our study revealed the heterogeneity in T-cell exhaustion in the TIME of HCC, along with differences in HBV integrations and molecular subtypes. These findings provide insight into the intricate relationship between high exhaustion, proliferation subtype, increased HBV integrations, and enhanced HBV-induced oncogenic potential in virus-related HCC.

IMPACT AND IMPLICATIONS: This study provides a comprehensive immune landscape of T-cell exhaustion using multi-omics analysis, offering critical insights into T cell heterogeneity in virus-related HCC. It establishes a strong association between higher HBV integration, enhanced oncogenic potential, T-cell exhaustion, and proliferation subtypes in HCC. Our results also establish a basis for personalized therapies tailored to the immune-exhaustion status within the TIME of each patient with HCC.

PMID:41113120 | PMC:PMC12529496 | DOI:10.1016/j.jhepr.2025.101518

Single-cell multi-omics analysis reveals cancer regulatory elements of transcriptional programs and clinical implications

Cell Death Dis. 2025 Oct 21;16(1):746. doi: 10.1038/s41419-025-08060-7.

ABSTRACT

The regulatory mechanisms governing transcriptional programs in the cancer genome remain elusive, particularly those concerning cell-type specificity. We carefully curated single-cell assay for transposase-accessible chromatin sequencing (scATAC-seq) and single-cell RNA sequencing (scRNA-seq) data from eight distinct carcinoma tissues, including breast, skin, colon, endometrium, lung, ovary, liver, and kidney. Using single-cell multi-omics analysis, we identified extensive open chromatin regions and constructed peak-gene link networks, which can reveal distinct cancer gene regulation and genetic risks. We further explored conserved epigenetic regulation across cell types within cancer and elucidated their functional implications. Moreover, we identified cell-type-associated transcription factors (TFs) that regulate key cellular functions, such as the TEAD family of TFs, which widely control cancer-related signaling pathways in tumor cells. In colon cancer, we further identified tumor-specific TFs that are more highly activated in tumor cells than in normal epithelial cells, including CEBPG, LEF1, SOX4, TCF7, and TEAD4, which are pivotal in driving malignant transcriptional programs and represent potential therapeutic targets, as corroborated by single-cell sequencing data from multiple sources and in vitro experiments. Our findings provide a comprehensive understanding of the regulatory dynamics underlying carcinomas and offer valuable insights into potential therapeutic interventions.

PMID:41120274 | PMC:PMC12541060 | DOI:10.1038/s41419-025-08060-7

Exploring Patient Perspectives, Engagement, and Output Quality in Doctor-Supervised Use of Artificial Intelligence During Informed Consent Consultation With ChatGPT and Retrieval Augmented Generation (RAG): Quantitative Exploratory Study

Background: Comprehensive preoperative education is essential for optimizing outcomes and ensuring informed consent in patients undergoing total hip arthroplasty (THA). Emerging artificial intelligence (AI) tools, such as ChatGPT, offer scalable support for patient education, but their clinical application requires rigorous evaluation to ensure accuracy, safety, and trust. Objective: This study assessed patients’ preferences and satisfaction with AI-assisted informed consent in THA, comparing traditional physician consultations to those supported by native ChatGPT and a customized version enhanced with retrieval-augmented generation (RAG). It also examined how state anxiety and general attitudes toward AI affect preferences for AI-supported consent and whether RAG integration improves ChatGPT response quality. Methods: A total of 36 patients scheduled for elective THA were assigned to one of three groups (12 each): (1) standard physician-only consultations (control), (2) physician-assisted consultations supported by native ChatGPT, and (3) supported by ChatGPT enhanced through RAG. Data collection involved standardized Likert scale questionnaires assessing patient satisfaction with the consent process, perceived informedness, anxiety levels, and attitudes toward AI. The ChatGPT responses were independently evaluated by physicians for relevance, accuracy, clarity, completeness, adherence to evidence-based guidelines, and appropriate length. Instances of hallucinations, factually incorrect or misleading outputs, were identified and rated by severity. Statistical analyses compared outcomes across groups and explored associations. Results: Patients interacting with the ChatGPT+RAG model reported significantly higher satisfaction levels with information delivery (P=.01) and perceived level of informedness (P=.01) than those using the native ChatGPT model. The mean number of patient questions in the control group was 20, compared with 39 in the native ChatGPT group (P=.06) and 52 in the ChatGPT+RAG group (P=.002). The majority of participants across all groups preferred a human clinician providing less accurate information over a more accurate AI-only assistant. These preferences were not influenced by sociodemographic variables (age, gender, and education), health literacy, state anxiety, or general attitudes toward AI. The ChatGPT+RAG model outperformed the native ChatGPT model across all evaluated response quality dimensions (all P<.01) and exhibited a significantly lower hallucination rate (5/52, 10% versus 15/39, 38%; P=.002). Conclusions: Integrating RAG with ChatGPT significantly improves the quality, clarity, and reliability of preoperative information, enhancing patient satisfaction and engagement beyond native ChatGPT. However, patients maintain a strong preference for physician-led informed consent, underscoring the role of AI chatbots as complementary tools rather than replacements. These findings support the cautious adoption of customized AI assistants to augment, not substitute, human interaction in surgical consent processes. Trial Registration:

Global Adoption, Promotion, Impact, and Deployment of AI in Patient Care, Health Care Delivery, Management, and Health Care Systems Leadership: Cross-Sectional Survey

Background: Artificial intelligence (AI) is increasingly being integrated into health care, offering a wide array of benefits. Current AI applications encompass patients’ diagnosis, treatment, data mining, and more to enhance patient care and quality of life. It is also democratizing access to expert support by providing timely and accurate disease diagnoses, better clinical management, quicker drug discovery, improved disease prevention, big data management, and health protection. Objective: The aim of the study is to document AI adoption in health care, assess participants’ perception on its usefulness in the management of health care delivery and leadership of health care systems, and identify characteristics of early adopters. Methods: We conducted a worldwide cross-sectional survey across all 6 inhabited continents using a self-administered questionnaire developed with the Qualtrics electronic data collection tool. This was piloted and reviewed to ensure completeness, accuracy, acceptability, cultural sensitivity, and relevance. Respondents were recruited by individualized email, following identification from professional associations or organizations, professional networks, and social media. Data were analyzed using SPSS (IBM Corp), with results presented as narrative, charts, and tables. Results: In total, 506 health care professionals completed the survey. While 92.3% (467/506) of respondents believed that AI has a role in patient care and health care management, only 76.5% (300/392) were willing to support AI adoption and embedding in their organization. Although top managers are mainly responsible for adoption processes, staff training remains low. AI is currently used mostly for diagnosis, patient care, and precision medicine. These uses of AI will continue in the near future, but in different ways. AI adoption was highest in Europe and lowest in Africa. Black or African American people were more likely to support AI adoption than White and Asian people. Poor knowledge of AI, fear of job loss, and resistance to change were the top barriers to AI adoption and embedding. Conclusions: AI use in health is global, but the adoption rate varies by geography and individual characteristics. AI adoption communication by executive health care management is poor, as is the level of training of health care staff. To improve AI adoption, management should improve communication with their teams, provide training on AI to their workers, and help individuals understand how AI works. Barriers such as ethical issues around data ownership and use should be addressed. African organizations should be proactive and invest in AI adoption early, so that they are not left behind in the AI revolution.

Assessing Large Language Models in Building a Structured Dataset From AskDocs Subreddit Data: Methodological Study

Background: In an era marked by the blooming reliance on digital platforms for healthcare consultation, the subreddit r/AskDocs has emerged as a pivotal forum. However, the vast, unstructured nature of forum data presents a formidable challenge; the extraction and meaningful analysis of such data require advanced tools that can navigate the complexities of language and context inherent in user-generated content. Objective: Our objective was to evaluate employing Large Language Models (LLMs) to systematically transform the rich, unstructured textual data from AskDocs into a structured dataset, an approach that aligns more closely with human cognitive processes compared to traditional data extraction methods. Methods: We developed a dataset of Reddit posts from r/AskDocs by extracting key information via human annotators. Then using specially engineered prompts we used state-of-the-art Large Language Models (LLMs) to extract data from posts and compared the results. The variation in the LLMs were further compared to the humans to show similarity. Results: Our findings indicate that LLMs not only match but, in several aspects, surpass even highly educated humans in extracting information, including both demographic and context details, from unstructured texts. Conclusions: This study not only validates the use of LLMs for analyzing digital healthcare communications but also opens new avenues for understanding online behaviors and interactions, signaling a shift towards more sophisticated methodologies in digital research and practice.

STAT+: European oncology experts roll out guidance for use of large language models in clinical care 

22 October 2025 at 22:13

BERLIN — The leading professional organization for European oncologists has rolled out its first set of guidance on how its members should use large language models, a type of artificial intelligence, in cancer medicine. 

“The oncology community cannot ignore the potential benefits which AI technology can provide to cancer patients,” the authors of the guidance wrote, while simultaneously acknowledging that there aren’t enough evaluations of the chatbots available to patients or tools available to doctors to address the risks associated with generative AI in medicine.

The guidance’s release — it was published last Saturday in the Annals of Oncology — coincided with the annual meeting for the European Society for Clinical Oncology in Berlin. The American Society of Clinical Oncology has issued its own set of principles for the responsible use of AI in cancer medicine, but has not released recommendations specific to large language models (LLMs).

Continue to STAT+ to read the full story…

© Adobe

❌