❌

Normal view

Asking the Right Questions: Benchmarking Large Language Models in the Development of Clinical Consultation Templates

arXiv:2508.01159v2 Announce Type: replace-cross Abstract: This study evaluates the capacity of large language models (LLMs) to generate structured clinical consultation templates for electronic consultation. Using 145 expert-crafted templates developed and routinely used by Stanford's eConsult team, we assess frontier models -- including o3, GPT-4o, Kimi K2, Claude 4 Sonnet, Llama 3 70B, and Gemini 2.5 Pro -- for their ability to produce clinically coherent, concise, and prioritized clinical question schemas. Through a multi-agent pipeline combining prompt optimization, semantic autograding, and prioritization analysis, we show that while models like o3 achieve high comprehensiveness (up to 92.2\%), they consistently generate excessively long templates and fail to correctly prioritize the most clinically important questions under length constraints. Performance varies across specialties, with significant degradation in narrative-driven fields such as psychiatry and pain medicine. Our findings demonstrate that LLMs can enhance structured clinical information exchange between physicians, while highlighting the need for more robust evaluation methods that capture a model's ability to prioritize clinically salient information within the time constraints of real-world physician communication.

SciAgent: A Unified Multi-Agent System for Generalistic Scientific Reasoning

arXiv:2511.08151v1 Announce Type: new Abstract: Recent advances in large language models have enabled AI systems to achieve expert-level performance on domain-specific scientific tasks, yet these systems remain narrow and handcrafted. We introduce SciAgent, a unified multi-agent system designed for generalistic scientific reasoning-the ability to adapt reasoning strategies across disciplines and difficulty levels. SciAgent organizes problem solving as a hierarchical process: a Coordinator Agent interprets each problem's domain and complexity, dynamically orchestrating specialized Worker Systems, each composed of interacting reasoning Sub-agents for symbolic deduction, conceptual modeling, numerical computation, and verification. These agents collaboratively assemble and refine reasoning pipelines tailored to each task. Across mathematics and physics Olympiads (IMO, IMC, IPhO, CPhO), SciAgent consistently attains or surpasses human gold-medalist performance, demonstrating both domain generality and reasoning adaptability. Additionally, SciAgent has been tested on the International Chemistry Olympiad (IChO) and selected problems from the Humanity's Last Exam (HLE) benchmark, further confirming the system's ability to generalize across diverse scientific domains. This work establishes SciAgent as a concrete step toward generalistic scientific intelligence-AI systems capable of coherent, cross-disciplinary reasoning at expert levels.

MemSearcher: Training LLMs to Reason, Search and Manage Memory via End-to-End Reinforcement Learning

arXiv:2511.02805v1 Announce Type: cross Abstract: Typical search agents concatenate the entire interaction history into the LLM context, preserving information integrity but producing long, noisy contexts, resulting in high computation and memory costs. In contrast, using only the current turn avoids this overhead but discards essential information. This trade-off limits the scalability of search agents. To address this challenge, we propose MemSearcher, an agent workflow that iteratively maintains a compact memory and combines the current turn with it. At each turn, MemSearcher fuses the user's question with the memory to generate reasoning traces, perform search actions, and update memory to retain only information essential for solving the task. This design stabilizes context length across multi-turn interactions, improving efficiency without sacrificing accuracy. To optimize this workflow, we introduce multi-context GRPO, an end-to-end RL framework that jointly optimize reasoning, search strategies, and memory management of MemSearcher Agents. Specifically, multi-context GRPO samples groups of trajectories under different contexts and propagates trajectory-level advantages across all conversations within them. Trained on the same dataset as Search-R1, MemSearcher achieves significant improvements over strong baselines on seven public benchmarks: +11% on Qwen2.5-3B-Instruct and +12% on Qwen2.5-7B-Instruct relative average gains. Notably, the 3B-based MemSearcher even outperforms 7B-based baselines, demonstrating that striking a balance between information integrity and efficiency yields both higher accuracy and lower computational overhead. The code and models will be publicly available at https://github.com/icip-cas/MemSearcher

Beyond Accuracy: Rethinking Hallucination and Regulatory Response in Generative AI

arXiv:2509.13345v2 Announce Type: replace-cross Abstract: Hallucination in generative AI is often treated as a technical failure to produce factually correct output. Yet this framing underrepresents the broader significance of hallucinated content in language models, which may appear fluent, persuasive, and contextually appropriate while conveying distortions that escape conventional accuracy checks. This paper critically examines how regulatory and evaluation frameworks have inherited a narrow view of hallucination, one that prioritises surface verifiability over deeper questions of meaning, influence, and impact. We propose a layered approach to understanding hallucination risks, encompassing epistemic instability, user misdirection, and social-scale effects. Drawing on interdisciplinary sources and examining instruments such as the EU AI Act and the GDPR, we show that current governance models struggle to address hallucination when it manifests as ambiguity, bias reinforcement, or normative convergence. Rather than improving factual precision alone, we argue for regulatory responses that account for languages generative nature, the asymmetries between system and user, and the shifting boundaries between information, persuasion, and harm.

Single-cell multi-omics analysis reveals cancer regulatory elements of transcriptional programs and clinical implications

Cell Death Dis. 2025 Oct 21;16(1):746. doi: 10.1038/s41419-025-08060-7.

ABSTRACT

The regulatory mechanisms governing transcriptional programs in the cancer genome remain elusive, particularly those concerning cell-type specificity. We carefully curated single-cell assay for transposase-accessible chromatin sequencing (scATAC-seq) and single-cell RNA sequencing (scRNA-seq) data from eight distinct carcinoma tissues, including breast, skin, colon, endometrium, lung, ovary, liver, and kidney. Using single-cell multi-omics analysis, we identified extensive open chromatin regions and constructed peak-gene link networks, which can reveal distinct cancer gene regulation and genetic risks. We further explored conserved epigenetic regulation across cell types within cancer and elucidated their functional implications. Moreover, we identified cell-type-associated transcription factors (TFs) that regulate key cellular functions, such as the TEAD family of TFs, which widely control cancer-related signaling pathways in tumor cells. In colon cancer, we further identified tumor-specific TFs that are more highly activated in tumor cells than in normal epithelial cells, including CEBPG, LEF1, SOX4, TCF7, and TEAD4, which are pivotal in driving malignant transcriptional programs and represent potential therapeutic targets, as corroborated by single-cell sequencing data from multiple sources and in vitro experiments. Our findings provide a comprehensive understanding of the regulatory dynamics underlying carcinomas and offer valuable insights into potential therapeutic interventions.

PMID:41120274 | PMC:PMC12541060 | DOI:10.1038/s41419-025-08060-7

Behavior Change Strategies in Digital Exercise Interventions for Adolescent Idiopathic Scoliosis: Scoping Review

Background: Adolescent idiopathic scoliosis is a common spinal deformity typically treated with exercise therapy. Despite the increasing use of digital technologies in interventions, there remains a gap in understanding how to effectively integrate behavior change techniques (BCTs) and behavior theories within these digital solutions. Objective: This review aims to identify the digital characteristics of interventions and the BCTs used, and to analyze potential theoretical mechanisms with the Theoretical Domains Framework and the capability, opportunity, motivation, and behavior model. Methods: We conducted a scoping review according to the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews) guidelines. A total of 5 databases, including PubMed, Web of Science, Embase, Cochrane Library, and CINAHL, were selected for screening eligible studies up to April 4, 2024. We included studies of any design type that involved patients with adolescent idiopathic scoliosis using digital interventions for exercise rehabilitation, including qualitative, quantitative, or mixed methods studies, and study protocols with detailed descriptions of digital interventions. Two researchers independently screened studies and extracted data into tables for descriptive analysis. The Mixed Methods Appraisal Tool was used to assess the quality of studies. Results: Out of the 3267 identified papers, 21 (0.64%) studies were included. The most frequently used technologies were videoconferencing (n=7) and instructional videos (n=5). The three most common BCT clusters were “Shaping Knowledge” (n=19), “Social Support” (n=16), and “Antecedents” (n=16). “Knowledge” was the most used mechanism of action (n=21), followed by “Skills” (n=16), “Environmental Context and Resources” (n=16), and “Social Influences” (n=16). The studies primarily addressed “Capability” and “Opportunity,” with less emphasis on “Motivation,” particularly “Automatic Motivation.” Conclusions: This review identified common digital technologies and their characteristics, analyzed potential mechanisms of behavior change in interventions, and provided recommendations for technology utilization. Future research should further evaluate the effectiveness of digital technologies while enhancing patient motivation and user experience. Trial Registration: PROSPERO CRD42024530851; https://www.crd.york.ac.uk/PROSPERO/view/CRD42024530851

Protein lipoylation in cancer: metabolic reprogramming and therapeutic potential

Cell Death Discovery, Published online: 02 September 2025; doi:10.1038/s41420-025-02718-z

Protein lipoylation in cancer: metabolic reprogramming and therapeutic potential

An organoid co-culture model for probing systemic anti-tumor immunity in lung cancer

Cell Stem Cell. 2025 Jun 6:S1934-5909(25)00191-2. doi: 10.1016/j.stem.2025.05.011. Online ahead of print.

ABSTRACT

Deciphering interactions between tumor micro- and systemic immune macroenvironments is essential for developing more effective cancer diagnosis and therapeutic strategies. Here, we established a gel-liquid interface (GLI) co-culture model of lung cancer organoids (LCOs) and paired peripheral-blood mononuclear cells (PBMCs), featuring enhanced interactions between immune cells and tumor organoids for optimized simulation of in vivo systemic anti-tumor immunity. By constructing a cohort of lung cancer patients, we demonstrated that the responses of GLI models under αPD1 treatment reflected the immunotherapy outcomes of the corresponding patients precisely. Furthermore, we dissected the various tumor immune processes mediated by PBMC-derived T cells within GLI models through functional multi-omics analyses, along with the characterization of circulating tumor-reactive T cells (GNLY+CD44+CD9+) with effector memory-like phenotypes as a potential indicator of immunotherapy efficacy. Our findings indicate that the GLI co-culture model can be used to develop diagnostic strategies for precision immunotherapies, as well as understanding the underlying mechanisms.

PMID:40513558 | DOI:10.1016/j.stem.2025.05.011

A statistical framework for multi-trait rare variant analysis in large-scale whole-genome sequencing studies

Nat Comput Sci. 2025 Feb 7. doi: 10.1038/s43588-024-00764-8. Online ahead of print.

ABSTRACT

Large-scale whole-genome sequencing (WGS) studies have improved our understanding of the contributions of coding and noncoding rare variants to complex human traits. Leveraging association effect sizes across multiple traits in WGS rare variant association analysis can improve statistical power over single-trait analysis, and also detect pleiotropic genes and regions. Existing multi-trait methods have limited ability to perform rare variant analysis of large-scale WGS data. We propose MultiSTAAR, a statistical framework and computationally scalable analytical pipeline for functionally informed multi-trait rare variant analysis in large-scale WGS studies. MultiSTAAR accounts for relatedness, population structure and correlation among phenotypes by jointly analyzing multiple traits, and further empowers rare variant association analysis by incorporating multiple functional annotations. We applied MultiSTAAR to jointly analyze three lipid traits in 61,838 multi-ethnic samples from the Trans-Omics for Precision Medicine (TOPMed) Program. We discovered and replicated new associations with lipid traits missed by single-trait analysis.

PMID:39920506 | DOI:10.1038/s43588-024-00764-8

Synthesis of portimines reveals the basis of their anti-cancer activity

Nature, Published online: 20 September 2023; doi:10.1038/s41586-023-06535-1

A scalable total synthesis of portimines enables structural reassignment of portimine B and in-depth functional evaluation of portimine A, revealing that portimine A induces translation inhibition selectively in human cancer cells and is efficacious in vivo tumour-clearance models.

Scientific discovery in the age of artificial intelligence

Nature, Published online: 02 August 2023; doi:10.1038/s41586-023-06221-2

The advances in artificial intelligence over the past decade are examined, with a discussion on how artificial intelligence systems can aid the scientific process and the central issues that remain despite advances.

Epigenetically regulated gene expression profiles decipher four molecular subtypes with prognostic and therapeutic implications in gastric cancer

Gastric cancer (GC) is one of the most common malignant tumors of the digestive tract which seriously endangers the health of human beings worldwide. Transcriptomic deregulation by epigenetic mechanisms plays ...

Predicting disease-free survival in colorectal cancer by circulating tumor DNA methylation markers

Recurrence represents a well-known poor prognostic factor for colorectal cancer (CRC) patients. This study aimed to establish an effective prognostic prediction model based on noninvasive circulating tumor DNA...
❌