❌

Normal view

MirrorMind: Empowering OmniScientist with the Expert Perspectives and Collective Knowledge of Human Scientists

arXiv:2511.16997v1 Announce Type: new Abstract: The emergence of AI Scientists has demonstrated remarkable potential in automating scientific research. However, current approaches largely conceptualize scientific discovery as a solitary optimization or search process, overlooking that knowledge production is inherently a social and historical endeavor. Human scientific insight stems from two distinct yet interconnected sources. First is the individual cognitive trajectory, where a researcher's unique insight is shaped by their evolving research history and stylistic preferences; another is the collective disciplinary memory, where knowledge is sedimented into vast, interconnected networks of citations and concepts. Existing LLMs still struggle to represent these structured, high-fidelity cognitive and social contexts. To bridge this gap, we introduce MirrorMind, a hierarchical cognitive architecture that integrates dual-memory representations within a three-level framework. The Individual Level constructs high-fidelity cognitive models of individual researchers by capturing their episodic, semantic, and persona memories; the Domain Level maps collective knowledge into structured disciplinary concept graphs; and the Interdisciplinary Level that acts as an orthogonal orchestration engine. Crucially, our architecture separates memory storage from agentic execution, enabling AI scientist agents to flexibly access individual memories for unique perspectives or collective structures to reason. We evaluate MirrorMind across four comprehensive tasks, including author-level cognitive simulation, complementary reasoning, cross-disciplinary collaboration promotion, and multi-agent scientific problem solving. The results show that by integrating individual cognitive depth with collective disciplinary breadth, MirrorMind moves beyond simple fact retrieval toward structural, personalized, and insight-generating scientific reasoning.

Hierarchical Retrieval with Out-Of-Vocabulary Queries: A Case Study on SNOMED CT

arXiv:2511.16698v1 Announce Type: cross Abstract: SNOMED CT is a biomedical ontology with a hierarchical representation of large-scale concepts. Knowledge retrieval in SNOMED CT is critical for its application, but often proves challenging due to language ambiguity, synonyms, polysemies and so on. This problem is exacerbated when the queries are out-of-vocabulary (OOV), i.e., having no equivalent matchings in the ontology. In this work, we focus on the problem of hierarchical concept retrieval from SNOMED CT with OOV queries, and propose an approach based on language model-based ontology embeddings. For evaluation, we construct OOV queries annotated against SNOMED CT concepts, testing the retrieval of the most direct subsumers and their less relevant ancestors. We find that our method outperforms the baselines including SBERT and two lexical matching methods. While evaluated against SNOMED CT, the approach is generalisable and can be extended to other ontologies. We release code, tools, and evaluation datasets at https://github.com/jonathondilworth/HR-OOV.

ConCISE: A Reference-Free Conciseness Evaluation Metric for LLM-Generated Answers

arXiv:2511.16846v1 Announce Type: cross Abstract: Large language models (LLMs) frequently generate responses that are lengthy and verbose, filled with redundant or unnecessary details. This diminishes clarity and user satisfaction, and it increases costs for model developers, especially with well-known proprietary models that charge based on the number of output tokens. In this paper, we introduce a novel reference-free metric for evaluating the conciseness of responses generated by LLMs. Our method quantifies non-essential content without relying on gold standard references and calculates the average of three calculations: i) a compression ratio between the original response and an LLM abstractive summary; ii) a compression ratio between the original response and an LLM extractive summary; and iii) wordremoval compression, where an LLM removes as many non-essential words as possible from the response while preserving its meaning, with the number of tokens removed indicating the conciseness score. Experimental results demonstrate that our proposed metric identifies redundancy in LLM outputs, offering a practical tool for automated evaluation of response brevity in conversational AI systems without the need for ground truth human annotations.

SMILE: A Composite Lexical-Semantic Metric for Question-Answering Evaluation

arXiv:2511.17432v1 Announce Type: cross Abstract: Traditional evaluation metrics for textual and visual question answering, like ROUGE, METEOR, and Exact Match (EM), focus heavily on n-gram based lexical similarity, often missing the deeper semantic understanding needed for accurate assessment. While measures like BERTScore and MoverScore leverage contextual embeddings to address this limitation, they lack flexibility in balancing sentence-level and keyword-level semantics and ignore lexical similarity, which remains important. Large Language Model (LLM) based evaluators, though powerful, come with drawbacks like high costs, bias, inconsistency, and hallucinations. To address these issues, we introduce SMILE: Semantic Metric Integrating Lexical Exactness, a novel approach that combines sentence-level semantic understanding with keyword-level semantic understanding and easy keyword matching. This composite method balances lexical precision and semantic relevance, offering a comprehensive evaluation. Extensive benchmarks across text, image, and video QA tasks show SMILE is highly correlated with human judgments and computationally lightweight, bridging the gap between lexical and semantic evaluation.

From Hypothesis to Publication: A Comprehensive Survey of AI-Driven Research Support Systems

arXiv:2503.01424v4 Announce Type: replace Abstract: Research is a fundamental process driving the advancement of human civilization, yet it demands substantial time and effort from researchers. In recent years, the rapid development of artificial intelligence (AI) technologies has inspired researchers to explore how AI can accelerate and enhance research. To monitor relevant advancements, this paper presents a systematic review of the progress in this domain. Specifically, we organize the relevant studies into three main categories: hypothesis formulation, hypothesis validation, and manuscript publication. Hypothesis formulation involves knowledge synthesis and hypothesis generation. Hypothesis validation includes the verification of scientific claims, theorem proving, and experiment validation. Manuscript publication encompasses manuscript writing and the peer review process. Furthermore, we identify and discuss the current challenges faced in these areas, as well as potential future directions for research. Finally, we also offer a comprehensive overview of existing benchmarks and tools across various domains that support the integration of AI into the research process. We hope this paper serves as an introduction for beginners and fosters future research. Resources have been made publicly available at https://github.com/zkzhou126/AI-for-Research.

Artificial Intelligence Index Report 2025

arXiv:2504.07139v3 Announce Type: replace Abstract: Welcome to the eighth edition of the AI Index report. The 2025 Index is our most comprehensive to date and arrives at an important moment, as AI's influence across society, the economy, and global governance continues to intensify. New in this year's report are in-depth analyses of the evolving landscape of AI hardware, novel estimates of inference costs, and new analyses of AI publication and patenting trends. We also introduce fresh data on corporate adoption of responsible AI practices, along with expanded coverage of AI's growing role in science and medicine. Since its founding in 2017 as an offshoot of the One Hundred Year Study of Artificial Intelligence, the AI Index has been committed to equipping policymakers, journalists, executives, researchers, and the public with accurate, rigorously validated, and globally sourced data. Our mission has always been to help these stakeholders make better-informed decisions about the development and deployment of AI. In a world where AI is discussed everywhere - from boardrooms to kitchen tables - this mission has never been more essential. The AI Index continues to lead in tracking and interpreting the most critical trends shaping the field - from the shifting geopolitical landscape and the rapid evolution of underlying technologies, to AI's expanding role in business, policymaking, and public life. Longitudinal tracking remains at the heart of our mission. In a domain advancing at breakneck speed, the Index provides essential context - helping us understand where AI stands today, how it got here, and where it may be headed next. Recognized globally as one of the most authoritative resources on artificial intelligence, the AI Index has been cited in major media outlets such as The New York Times, Bloomberg, and The Guardian; referenced in hundreds of academic papers; and used by policymakers and government agencies around the world.

LLM-Agent-UMF: LLM-based Agent Unified Modeling Framework for Seamless Design of Multi Active/Passive Core-Agent Architectures

arXiv:2409.11393v3 Announce Type: replace-cross Abstract: In an era where vast amounts of data are collected and processed from diverse sources, there is a growing demand for sophisticated AI systems capable of intelligently fusing and analyzing this information. To address these challenges, researchers have turned towards integrating tools into LLM-powered agents to enhance the overall information fusion process. However, the conjunction of these technologies and the proposed enhancements in several state-of-the-art works followed a non-unified software architecture, resulting in a lack of modularity and terminological inconsistencies among researchers. To address these issues, we propose a novel LLM-based Agent Unified Modeling Framework (LLM-Agent-UMF) that establishes a clear foundation for agent development from both functional and software architectural perspectives, developed and evaluated using the Architecture Tradeoff and Risk Analysis Framework (ATRAF). Our framework clearly distinguishes between the different components of an LLM-based agent, setting LLMs and tools apart from a new element, the core-agent, which plays the role of central coordinator. This pivotal entity comprises five modules: planning, memory, profile, action, and security -- the latter often neglected in previous works. By classifying core-agents into passive and active types based on their authoritative natures, we propose various multi-core agent architectures that combine unique characteristics of distinctive agents to tackle complex tasks more efficiently. We evaluate our framework by applying it to thirteen state-of-the-art agents, thereby demonstrating its alignment with their functionalities and clarifying overlooked architectural aspects. Moreover, we thoroughly assess five architecture variants of our framework by designing new agent architectures that combine characteristics of state-of-the-art agents to address specific goals. ...

Genomic Next-Token Predictors are In-Context Learners

arXiv:2511.12797v2 Announce Type: replace-cross Abstract: In-context learning (ICL) -- the capacity of a model to infer and apply abstract patterns from examples provided within its input -- has been extensively studied in large language models trained for next-token prediction on human text. In fact, prior work often attributes this emergent behavior to distinctive statistical properties in human language. This raises a fundamental question: can ICL arise organically in other sequence domains purely through large-scale predictive training? To explore this, we turn to genomic sequences, an alternative symbolic domain rich in statistical structure. Specifically, we study the Evo2 genomic model, trained predominantly on next-nucleotide (A/T/C/G) prediction, at a scale comparable to mid-sized LLMs. We develop a controlled experimental framework comprising symbolic reasoning tasks instantiated in both linguistic and genomic forms, enabling direct comparison of ICL across genomic and linguistic models. Our results show that genomic models, like their linguistic counterparts, exhibit log-linear gains in pattern induction as the number of in-context demonstrations increases. To the best of our knowledge, this is the first evidence of organically emergent ICL in genomic sequences, supporting the hypothesis that ICL arises as a consequence of large-scale predictive modeling over rich data. These findings extend emergent meta-learning beyond language, pointing toward a unified, modality-agnostic view of in-context learning.

On the public dissemination and open sourcing of ultrasound resources, datasets and deep learning models

npj Digital Medicine, Published online: 24 November 2025; doi:10.1038/s41746-025-02162-4

On the public dissemination and open sourcing of ultrasound resources, datasets and deep learning models

Multimodal analysis of whole slide images in colorectal cancer

npj Digital Medicine, Published online: 24 November 2025; doi:10.1038/s41746-025-02095-y

Multimodal analysis of whole slide images in colorectal cancer

Integrative analysis of genomic and transcriptomic data informs precancer progression in the pancreas

bioRxiv [Preprint]. 2025 Nov 4:2025.11.03.686234. doi: 10.1101/2025.11.03.686234.

ABSTRACT

Pancreatic ductal adenocarcinoma (PDAC) arises from heterogeneous precursor lesions, including intraductal papillary mucinous neoplasms (IPMNs), but the features distinguishing indolent from progressive lesions remain unclear. We performed an integrative analysis of transcriptomic, genomic, and microenvironmental profiles of IPMNs to define multi-omic phenotypes. Using transfer learning, we projected IPMN-derived transcriptional programs onto spatial transcriptomic datasets from IPMNs and pancreatic intraepithelial neoplasias (PanINs). We identified two major phenotypes: one associated with cancer-associated fibroblasts and epithelial-to-mesenchymal transition, shared across IPMN, PanIN, and PDAC; and a second, glycolysis-enriched phenotype with a unique somatic mutation profile specific to IPMN. Spatial mapping further revealed grade-specific enrichment of transcriptional programs and distinct interactions with stromal and immune subtypes, underscoring the role of the precancer microenvironment in progression. These findings establish multi-omic phenotypes that unify genetic, transcriptional, and microenvironmental heterogeneity, providing a framework for distinguishing progressive from indolent precancers and a web-based public atlas for future exploration of these data and transcriptional phenotypes.

PMID:41279473 | PMC:PMC12637499 | DOI:10.1101/2025.11.03.686234

Knowledge-informed multimodal cfDNA analysis improves sensitivity and generalization in cancer detection

bioRxiv [Preprint]. 2025 Oct 21:2025.10.20.683167. doi: 10.1101/2025.10.20.683167.

ABSTRACT

Liquid biopsy offers a minimally invasive opportunity to detect and monitor cancers through analysis of cell-free DNA (cfDNA). However, current approaches face challenges of limited sensitivity at low tumor fractions, technical variability, and poor generalization across cohorts. Tumor-informed targeted methods offer high specificity but suffer from low sensitivity due to random sampling, tumor evolution and adaptation (including resistance mechanisms), and other sources of heterogeneity. Conversely, tumor-naive genome-wide methods can increase sensitivity but often sacrifice specificity, particularly at low tumor fractions. We developed Fragmentomics Analysis for Tumor Evaluation with AI (Fate-AI), a multimodal framework that integrates fragmentomic and methylation-derived features from low-pass whole-genome sequencing (LPWGS) and cell-free methylated DNA immunoprecipitation and high-throughput sequencing (cfMeDIP-seq). It employs a knowledge-informed strategy to select recurrently altered genomic regions and tissue-specific methylation loci to combine the advantages of tumor-naive approaches with the specificity of tumor-informed approaches. This approach derives robust per-sample normalized features that mitigate batch effects and enhance cross-cohort reproducibility. We evaluated Fate-AI on a total of 1,219 plasma samples spanning ten cancer types and healthy controls from multiple laboratories and sequencing centers, including 432 newly profiled cases (280 with both cfMeDIP-seq and LPWGS) together with 787 samples from four independent public datasets. Fate-AI achieved superior sensitivity and specificity compared to state-of-the-art methods, detecting tumor-derived signals at fractions as low as 10-5 in experimental dilutions. Fate-AI scores correlated with disease stage and tracked longitudinal progression, anticipating relapse months before clinical progression. Furthermore, Fate-AI enabled tissue-of-origin classification, with AUCs ranging from 0.84 to 0.97 across six cancer types. Collectively, our results demonstrate that Fate-AI provides a sensitive, generalizable, and clinically actionable platform for early detection, minimal residual disease monitoring, and tissue-of-origin classification, supporting its potential as a liquid biopsy framework in precision oncology.

PMID:41278930 | PMC:PMC12633305 | DOI:10.1101/2025.10.20.683167

Health care Experiences of Educated Young Adults With Blindness in the Digital Age: Qualitative Study

Background: The rapid advancement of digital health technologies (DHTs) offers substantial potential for improving healthcare access, yet it simultaneously risks exacerbating existing inequities for marginalized populations. Previous research on the digital divide has often treated individuals with blindness as a homogenous group, primarily focusing on barriers related to digital access and skills. However, less is known about the nuanced experiences of specific subgroups, such as educated and digitally literate young adults. This study focuses on this demographic to understand how their advanced digital capabilities interact with systemic and infrastructural barriers in healthcare. Objective: This qualitative study aimed to explore the lived healthcare experiences of educated young adults with blindness in China, specifically identifying how DHTs simultaneously contribute to their empowerment and exclusion. Methods: Eligible participants were educated young adults with blindness in China (aged 18-30 years, Mandarin speakers, smartphone users, and holding or pursuing higher education). A total of 12 semi-structured interviews were conducted in Mandarin during September 2024. All interviews were audio-recorded and transcribed verbatim. An inductive thematic analysis was employed to interpret the data and identify key themes. Results: Participants’ experiences highlighted an “empowered but excluded” dynamic. Seven key themes emerged, categorized into empowerment and exclusion. Empowerment themes included: (1) digital platforms empowering self-management and healthcare access, where DHTs enabled independent appointment booking and access to comprehensive health information; and (2) digital platforms empowering for finding medical visit companions, facilitating the discovery of companions for physical and emotional support. Exclusion themes comprised: (3) inaccessible online appointment systems, due to non-inclusive designs; (4) inaccessible healthcare environments and information formats, stemming from non-accessible self-service machines and written materials; (5) lack of provider competencies in respecting patient autonomy, as providers often assumed digital incompetence; (6) data privacy and security concerns, heightened by increased digitalization and reliance on assistive tools; and (7) challenges related to the quality and consistency of online companion support, highlighting the limitations of platform-based assistance. Conclusions: Our findings reveal an “empowered but excluded” dynamic: the potential for digital empowerment and enhanced independence is often curtailed by systematic barriers. Addressing this necessitates a multifaceted approach: enhancing technological accessibility through robust standards adherence and inclusive co-design processes; improving healthcare provider competencies in patient-centered care via targeted training; and empowering educated young blind adults by building their capacity for self-determination to achieve equitable healthcare access.

Considerations for Patient Privacy of Large Language Models in Health Care: Scoping Review

Background: The application of large language models (LLMs) in health care holds significant potential for enhancing patient care and advancing medical research. However, the protection of patient privacy remains a critical issue, especially when handling patient health information (PHI). Objective: This scoping review aims to evaluate the adequacy of current approaches and identify areas in need of improvement to ensure robust patient privacy protection in the existing studies about PHI-LLMs within the health care domain. Methods: A search of the literature published from January 1, 2022, to July 20, 2025, was performed on July 20, 2025, using 2 databases (PubMed and Embase). This scoping review focused on the following three research questions: (1) What studies on the development and application of LLMs using PHI currently exist within the health care domain? (2) What patient privacy considerations are addressed in existing PHI-LLMs research, and are these measures sufficient? (3) How can future research on the development and application of LLMs using PHI better protect patient privacy? Studies were included if they focused on the development and application of LLMs within health care using PHI, encompassing activities such as model construction, fine-tuning, optimization, testing, and performance comparison. Eligible literature comprised original research articles written in English. Conversely, studies were excluded if they used publicly available datasets, under the assumption that such data have been adequately deidentified. Additionally, non-English publications, reviews, abstracts, incomplete reports, and preprints were excluded from the review due to the lack of rigorous peer review. Results: This study systematically identified 9823 studies on PHI-LLM and included 464 studies published between 2022 and 2025. Among the 464 studies, (1) a small number of studies neglected ethical review (n=45, 9.7%) and patient informed consent (n=148, 31.9%) during the research process, (2) more than a third of the studies (n=178, 38.4%) failed to report whether to implement effective measures to protect PHI, and (3) there was a significant lack of transparency and comprehensive detail in anonymization and deidentification methods. Conclusions: We propose comprehensive recommendations across 3 phases—study design, implementation, and reporting—to strengthen patient privacy protection and transparency in PHI-LLM. This study emphasizes the urgent need for the development of stricter regulatory frameworks and the adoption of advanced privacy protection technologies to effectively safeguard PHI. It is anticipated that future applications of LLMs in the health care field will achieve a balance between innovation and robust patient privacy protection, thereby enhancing ethical standards and scientific credibility.

Clinical validation of a three-marker methylation panel to detect CIN3+ in vaginal self-samples in the Dutch population-based screening programme

The use of vaginal self-sampling for cervical cancer screening is promising and increasing. However, triage cytology cannot be performed on vaginal self-sampling material after a high-risk human papilloma viru...

Beyond GeneGPT: A Multi-Agent Architecture with Open-Source LLMs for Enhanced Genomic Question Answering

arXiv:2511.15061v1 Announce Type: new Abstract: Genomic question answering often requires complex reasoning and integration across diverse biomedical sources. GeneGPT addressed this challenge by combining domain-specific APIs with OpenAI's code-davinci-002 large language model to enable natural language interaction with genomic databases. However, its reliance on a proprietary model limits scalability, increases operational costs, and raises concerns about data privacy and generalization. In this work, we revisit and reproduce GeneGPT in a pilot study using open source models, including Llama 3.1, Qwen2.5, and Qwen2.5 Coder, within a monolithic architecture; this allows us to identify the limitations of this approach. Building on this foundation, we then develop OpenBioLLM, a modular multi-agent framework that extends GeneGPT by introducing agent specialization for tool routing, query generation, and response validation. This enables coordinated reasoning and role-based task execution. OpenBioLLM matches or outperforms GeneGPT on over 90% of the benchmark tasks, achieving average scores of 0.849 on Gene-Turing and 0.830 on GeneHop, while using smaller open-source models without additional fine-tuning or tool-specific pretraining. OpenBioLLM's modular multi-agent design reduces latency by 40-50% across benchmark tasks, significantly improving efficiency without compromising model capability. The results of our comprehensive evaluation highlight the potential of open-source multi-agent systems for genomic question answering. Code and resources are available at https://github.com/ielab/OpenBioLLM.

Exploring the use of AI authors and reviewers at Agents4Science

arXiv:2511.15534v1 Announce Type: new Abstract: There is growing interest in using AI agents for scientific research, yet fundamental questions remain about their capabilities as scientists and reviewers. To explore these questions, we organized Agents4Science, the first conference in which AI agents serve as both primary authors and reviewers, with humans as co-authors and co-reviewers. Here, we discuss the key learnings from the conference and their implications for human-AI collaboration in science.

Uncertainty Makes It Stable: Curiosity-Driven Quantized Mixture-of-Experts

arXiv:2511.11743v2 Announce Type: replace-cross Abstract: Deploying deep neural networks on resource-constrained devices faces two critical challenges: maintaining accuracy under aggressive quantization while ensuring predictable inference latency. We present a curiosity-driven quantized Mixture-of-Experts framework that addresses both through Bayesian epistemic uncertainty-based routing across heterogeneous experts (BitNet ternary, 1-16 bit BitLinear, post-training quantization). Evaluated on audio classification benchmarks (ESC-50, Quinn, UrbanSound8K), our 4-bit quantization maintains 99.9 percent of 16-bit accuracy (0.858 vs 0.859 F1) with 4x compression and 41 percent energy savings versus 8-bit. Crucially, curiosity-driven routing reduces MoE latency variance by 82 percent (p = 0.008, Levene's test) from 230 ms to 29 ms standard deviation, enabling stable inference for battery-constrained devices. Statistical analysis confirms 4-bit/8-bit achieve practical equivalence with full precision (p > 0.05), while MoE architectures introduce 11 percent latency overhead (p

Pan-cancer prevalence, risk, and clinical and demographic characteristics of Lynch Syndrome-associated variants in BioBank Japan

Commun Med (Lond). 2025 Nov 13. doi: 10.1038/s43856-025-01231-9. Online ahead of print.

ABSTRACT

BACKGROUND: Although germline testing for DNA mismatch repair (MMR) genes is routinely performed, clinical guidelines highlight evidence gaps due to limited populations and biases. We examined germline pathogenic variants of MMR genes (MLH1, MSH2, MSH6, and PMS2) in 112,927 unselected individuals from BioBank Japan.

METHODS: We analyzed 74,085 cancer patients with 23 cancer types and 38,842 controls matched by sex, age, and hospital area from BioBank Japan, collected between April 2003 and March 2018. Germline pathogenic variants in the coding regions and 2 bp flanking intronic sequences of MMR genes were identified using a multiplex PCR-based target sequencing method. We examined associations with cancer types and demographic characterization of the pathogenic variants, comparing findings to existing clinical guidelines.

RESULTS: Here we show 228 pathogenic variants identified in MMR genes, with pathogenic MSH6 variants most frequently observed in endometrial cancer and 12 other significant associations. Twelve other significant associations are noted across a broad range of odds ratios, whereas pancreatic cancer exhibits no such association. Pathogenic variant carriers are diagnosed up to 12.4 years earlier than non-carriers, and colorectal and gastric cancers are diagnosed up to 16.4 years later than indicated by the guidelines. Higher carrier frequencies are observed in patients with both colorectal and endometrial cancers (24.8%) and in those with endometrial cancer and a family history of endometrial (26.0%) or colorectal (16.1%) cancers.

CONCLUSIONS: This study provides critical insights for clinical guidelines on the associations between cancer types, age at diagnosis, and carrier frequency.

PMID:41258140 | DOI:10.1038/s43856-025-01231-9

❌