❌

Normal view

MirrorMind: Empowering OmniScientist with the Expert Perspectives and Collective Knowledge of Human Scientists

arXiv:2511.16997v1 Announce Type: new Abstract: The emergence of AI Scientists has demonstrated remarkable potential in automating scientific research. However, current approaches largely conceptualize scientific discovery as a solitary optimization or search process, overlooking that knowledge production is inherently a social and historical endeavor. Human scientific insight stems from two distinct yet interconnected sources. First is the individual cognitive trajectory, where a researcher's unique insight is shaped by their evolving research history and stylistic preferences; another is the collective disciplinary memory, where knowledge is sedimented into vast, interconnected networks of citations and concepts. Existing LLMs still struggle to represent these structured, high-fidelity cognitive and social contexts. To bridge this gap, we introduce MirrorMind, a hierarchical cognitive architecture that integrates dual-memory representations within a three-level framework. The Individual Level constructs high-fidelity cognitive models of individual researchers by capturing their episodic, semantic, and persona memories; the Domain Level maps collective knowledge into structured disciplinary concept graphs; and the Interdisciplinary Level that acts as an orthogonal orchestration engine. Crucially, our architecture separates memory storage from agentic execution, enabling AI scientist agents to flexibly access individual memories for unique perspectives or collective structures to reason. We evaluate MirrorMind across four comprehensive tasks, including author-level cognitive simulation, complementary reasoning, cross-disciplinary collaboration promotion, and multi-agent scientific problem solving. The results show that by integrating individual cognitive depth with collective disciplinary breadth, MirrorMind moves beyond simple fact retrieval toward structural, personalized, and insight-generating scientific reasoning.

Hierarchical Retrieval with Out-Of-Vocabulary Queries: A Case Study on SNOMED CT

arXiv:2511.16698v1 Announce Type: cross Abstract: SNOMED CT is a biomedical ontology with a hierarchical representation of large-scale concepts. Knowledge retrieval in SNOMED CT is critical for its application, but often proves challenging due to language ambiguity, synonyms, polysemies and so on. This problem is exacerbated when the queries are out-of-vocabulary (OOV), i.e., having no equivalent matchings in the ontology. In this work, we focus on the problem of hierarchical concept retrieval from SNOMED CT with OOV queries, and propose an approach based on language model-based ontology embeddings. For evaluation, we construct OOV queries annotated against SNOMED CT concepts, testing the retrieval of the most direct subsumers and their less relevant ancestors. We find that our method outperforms the baselines including SBERT and two lexical matching methods. While evaluated against SNOMED CT, the approach is generalisable and can be extended to other ontologies. We release code, tools, and evaluation datasets at https://github.com/jonathondilworth/HR-OOV.

ConCISE: A Reference-Free Conciseness Evaluation Metric for LLM-Generated Answers

arXiv:2511.16846v1 Announce Type: cross Abstract: Large language models (LLMs) frequently generate responses that are lengthy and verbose, filled with redundant or unnecessary details. This diminishes clarity and user satisfaction, and it increases costs for model developers, especially with well-known proprietary models that charge based on the number of output tokens. In this paper, we introduce a novel reference-free metric for evaluating the conciseness of responses generated by LLMs. Our method quantifies non-essential content without relying on gold standard references and calculates the average of three calculations: i) a compression ratio between the original response and an LLM abstractive summary; ii) a compression ratio between the original response and an LLM extractive summary; and iii) wordremoval compression, where an LLM removes as many non-essential words as possible from the response while preserving its meaning, with the number of tokens removed indicating the conciseness score. Experimental results demonstrate that our proposed metric identifies redundancy in LLM outputs, offering a practical tool for automated evaluation of response brevity in conversational AI systems without the need for ground truth human annotations.

SMILE: A Composite Lexical-Semantic Metric for Question-Answering Evaluation

arXiv:2511.17432v1 Announce Type: cross Abstract: Traditional evaluation metrics for textual and visual question answering, like ROUGE, METEOR, and Exact Match (EM), focus heavily on n-gram based lexical similarity, often missing the deeper semantic understanding needed for accurate assessment. While measures like BERTScore and MoverScore leverage contextual embeddings to address this limitation, they lack flexibility in balancing sentence-level and keyword-level semantics and ignore lexical similarity, which remains important. Large Language Model (LLM) based evaluators, though powerful, come with drawbacks like high costs, bias, inconsistency, and hallucinations. To address these issues, we introduce SMILE: Semantic Metric Integrating Lexical Exactness, a novel approach that combines sentence-level semantic understanding with keyword-level semantic understanding and easy keyword matching. This composite method balances lexical precision and semantic relevance, offering a comprehensive evaluation. Extensive benchmarks across text, image, and video QA tasks show SMILE is highly correlated with human judgments and computationally lightweight, bridging the gap between lexical and semantic evaluation.

From Hypothesis to Publication: A Comprehensive Survey of AI-Driven Research Support Systems

arXiv:2503.01424v4 Announce Type: replace Abstract: Research is a fundamental process driving the advancement of human civilization, yet it demands substantial time and effort from researchers. In recent years, the rapid development of artificial intelligence (AI) technologies has inspired researchers to explore how AI can accelerate and enhance research. To monitor relevant advancements, this paper presents a systematic review of the progress in this domain. Specifically, we organize the relevant studies into three main categories: hypothesis formulation, hypothesis validation, and manuscript publication. Hypothesis formulation involves knowledge synthesis and hypothesis generation. Hypothesis validation includes the verification of scientific claims, theorem proving, and experiment validation. Manuscript publication encompasses manuscript writing and the peer review process. Furthermore, we identify and discuss the current challenges faced in these areas, as well as potential future directions for research. Finally, we also offer a comprehensive overview of existing benchmarks and tools across various domains that support the integration of AI into the research process. We hope this paper serves as an introduction for beginners and fosters future research. Resources have been made publicly available at https://github.com/zkzhou126/AI-for-Research.

Artificial Intelligence Index Report 2025

arXiv:2504.07139v3 Announce Type: replace Abstract: Welcome to the eighth edition of the AI Index report. The 2025 Index is our most comprehensive to date and arrives at an important moment, as AI's influence across society, the economy, and global governance continues to intensify. New in this year's report are in-depth analyses of the evolving landscape of AI hardware, novel estimates of inference costs, and new analyses of AI publication and patenting trends. We also introduce fresh data on corporate adoption of responsible AI practices, along with expanded coverage of AI's growing role in science and medicine. Since its founding in 2017 as an offshoot of the One Hundred Year Study of Artificial Intelligence, the AI Index has been committed to equipping policymakers, journalists, executives, researchers, and the public with accurate, rigorously validated, and globally sourced data. Our mission has always been to help these stakeholders make better-informed decisions about the development and deployment of AI. In a world where AI is discussed everywhere - from boardrooms to kitchen tables - this mission has never been more essential. The AI Index continues to lead in tracking and interpreting the most critical trends shaping the field - from the shifting geopolitical landscape and the rapid evolution of underlying technologies, to AI's expanding role in business, policymaking, and public life. Longitudinal tracking remains at the heart of our mission. In a domain advancing at breakneck speed, the Index provides essential context - helping us understand where AI stands today, how it got here, and where it may be headed next. Recognized globally as one of the most authoritative resources on artificial intelligence, the AI Index has been cited in major media outlets such as The New York Times, Bloomberg, and The Guardian; referenced in hundreds of academic papers; and used by policymakers and government agencies around the world.

LLM-Agent-UMF: LLM-based Agent Unified Modeling Framework for Seamless Design of Multi Active/Passive Core-Agent Architectures

arXiv:2409.11393v3 Announce Type: replace-cross Abstract: In an era where vast amounts of data are collected and processed from diverse sources, there is a growing demand for sophisticated AI systems capable of intelligently fusing and analyzing this information. To address these challenges, researchers have turned towards integrating tools into LLM-powered agents to enhance the overall information fusion process. However, the conjunction of these technologies and the proposed enhancements in several state-of-the-art works followed a non-unified software architecture, resulting in a lack of modularity and terminological inconsistencies among researchers. To address these issues, we propose a novel LLM-based Agent Unified Modeling Framework (LLM-Agent-UMF) that establishes a clear foundation for agent development from both functional and software architectural perspectives, developed and evaluated using the Architecture Tradeoff and Risk Analysis Framework (ATRAF). Our framework clearly distinguishes between the different components of an LLM-based agent, setting LLMs and tools apart from a new element, the core-agent, which plays the role of central coordinator. This pivotal entity comprises five modules: planning, memory, profile, action, and security -- the latter often neglected in previous works. By classifying core-agents into passive and active types based on their authoritative natures, we propose various multi-core agent architectures that combine unique characteristics of distinctive agents to tackle complex tasks more efficiently. We evaluate our framework by applying it to thirteen state-of-the-art agents, thereby demonstrating its alignment with their functionalities and clarifying overlooked architectural aspects. Moreover, we thoroughly assess five architecture variants of our framework by designing new agent architectures that combine characteristics of state-of-the-art agents to address specific goals. ...

Genomic Next-Token Predictors are In-Context Learners

arXiv:2511.12797v2 Announce Type: replace-cross Abstract: In-context learning (ICL) -- the capacity of a model to infer and apply abstract patterns from examples provided within its input -- has been extensively studied in large language models trained for next-token prediction on human text. In fact, prior work often attributes this emergent behavior to distinctive statistical properties in human language. This raises a fundamental question: can ICL arise organically in other sequence domains purely through large-scale predictive training? To explore this, we turn to genomic sequences, an alternative symbolic domain rich in statistical structure. Specifically, we study the Evo2 genomic model, trained predominantly on next-nucleotide (A/T/C/G) prediction, at a scale comparable to mid-sized LLMs. We develop a controlled experimental framework comprising symbolic reasoning tasks instantiated in both linguistic and genomic forms, enabling direct comparison of ICL across genomic and linguistic models. Our results show that genomic models, like their linguistic counterparts, exhibit log-linear gains in pattern induction as the number of in-context demonstrations increases. To the best of our knowledge, this is the first evidence of organically emergent ICL in genomic sequences, supporting the hypothesis that ICL arises as a consequence of large-scale predictive modeling over rich data. These findings extend emergent meta-learning beyond language, pointing toward a unified, modality-agnostic view of in-context learning.

On the public dissemination and open sourcing of ultrasound resources, datasets and deep learning models

npj Digital Medicine, Published online: 24 November 2025; doi:10.1038/s41746-025-02162-4

On the public dissemination and open sourcing of ultrasound resources, datasets and deep learning models

Multimodal analysis of whole slide images in colorectal cancer

npj Digital Medicine, Published online: 24 November 2025; doi:10.1038/s41746-025-02095-y

Multimodal analysis of whole slide images in colorectal cancer

Integrative analysis of genomic and transcriptomic data informs precancer progression in the pancreas

bioRxiv [Preprint]. 2025 Nov 4:2025.11.03.686234. doi: 10.1101/2025.11.03.686234.

ABSTRACT

Pancreatic ductal adenocarcinoma (PDAC) arises from heterogeneous precursor lesions, including intraductal papillary mucinous neoplasms (IPMNs), but the features distinguishing indolent from progressive lesions remain unclear. We performed an integrative analysis of transcriptomic, genomic, and microenvironmental profiles of IPMNs to define multi-omic phenotypes. Using transfer learning, we projected IPMN-derived transcriptional programs onto spatial transcriptomic datasets from IPMNs and pancreatic intraepithelial neoplasias (PanINs). We identified two major phenotypes: one associated with cancer-associated fibroblasts and epithelial-to-mesenchymal transition, shared across IPMN, PanIN, and PDAC; and a second, glycolysis-enriched phenotype with a unique somatic mutation profile specific to IPMN. Spatial mapping further revealed grade-specific enrichment of transcriptional programs and distinct interactions with stromal and immune subtypes, underscoring the role of the precancer microenvironment in progression. These findings establish multi-omic phenotypes that unify genetic, transcriptional, and microenvironmental heterogeneity, providing a framework for distinguishing progressive from indolent precancers and a web-based public atlas for future exploration of these data and transcriptional phenotypes.

PMID:41279473 | PMC:PMC12637499 | DOI:10.1101/2025.11.03.686234

Knowledge-informed multimodal cfDNA analysis improves sensitivity and generalization in cancer detection

bioRxiv [Preprint]. 2025 Oct 21:2025.10.20.683167. doi: 10.1101/2025.10.20.683167.

ABSTRACT

Liquid biopsy offers a minimally invasive opportunity to detect and monitor cancers through analysis of cell-free DNA (cfDNA). However, current approaches face challenges of limited sensitivity at low tumor fractions, technical variability, and poor generalization across cohorts. Tumor-informed targeted methods offer high specificity but suffer from low sensitivity due to random sampling, tumor evolution and adaptation (including resistance mechanisms), and other sources of heterogeneity. Conversely, tumor-naive genome-wide methods can increase sensitivity but often sacrifice specificity, particularly at low tumor fractions. We developed Fragmentomics Analysis for Tumor Evaluation with AI (Fate-AI), a multimodal framework that integrates fragmentomic and methylation-derived features from low-pass whole-genome sequencing (LPWGS) and cell-free methylated DNA immunoprecipitation and high-throughput sequencing (cfMeDIP-seq). It employs a knowledge-informed strategy to select recurrently altered genomic regions and tissue-specific methylation loci to combine the advantages of tumor-naive approaches with the specificity of tumor-informed approaches. This approach derives robust per-sample normalized features that mitigate batch effects and enhance cross-cohort reproducibility. We evaluated Fate-AI on a total of 1,219 plasma samples spanning ten cancer types and healthy controls from multiple laboratories and sequencing centers, including 432 newly profiled cases (280 with both cfMeDIP-seq and LPWGS) together with 787 samples from four independent public datasets. Fate-AI achieved superior sensitivity and specificity compared to state-of-the-art methods, detecting tumor-derived signals at fractions as low as 10-5 in experimental dilutions. Fate-AI scores correlated with disease stage and tracked longitudinal progression, anticipating relapse months before clinical progression. Furthermore, Fate-AI enabled tissue-of-origin classification, with AUCs ranging from 0.84 to 0.97 across six cancer types. Collectively, our results demonstrate that Fate-AI provides a sensitive, generalizable, and clinically actionable platform for early detection, minimal residual disease monitoring, and tissue-of-origin classification, supporting its potential as a liquid biopsy framework in precision oncology.

PMID:41278930 | PMC:PMC12633305 | DOI:10.1101/2025.10.20.683167

A new AI benchmark tests whether chatbots protect human well-being

25 November 2025 at 00:15
Most AI benchmarks measure intelligence and instruction-following rather than psychological safety. Humane Bench evaluates models based on core principles of human flourishing, prioritizing well-being, and respecting user attention.
❌