Normal view
-
cs.AI, q-bio.NC updates on arXiv.org
-
MirrorMind: Empowering OmniScientist with the Expert Perspectives and Collective Knowledge of Human Scientists
arXiv:2511.16997v1 Announce Type: new Abstract: The emergence of AI Scientists has demonstrated remarkable potential in automating scientific research. However, current approaches largely conceptualize scientific discovery as a solitary optimization or search process, overlooking that knowledge production is inherently a social and historical endeavor. Human scientific insight stems from two distinct yet interconnected sources. First is the individual cognitive trajectory, where a researcher's
-
cs.AI, q-bio.NC updates on arXiv.org
-
Hierarchical Retrieval with Out-Of-Vocabulary Queries: A Case Study on SNOMED CT
arXiv:2511.16698v1 Announce Type: cross Abstract: SNOMED CT is a biomedical ontology with a hierarchical representation of large-scale concepts. Knowledge retrieval in SNOMED CT is critical for its application, but often proves challenging due to language ambiguity, synonyms, polysemies and so on. This problem is exacerbated when the queries are out-of-vocabulary (OOV), i.e., having no equivalent matchings in the ontology. In this work, we focus on the problem of hierarchical concept retrieval
Hierarchical Retrieval with Out-Of-Vocabulary Queries: A Case Study on SNOMED CT
-
cs.AI, q-bio.NC updates on arXiv.org
-
ConCISE: A Reference-Free Conciseness Evaluation Metric for LLM-Generated Answers
arXiv:2511.16846v1 Announce Type: cross Abstract: Large language models (LLMs) frequently generate responses that are lengthy and verbose, filled with redundant or unnecessary details. This diminishes clarity and user satisfaction, and it increases costs for model developers, especially with well-known proprietary models that charge based on the number of output tokens. In this paper, we introduce a novel reference-free metric for evaluating the conciseness of responses generated by LLMs. Our m
ConCISE: A Reference-Free Conciseness Evaluation Metric for LLM-Generated Answers
-
cs.AI, q-bio.NC updates on arXiv.org
-
SMILE: A Composite Lexical-Semantic Metric for Question-Answering Evaluation
arXiv:2511.17432v1 Announce Type: cross Abstract: Traditional evaluation metrics for textual and visual question answering, like ROUGE, METEOR, and Exact Match (EM), focus heavily on n-gram based lexical similarity, often missing the deeper semantic understanding needed for accurate assessment. While measures like BERTScore and MoverScore leverage contextual embeddings to address this limitation, they lack flexibility in balancing sentence-level and keyword-level semantics and ignore lexical si
SMILE: A Composite Lexical-Semantic Metric for Question-Answering Evaluation
-
cs.AI, q-bio.NC updates on arXiv.org
-
From Hypothesis to Publication: A Comprehensive Survey of AI-Driven Research Support Systems
arXiv:2503.01424v4 Announce Type: replace Abstract: Research is a fundamental process driving the advancement of human civilization, yet it demands substantial time and effort from researchers. In recent years, the rapid development of artificial intelligence (AI) technologies has inspired researchers to explore how AI can accelerate and enhance research. To monitor relevant advancements, this paper presents a systematic review of the progress in this domain. Specifically, we organize the relev
From Hypothesis to Publication: A Comprehensive Survey of AI-Driven Research Support Systems
-
cs.AI, q-bio.NC updates on arXiv.org
-
Artificial Intelligence Index Report 2025
arXiv:2504.07139v3 Announce Type: replace Abstract: Welcome to the eighth edition of the AI Index report. The 2025 Index is our most comprehensive to date and arrives at an important moment, as AI's influence across society, the economy, and global governance continues to intensify. New in this year's report are in-depth analyses of the evolving landscape of AI hardware, novel estimates of inference costs, and new analyses of AI publication and patenting trends. We also introduce fresh data on
Artificial Intelligence Index Report 2025
-
cs.AI, q-bio.NC updates on arXiv.org
-
LLM-Agent-UMF: LLM-based Agent Unified Modeling Framework for Seamless Design of Multi Active/Passive Core-Agent Architectures
arXiv:2409.11393v3 Announce Type: replace-cross Abstract: In an era where vast amounts of data are collected and processed from diverse sources, there is a growing demand for sophisticated AI systems capable of intelligently fusing and analyzing this information. To address these challenges, researchers have turned towards integrating tools into LLM-powered agents to enhance the overall information fusion process. However, the conjunction of these technologies and the proposed enhancements in s
LLM-Agent-UMF: LLM-based Agent Unified Modeling Framework for Seamless Design of Multi Active/Passive Core-Agent Architectures
-
cs.AI, q-bio.NC updates on arXiv.org
-
Genomic Next-Token Predictors are In-Context Learners
arXiv:2511.12797v2 Announce Type: replace-cross Abstract: In-context learning (ICL) -- the capacity of a model to infer and apply abstract patterns from examples provided within its input -- has been extensively studied in large language models trained for next-token prediction on human text. In fact, prior work often attributes this emergent behavior to distinctive statistical properties in human language. This raises a fundamental question: can ICL arise organically in other sequence domains
Genomic Next-Token Predictors are In-Context Learners
-
npj Digital Medicine
-
On the public dissemination and open sourcing of ultrasound resources, datasets and deep learning models
npj Digital Medicine, Published online: 24 November 2025; doi:10.1038/s41746-025-02162-4On the public dissemination and open sourcing of ultrasound resources, datasets and deep learning models
On the public dissemination and open sourcing of ultrasound resources, datasets and deep learning models
npj Digital Medicine, Published online: 24 November 2025; doi:10.1038/s41746-025-02162-4
On the public dissemination and open sourcing of ultrasound resources, datasets and deep learning models-
npj Digital Medicine
-
Multimodal analysis of whole slide images in colorectal cancer
npj Digital Medicine, Published online: 24 November 2025; doi:10.1038/s41746-025-02095-yMultimodal analysis of whole slide images in colorectal cancer
Multimodal analysis of whole slide images in colorectal cancer
npj Digital Medicine, Published online: 24 November 2025; doi:10.1038/s41746-025-02095-y
Multimodal analysis of whole slide images in colorectal cancer-
(Multiomics OR Omics) AND (Pancreatic)
-
Integrative analysis of genomic and transcriptomic data informs precancer progression in the pancreas
bioRxiv [Preprint]. 2025 Nov 4:2025.11.03.686234. doi: 10.1101/2025.11.03.686234.ABSTRACTPancreatic ductal adenocarcinoma (PDAC) arises from heterogeneous precursor lesions, including intraductal papillary mucinous neoplasms (IPMNs), but the features distinguishing indolent from progressive lesions remain unclear. We performed an integrative analysis of transcriptomic, genomic, and microenvironmental profiles of IPMNs to define multi-omic phenotypes. Using transfer learning, we projected IPMN-de
Integrative analysis of genomic and transcriptomic data informs precancer progression in the pancreas
bioRxiv [Preprint]. 2025 Nov 4:2025.11.03.686234. doi: 10.1101/2025.11.03.686234.
ABSTRACT
Pancreatic ductal adenocarcinoma (PDAC) arises from heterogeneous precursor lesions, including intraductal papillary mucinous neoplasms (IPMNs), but the features distinguishing indolent from progressive lesions remain unclear. We performed an integrative analysis of transcriptomic, genomic, and microenvironmental profiles of IPMNs to define multi-omic phenotypes. Using transfer learning, we projected IPMN-derived transcriptional programs onto spatial transcriptomic datasets from IPMNs and pancreatic intraepithelial neoplasias (PanINs). We identified two major phenotypes: one associated with cancer-associated fibroblasts and epithelial-to-mesenchymal transition, shared across IPMN, PanIN, and PDAC; and a second, glycolysis-enriched phenotype with a unique somatic mutation profile specific to IPMN. Spatial mapping further revealed grade-specific enrichment of transcriptional programs and distinct interactions with stromal and immune subtypes, underscoring the role of the precancer microenvironment in progression. These findings establish multi-omic phenotypes that unify genetic, transcriptional, and microenvironmental heterogeneity, providing a framework for distinguishing progressive from indolent precancers and a web-based public atlas for future exploration of these data and transcriptional phenotypes.
PMID:41279473 | PMC:PMC12637499 | DOI:10.1101/2025.11.03.686234
-
MRD
-
Knowledge-informed multimodal cfDNA analysis improves sensitivity and generalization in cancer detection
bioRxiv [Preprint]. 2025 Oct 21:2025.10.20.683167. doi: 10.1101/2025.10.20.683167.ABSTRACTLiquid biopsy offers a minimally invasive opportunity to detect and monitor cancers through analysis of cell-free DNA (cfDNA). However, current approaches face challenges of limited sensitivity at low tumor fractions, technical variability, and poor generalization across cohorts. Tumor-informed targeted methods offer high specificity but suffer from low sensitivity due to random sampling, tumor evolution an
Knowledge-informed multimodal cfDNA analysis improves sensitivity and generalization in cancer detection
bioRxiv [Preprint]. 2025 Oct 21:2025.10.20.683167. doi: 10.1101/2025.10.20.683167.
ABSTRACT
Liquid biopsy offers a minimally invasive opportunity to detect and monitor cancers through analysis of cell-free DNA (cfDNA). However, current approaches face challenges of limited sensitivity at low tumor fractions, technical variability, and poor generalization across cohorts. Tumor-informed targeted methods offer high specificity but suffer from low sensitivity due to random sampling, tumor evolution and adaptation (including resistance mechanisms), and other sources of heterogeneity. Conversely, tumor-naive genome-wide methods can increase sensitivity but often sacrifice specificity, particularly at low tumor fractions. We developed Fragmentomics Analysis for Tumor Evaluation with AI (Fate-AI), a multimodal framework that integrates fragmentomic and methylation-derived features from low-pass whole-genome sequencing (LPWGS) and cell-free methylated DNA immunoprecipitation and high-throughput sequencing (cfMeDIP-seq). It employs a knowledge-informed strategy to select recurrently altered genomic regions and tissue-specific methylation loci to combine the advantages of tumor-naive approaches with the specificity of tumor-informed approaches. This approach derives robust per-sample normalized features that mitigate batch effects and enhance cross-cohort reproducibility. We evaluated Fate-AI on a total of 1,219 plasma samples spanning ten cancer types and healthy controls from multiple laboratories and sequencing centers, including 432 newly profiled cases (280 with both cfMeDIP-seq and LPWGS) together with 787 samples from four independent public datasets. Fate-AI achieved superior sensitivity and specificity compared to state-of-the-art methods, detecting tumor-derived signals at fractions as low as 10-5 in experimental dilutions. Fate-AI scores correlated with disease stage and tracked longitudinal progression, anticipating relapse months before clinical progression. Furthermore, Fate-AI enabled tissue-of-origin classification, with AUCs ranging from 0.84 to 0.97 across six cancer types. Collectively, our results demonstrate that Fate-AI provides a sensitive, generalizable, and clinically actionable platform for early detection, minimal residual disease monitoring, and tissue-of-origin classification, supporting its potential as a liquid biopsy framework in precision oncology.
PMID:41278930 | PMC:PMC12633305 | DOI:10.1101/2025.10.20.683167