❌

Normal view

ConCISE: A Reference-Free Conciseness Evaluation Metric for LLM-Generated Answers

arXiv:2511.16846v1 Announce Type: cross Abstract: Large language models (LLMs) frequently generate responses that are lengthy and verbose, filled with redundant or unnecessary details. This diminishes clarity and user satisfaction, and it increases costs for model developers, especially with well-known proprietary models that charge based on the number of output tokens. In this paper, we introduce a novel reference-free metric for evaluating the conciseness of responses generated by LLMs. Our method quantifies non-essential content without relying on gold standard references and calculates the average of three calculations: i) a compression ratio between the original response and an LLM abstractive summary; ii) a compression ratio between the original response and an LLM extractive summary; and iii) wordremoval compression, where an LLM removes as many non-essential words as possible from the response while preserving its meaning, with the number of tokens removed indicating the conciseness score. Experimental results demonstrate that our proposed metric identifies redundancy in LLM outputs, offering a practical tool for automated evaluation of response brevity in conversational AI systems without the need for ground truth human annotations.

From Hypothesis to Publication: A Comprehensive Survey of AI-Driven Research Support Systems

arXiv:2503.01424v4 Announce Type: replace Abstract: Research is a fundamental process driving the advancement of human civilization, yet it demands substantial time and effort from researchers. In recent years, the rapid development of artificial intelligence (AI) technologies has inspired researchers to explore how AI can accelerate and enhance research. To monitor relevant advancements, this paper presents a systematic review of the progress in this domain. Specifically, we organize the relevant studies into three main categories: hypothesis formulation, hypothesis validation, and manuscript publication. Hypothesis formulation involves knowledge synthesis and hypothesis generation. Hypothesis validation includes the verification of scientific claims, theorem proving, and experiment validation. Manuscript publication encompasses manuscript writing and the peer review process. Furthermore, we identify and discuss the current challenges faced in these areas, as well as potential future directions for research. Finally, we also offer a comprehensive overview of existing benchmarks and tools across various domains that support the integration of AI into the research process. We hope this paper serves as an introduction for beginners and fosters future research. Resources have been made publicly available at https://github.com/zkzhou126/AI-for-Research.

Artificial Intelligence Index Report 2025

arXiv:2504.07139v3 Announce Type: replace Abstract: Welcome to the eighth edition of the AI Index report. The 2025 Index is our most comprehensive to date and arrives at an important moment, as AI's influence across society, the economy, and global governance continues to intensify. New in this year's report are in-depth analyses of the evolving landscape of AI hardware, novel estimates of inference costs, and new analyses of AI publication and patenting trends. We also introduce fresh data on corporate adoption of responsible AI practices, along with expanded coverage of AI's growing role in science and medicine. Since its founding in 2017 as an offshoot of the One Hundred Year Study of Artificial Intelligence, the AI Index has been committed to equipping policymakers, journalists, executives, researchers, and the public with accurate, rigorously validated, and globally sourced data. Our mission has always been to help these stakeholders make better-informed decisions about the development and deployment of AI. In a world where AI is discussed everywhere - from boardrooms to kitchen tables - this mission has never been more essential. The AI Index continues to lead in tracking and interpreting the most critical trends shaping the field - from the shifting geopolitical landscape and the rapid evolution of underlying technologies, to AI's expanding role in business, policymaking, and public life. Longitudinal tracking remains at the heart of our mission. In a domain advancing at breakneck speed, the Index provides essential context - helping us understand where AI stands today, how it got here, and where it may be headed next. Recognized globally as one of the most authoritative resources on artificial intelligence, the AI Index has been cited in major media outlets such as The New York Times, Bloomberg, and The Guardian; referenced in hundreds of academic papers; and used by policymakers and government agencies around the world.

Genomic Next-Token Predictors are In-Context Learners

arXiv:2511.12797v2 Announce Type: replace-cross Abstract: In-context learning (ICL) -- the capacity of a model to infer and apply abstract patterns from examples provided within its input -- has been extensively studied in large language models trained for next-token prediction on human text. In fact, prior work often attributes this emergent behavior to distinctive statistical properties in human language. This raises a fundamental question: can ICL arise organically in other sequence domains purely through large-scale predictive training? To explore this, we turn to genomic sequences, an alternative symbolic domain rich in statistical structure. Specifically, we study the Evo2 genomic model, trained predominantly on next-nucleotide (A/T/C/G) prediction, at a scale comparable to mid-sized LLMs. We develop a controlled experimental framework comprising symbolic reasoning tasks instantiated in both linguistic and genomic forms, enabling direct comparison of ICL across genomic and linguistic models. Our results show that genomic models, like their linguistic counterparts, exhibit log-linear gains in pattern induction as the number of in-context demonstrations increases. To the best of our knowledge, this is the first evidence of organically emergent ICL in genomic sequences, supporting the hypothesis that ICL arises as a consequence of large-scale predictive modeling over rich data. These findings extend emergent meta-learning beyond language, pointing toward a unified, modality-agnostic view of in-context learning.

On the public dissemination and open sourcing of ultrasound resources, datasets and deep learning models

npj Digital Medicine, Published online: 24 November 2025; doi:10.1038/s41746-025-02162-4

On the public dissemination and open sourcing of ultrasound resources, datasets and deep learning models

Multimodal analysis of whole slide images in colorectal cancer

npj Digital Medicine, Published online: 24 November 2025; doi:10.1038/s41746-025-02095-y

Multimodal analysis of whole slide images in colorectal cancer

Integrative analysis of genomic and transcriptomic data informs precancer progression in the pancreas

bioRxiv [Preprint]. 2025 Nov 4:2025.11.03.686234. doi: 10.1101/2025.11.03.686234.

ABSTRACT

Pancreatic ductal adenocarcinoma (PDAC) arises from heterogeneous precursor lesions, including intraductal papillary mucinous neoplasms (IPMNs), but the features distinguishing indolent from progressive lesions remain unclear. We performed an integrative analysis of transcriptomic, genomic, and microenvironmental profiles of IPMNs to define multi-omic phenotypes. Using transfer learning, we projected IPMN-derived transcriptional programs onto spatial transcriptomic datasets from IPMNs and pancreatic intraepithelial neoplasias (PanINs). We identified two major phenotypes: one associated with cancer-associated fibroblasts and epithelial-to-mesenchymal transition, shared across IPMN, PanIN, and PDAC; and a second, glycolysis-enriched phenotype with a unique somatic mutation profile specific to IPMN. Spatial mapping further revealed grade-specific enrichment of transcriptional programs and distinct interactions with stromal and immune subtypes, underscoring the role of the precancer microenvironment in progression. These findings establish multi-omic phenotypes that unify genetic, transcriptional, and microenvironmental heterogeneity, providing a framework for distinguishing progressive from indolent precancers and a web-based public atlas for future exploration of these data and transcriptional phenotypes.

PMID:41279473 | PMC:PMC12637499 | DOI:10.1101/2025.11.03.686234

Knowledge-informed multimodal cfDNA analysis improves sensitivity and generalization in cancer detection

bioRxiv [Preprint]. 2025 Oct 21:2025.10.20.683167. doi: 10.1101/2025.10.20.683167.

ABSTRACT

Liquid biopsy offers a minimally invasive opportunity to detect and monitor cancers through analysis of cell-free DNA (cfDNA). However, current approaches face challenges of limited sensitivity at low tumor fractions, technical variability, and poor generalization across cohorts. Tumor-informed targeted methods offer high specificity but suffer from low sensitivity due to random sampling, tumor evolution and adaptation (including resistance mechanisms), and other sources of heterogeneity. Conversely, tumor-naive genome-wide methods can increase sensitivity but often sacrifice specificity, particularly at low tumor fractions. We developed Fragmentomics Analysis for Tumor Evaluation with AI (Fate-AI), a multimodal framework that integrates fragmentomic and methylation-derived features from low-pass whole-genome sequencing (LPWGS) and cell-free methylated DNA immunoprecipitation and high-throughput sequencing (cfMeDIP-seq). It employs a knowledge-informed strategy to select recurrently altered genomic regions and tissue-specific methylation loci to combine the advantages of tumor-naive approaches with the specificity of tumor-informed approaches. This approach derives robust per-sample normalized features that mitigate batch effects and enhance cross-cohort reproducibility. We evaluated Fate-AI on a total of 1,219 plasma samples spanning ten cancer types and healthy controls from multiple laboratories and sequencing centers, including 432 newly profiled cases (280 with both cfMeDIP-seq and LPWGS) together with 787 samples from four independent public datasets. Fate-AI achieved superior sensitivity and specificity compared to state-of-the-art methods, detecting tumor-derived signals at fractions as low as 10-5 in experimental dilutions. Fate-AI scores correlated with disease stage and tracked longitudinal progression, anticipating relapse months before clinical progression. Furthermore, Fate-AI enabled tissue-of-origin classification, with AUCs ranging from 0.84 to 0.97 across six cancer types. Collectively, our results demonstrate that Fate-AI provides a sensitive, generalizable, and clinically actionable platform for early detection, minimal residual disease monitoring, and tissue-of-origin classification, supporting its potential as a liquid biopsy framework in precision oncology.

PMID:41278930 | PMC:PMC12633305 | DOI:10.1101/2025.10.20.683167

Health care Experiences of Educated Young Adults With Blindness in the Digital Age: Qualitative Study

Background: The rapid advancement of digital health technologies (DHTs) offers substantial potential for improving healthcare access, yet it simultaneously risks exacerbating existing inequities for marginalized populations. Previous research on the digital divide has often treated individuals with blindness as a homogenous group, primarily focusing on barriers related to digital access and skills. However, less is known about the nuanced experiences of specific subgroups, such as educated and digitally literate young adults. This study focuses on this demographic to understand how their advanced digital capabilities interact with systemic and infrastructural barriers in healthcare. Objective: This qualitative study aimed to explore the lived healthcare experiences of educated young adults with blindness in China, specifically identifying how DHTs simultaneously contribute to their empowerment and exclusion. Methods: Eligible participants were educated young adults with blindness in China (aged 18-30 years, Mandarin speakers, smartphone users, and holding or pursuing higher education). A total of 12 semi-structured interviews were conducted in Mandarin during September 2024. All interviews were audio-recorded and transcribed verbatim. An inductive thematic analysis was employed to interpret the data and identify key themes. Results: Participants’ experiences highlighted an “empowered but excluded” dynamic. Seven key themes emerged, categorized into empowerment and exclusion. Empowerment themes included: (1) digital platforms empowering self-management and healthcare access, where DHTs enabled independent appointment booking and access to comprehensive health information; and (2) digital platforms empowering for finding medical visit companions, facilitating the discovery of companions for physical and emotional support. Exclusion themes comprised: (3) inaccessible online appointment systems, due to non-inclusive designs; (4) inaccessible healthcare environments and information formats, stemming from non-accessible self-service machines and written materials; (5) lack of provider competencies in respecting patient autonomy, as providers often assumed digital incompetence; (6) data privacy and security concerns, heightened by increased digitalization and reliance on assistive tools; and (7) challenges related to the quality and consistency of online companion support, highlighting the limitations of platform-based assistance. Conclusions: Our findings reveal an “empowered but excluded” dynamic: the potential for digital empowerment and enhanced independence is often curtailed by systematic barriers. Addressing this necessitates a multifaceted approach: enhancing technological accessibility through robust standards adherence and inclusive co-design processes; improving healthcare provider competencies in patient-centered care via targeted training; and empowering educated young blind adults by building their capacity for self-determination to achieve equitable healthcare access.

Impact of Digital Interventions on the Treatment Burden of Patients With Chronic Conditions: Systematic Review

Background: Digital interventions can provide cost-effective, quality health care for patients with chronic conditions. Patients with chronic conditions often are burdened by a substantial load of adhering to a treatment regimen and suffer from impacts on their function and well-being. This treatment burden has consequences for treatment adherence and disease outcomes. Digital interventions have the potential to alleviate the burden, but they also may cause new challenges and an increased workload for the patient. Previous reviews have examined digital interventions or treatment burden separately, but there is a lack of systematic reviews on the intersection of digital interventions, treatment burden, and chronic conditions. Objective: This systematic review aimed to evaluate the evidence of how digital interventions impact the treatment burden experienced by people with chronic conditions, and to assess the quality of this evidence. Methods: We searched databases PubMed, Scopus, Web of Science, ACM, PubMed Central, and CINAHL for articles published between January 1, 2013, and June 17, 2025. We included studies that had key topics related to chronic conditions, treatment burden, and digital interventions. A total of 2 reviewers independently screened the articles in 2 stages, extracted data on study design, participant characteristics, intervention type, and treatment burden outcomes from included articles, and assessed their quality using the Critical Appraisal tools from the Joanna Briggs Institute. A convergent integrated approach was used for data synthesis and integration, where quantitative data were converted into qualitative data, and the qualitative and quantitative evidence were analyzed and categorized together. Results: We included 46 relevant studies in total. We categorized the interventions into 4 types: Telehealth, informational resources, self-management tools, and facilitated tools. The results of this study indicate that digital interventions mostly support patients with chronic conditions with their treatment burden, with minor concerns of increasing treatment burden. The main benefits are support with self-management, informational support, and easier ways to contact health care professionals. The main concerns were accessibility issues, time-consuming tools, and causing fear and anxiety. Conclusions: Our findings demonstrate how treatment burden is a relevant concept for future digital health care research and practice. Digital interventions can help patients with their treatment burden by supporting self-management, improving access to health care, improving patients’ experience, and addressing relevant concerns. More research is needed about conditions with low or medium initial treatment burden.

Clinical validation of a three-marker methylation panel to detect CIN3+ in vaginal self-samples in the Dutch population-based screening programme

The use of vaginal self-sampling for cervical cancer screening is promising and increasing. However, triage cytology cannot be performed on vaginal self-sampling material after a high-risk human papilloma viru...

Accelerating Local AI on Consumer GPUs: A Hardware-Aware Dynamic Strategy for YOLOv10s

arXiv:2509.07928v2 Announce Type: replace-cross Abstract: As local AI grows in popularity, there is a critical gap between the benchmark performance of object detectors and their practical viability on consumer-grade hardware. While models like YOLOv10s promise real-time speeds, these metrics are typically achieved on high-power, desktop-class GPUs. This paper reveals that on resource-constrained systems, such as laptops with RTX 4060 GPUs, performance is not compute-bound but is instead dominated by system-level bottlenecks, as illustrated by a simple bottleneck test. To overcome this hardware-level constraint, we introduce a Two-Pass Adaptive Inference algorithm, a model-independent approach that requires no architectural changes. This study mainly focuses on adaptive inference strategies and undertakes a comparative analysis of architectural early-exit and resolution-adaptive routing, highlighting their respective trade-offs within a unified evaluation framework. The system uses a fast, low-resolution pass and only escalates to a high-resolution model pass when detection confidence is low. On a 5000-image COCO dataset, our method achieves a 1.85x speedup over a PyTorch Early-Exit baseline, with a modest mAP loss of 5.51%. This work provides a practical and reproducible blueprint for deploying high-performance, real-time AI on consumer-grade devices by shifting the focus from pure model optimization to hardware-aware inference strategies that maximize throughput.

Uncertainty Makes It Stable: Curiosity-Driven Quantized Mixture-of-Experts

arXiv:2511.11743v2 Announce Type: replace-cross Abstract: Deploying deep neural networks on resource-constrained devices faces two critical challenges: maintaining accuracy under aggressive quantization while ensuring predictable inference latency. We present a curiosity-driven quantized Mixture-of-Experts framework that addresses both through Bayesian epistemic uncertainty-based routing across heterogeneous experts (BitNet ternary, 1-16 bit BitLinear, post-training quantization). Evaluated on audio classification benchmarks (ESC-50, Quinn, UrbanSound8K), our 4-bit quantization maintains 99.9 percent of 16-bit accuracy (0.858 vs 0.859 F1) with 4x compression and 41 percent energy savings versus 8-bit. Crucially, curiosity-driven routing reduces MoE latency variance by 82 percent (p = 0.008, Levene's test) from 230 ms to 29 ms standard deviation, enabling stable inference for battery-constrained devices. Statistical analysis confirms 4-bit/8-bit achieve practical equivalence with full precision (p > 0.05), while MoE architectures introduce 11 percent latency overhead (p

Pan-cancer prevalence, risk, and clinical and demographic characteristics of Lynch Syndrome-associated variants in BioBank Japan

Commun Med (Lond). 2025 Nov 13. doi: 10.1038/s43856-025-01231-9. Online ahead of print.

ABSTRACT

BACKGROUND: Although germline testing for DNA mismatch repair (MMR) genes is routinely performed, clinical guidelines highlight evidence gaps due to limited populations and biases. We examined germline pathogenic variants of MMR genes (MLH1, MSH2, MSH6, and PMS2) in 112,927 unselected individuals from BioBank Japan.

METHODS: We analyzed 74,085 cancer patients with 23 cancer types and 38,842 controls matched by sex, age, and hospital area from BioBank Japan, collected between April 2003 and March 2018. Germline pathogenic variants in the coding regions and 2 bp flanking intronic sequences of MMR genes were identified using a multiplex PCR-based target sequencing method. We examined associations with cancer types and demographic characterization of the pathogenic variants, comparing findings to existing clinical guidelines.

RESULTS: Here we show 228 pathogenic variants identified in MMR genes, with pathogenic MSH6 variants most frequently observed in endometrial cancer and 12 other significant associations. Twelve other significant associations are noted across a broad range of odds ratios, whereas pancreatic cancer exhibits no such association. Pathogenic variant carriers are diagnosed up to 12.4 years earlier than non-carriers, and colorectal and gastric cancers are diagnosed up to 16.4 years later than indicated by the guidelines. Higher carrier frequencies are observed in patients with both colorectal and endometrial cancers (24.8%) and in those with endometrial cancer and a family history of endometrial (26.0%) or colorectal (16.1%) cancers.

CONCLUSIONS: This study provides critical insights for clinical guidelines on the associations between cancer types, age at diagnosis, and carrier frequency.

PMID:41258140 | DOI:10.1038/s43856-025-01231-9

Latent plasticity of the human pancreas across development, health, and disease

bioRxiv [Preprint]. 2025 Oct 3:2025.10.01.679230. doi: 10.1101/2025.10.01.679230.

ABSTRACT

The pancreas plays a central role in major human diseases, yet our understanding of its cellular diversity and plasticity remains incomplete. Here, we present a single-cell multiomics atlas of the human pancreas, profiling over four million cells and nuclei from 57 donors across fetal development, adult homeostasis, and type 2 diabetes (T2D). Integrating sc/snRNA-seq, snATAC-seq, VASA-seq, spatial transcriptomics (Xenium), and multiplexed proteomics (CODEX), we resolve gene expression, chromatin accessibility, and spatial organization at high resolution. We identify transcriptionally plastic centroacinar-like cells (pCACs) in adults with fetal-like features, delineate endocrine and exocrine lineage trajectories during development, and uncover HNF1A-defined beta cell epigenetic states. In T2D, we observe shifts in beta cell subtypes and altered regulatory programs. Glucose perturbation of healthy islets reveals cell-type-specific adaptation and stress responses. This atlas provides a foundational framework to understand pancreas biology and the role of cellular plasticity in regeneration and disease.

PMID:41256699 | PMC:PMC12622017 | DOI:10.1101/2025.10.01.679230

SMMILe enables accurate spatial quantification in digital pathology using multiple-instance learning

Nature Cancer, Published online: 19 November 2025; doi:10.1038/s43018-025-01060-8

Gao et al. present SMMILe, a multiple-instance learning-based tool that leverages whole-slide images for accurate spatial quantification without compromising on classification performance, and show it outperforms state-of-the-art methods.

Foundation Models in Medical Imaging: A Review and Outlook

arXiv:2506.09095v4 Announce Type: replace-cross Abstract: Foundation models (FMs) are changing the way medical images are analyzed by learning from large collections of unlabeled data. Instead of relying on manually annotated examples, FMs are pre-trained to learn general-purpose visual features that can later be adapted to specific clinical tasks with little additional supervision. In this review, we examine how FMs are being developed and applied in pathology, radiology, and ophthalmology, drawing on evidence from over 150 studies. We explain the core components of FM pipelines, including model architectures, self-supervised learning methods, and strategies for downstream adaptation. We also review how FMs are being used in each imaging domain and compare design choices across applications. Finally, we discuss key challenges and open questions to guide future research.
❌