Normal view
-
cs.AI, q-bio.NC updates on arXiv.org
-
GlobalDentBench: A Multinational Benchmark for Evaluating LLM Clinical Reasoning in Dentistry with Expert Calibration
arXiv:2605.24636v2 Announce Type: new Abstract: While large language models (LLMs) hold transformative potential for medicine, their reasoning robustness and safety in real-world clinical scenarios remain critically underexplored, particularly in dentistry. Here we introduce GlobalDentBench, the first multinational dental benchmark, featuring a taxonomy that encompasses 14 dental specialties across 88 countries and regions spanning six continents. The benchmark comprises 8,978 expert-validated
-
cs.AI, q-bio.NC updates on arXiv.org
-
Position: Science of AI Evaluation Requires Item-level Benchmark Data
arXiv:2604.03244v1 Announce Type: new Abstract: AI evaluations have become the primary evidence for deploying generative AI systems across high-stakes domains. However, current evaluation paradigms often exhibit systemic validity failures. These issues, ranging from unjustified design choices to misaligned metrics, remain intractable without a principled framework for gathering validity evidence and conducting granular diagnostic analysis. In this position paper, we argue that item-level AI ben
Position: Science of AI Evaluation Requires Item-level Benchmark Data
-
Omics in Hepatocellular
-
Hypoxia-related and immune phenotype-related fusion model for non-invasive prognostication of hepatocellular carcinoma treated by TACE: a multicentre study
Gut. 2026 Mar 30:gutjnl-2025-337938. doi: 10.1136/gutjnl-2025-337938. Online ahead of print.ABSTRACTBACKGROUND: Survival outcomes after transarterial chemoembolisation (TACE) vary in hepatocellular carcinoma (HCC) patients, and existing prognostic scores and imaging models often lack generalisability and biological interpretability.OBJECTIVE: To develop and validate a multimodal prognostication model for HCC that allows for a precise assessment of survival outcomes of HCC patients receiving TACE
Hypoxia-related and immune phenotype-related fusion model for non-invasive prognostication of hepatocellular carcinoma treated by TACE: a multicentre study
Gut. 2026 Mar 30:gutjnl-2025-337938. doi: 10.1136/gutjnl-2025-337938. Online ahead of print.
ABSTRACT
BACKGROUND: Survival outcomes after transarterial chemoembolisation (TACE) vary in hepatocellular carcinoma (HCC) patients, and existing prognostic scores and imaging models often lack generalisability and biological interpretability.
OBJECTIVE: To develop and validate a multimodal prognostication model for HCC that allows for a precise assessment of survival outcomes of HCC patients receiving TACE therapy.
DESIGN: This study enrolled 1448 HCC patients, including a TACE cohort (n=1349), a biomarker subset from a randomised trial (n=41), a single-cell RNA sequencing cohort and The Cancer Genome Atlas (TCGA) HCC cohort (n=50). Pre-treatment contrast-enhanced CT images were used to construct deep learning and conventional radiomic models. The early-fusion and late-fusion models (LFMs) were compared, and a clinical-radiologic model (CRM) was formed by integrating the better-performing LFM with clinical variables. Using TCGA data and single-cell transcriptomic profiles, the differences between high-score and low-score groups in tumour immune microenvironment, cellular functional states and key signalling pathways were investigated.
RESULTS: The CRM effectively stratified patients' survival across multiple independent cohorts and achieved more granular risk stratification than the existing clinical models. Multi-omic analyses revealed that in the LFM high-score group, myelocytomatosis oncogene was activated, epithelial-mesenchymal transition enhanced, glycolysis upregulated and hypoxia pathway activated. Single-cell transcriptomic data confirmed that virtually all cell types in high-risk patients scored high in hypoxia, and cytotoxic T cells had a reduced cytotoxic activity.
CONCLUSION: The CRM model can non-invasively predict the prognosis of HCC patients treated by TACE therapy.
PMID:41856522 | DOI:10.1136/gutjnl-2025-337938
-
Cell
-
Divergent tumor immunity determined by bacteria-cancer cell engagement
In a preclinical breast cancer metastasis model, the same bacteria strain, when present intracellularly versus extracellularly, exerts opposing effects on tumor immunity by inducing divergent neutrophil states, highlighting the intricacy in bacterial-host engagement for shaping tumor immunity.
Divergent tumor immunity determined by bacteria-cancer cell engagement
-
cs.AI, q-bio.NC updates on arXiv.org
-
daVinci-Env: Open SWE Environment Synthesis at Scale
arXiv:2603.13023v1 Announce Type: cross Abstract: Training capable software engineering (SWE) agents demands large-scale, executable, and verifiable environments that provide dynamic feedback loops for iterative code editing, test execution, and solution refinement. However, existing open-source datasets remain limited in scale and repository diversity, while industrial solutions are opaque with unreleased infrastructure, creating a prohibitive barrier for most academic research groups. We pres
daVinci-Env: Open SWE Environment Synthesis at Scale
-
cs.AI, q-bio.NC updates on arXiv.org
-
MindfulAgents: Personalizing Mindfulness Meditation via an Expert-Aligned Multi-Agent System
arXiv:2603.06926v1 Announce Type: cross Abstract: Mindfulness meditation is a widely accessible and evidence-based method for supporting mental health. Despite the proliferation of mindfulness meditation apps, sustaining user engagement remains a persistent challenge. Personalizing the meditation experience is a promising strategy to improve engagement, but it often requires costly and unscalable manual effort. We present MindfulAgents, a multi-agent system powered by large language models that
MindfulAgents: Personalizing Mindfulness Meditation via an Expert-Aligned Multi-Agent System
-
cs.AI, q-bio.NC updates on arXiv.org
-
Preference Leakage: A Contamination Problem in LLM-as-a-judge
arXiv:2502.01534v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) as judges and LLM-based data synthesis have emerged as two fundamental LLM-driven data annotation methods in model development. While their combination significantly enhances the efficiency of model training and evaluation, little attention has been given to the potential contamination brought by this new model development paradigm. In this work, we expose preference leakage, a contamination problem in LLM-as