❌

Reading view

Deep learning predicts gene rearrangements from histopathology in large B-cell lymphoma

npj Digital Medicine, Published online: 12 September 2026; doi:10.1038/s41746-026-03238-5

Deep learning predicts gene rearrangements from histopathology in large B-cell lymphoma
  •  

GlobalDentBench: A Multinational Benchmark for Evaluating LLM Clinical Reasoning in Dentistry with Expert Calibration

arXiv:2605.24636v2 Announce Type: new Abstract: While large language models (LLMs) hold transformative potential for medicine, their reasoning robustness and safety in real-world clinical scenarios remain critically underexplored, particularly in dentistry. Here we introduce GlobalDentBench, the first multinational dental benchmark, featuring a taxonomy that encompasses 14 dental specialties across 88 countries and regions spanning six continents. The benchmark comprises 8,978 expert-validated questions across three formats (multiple-choice, short-answer, and case-based questions) and assesses three progressive reasoning levels: knowledge recall (L1), routine reasoning (L2), and individualized reasoning (L3). To ensure data quality, the automated construction framework was calibrated by six senior dentists, achieving expert agreement rates of 99.98% for multiple-choice and short-answer questions and 96.78% for the more complex case-based questions. Evaluation of 12 frontier LLMs on GlobalDentBench revealed a sharp, stepwise performance degradation with increasing reasoning complexity. Specifically, accuracy plummeted from 81.34% on multiple-choice to 64.53% on short-answer and 22.34% on case-based questions, while declining markedly from 74.01% at L1 to 55.64% at L2 and 35.71% at L3. More critically, risk analysis of real-world dental cases demonstrated an alarming overall unsafe rate of 31.01% in LLM-generated clinical recommendations, with 4.51% posing risks of irreversible patient harm and risks particularly pronounced in specialties such as orthodontics. These findings expose fundamental limitations in the medical reasoning and safety of current LLMs. Consequently, GlobalDentBench provides a scalable foundation for trustworthy clinical AI evaluation, underscoring the urgent need for rigorous validation before the safe deployment of these models in healthcare.
  •  

Position: Science of AI Evaluation Requires Item-level Benchmark Data

arXiv:2604.03244v1 Announce Type: new Abstract: AI evaluations have become the primary evidence for deploying generative AI systems across high-stakes domains. However, current evaluation paradigms often exhibit systemic validity failures. These issues, ranging from unjustified design choices to misaligned metrics, remain intractable without a principled framework for gathering validity evidence and conducting granular diagnostic analysis. In this position paper, we argue that item-level AI benchmark data is essential for establishing a rigorous science of AI evaluation. Item-level analysis enables fine-grained diagnostics and principled validation of benchmarks. We substantiate this position by dissecting current validity failures and revisiting evaluation paradigms across computer science and psychometrics. Through illustrative analyses of item properties and latent constructs, we demonstrate the unique insights afforded by item-level data. To catalyze community-wide adoption, we introduce OpenEval, a growing repository of item-level benchmark data designed supporting evidence-centered AI evaluation.
  •  

Hypoxia-related and immune phenotype-related fusion model for non-invasive prognostication of hepatocellular carcinoma treated by TACE: a multicentre study

Gut. 2026 Mar 30:gutjnl-2025-337938. doi: 10.1136/gutjnl-2025-337938. Online ahead of print.

ABSTRACT

BACKGROUND: Survival outcomes after transarterial chemoembolisation (TACE) vary in hepatocellular carcinoma (HCC) patients, and existing prognostic scores and imaging models often lack generalisability and biological interpretability.

OBJECTIVE: To develop and validate a multimodal prognostication model for HCC that allows for a precise assessment of survival outcomes of HCC patients receiving TACE therapy.

DESIGN: This study enrolled 1448 HCC patients, including a TACE cohort (n=1349), a biomarker subset from a randomised trial (n=41), a single-cell RNA sequencing cohort and The Cancer Genome Atlas (TCGA) HCC cohort (n=50). Pre-treatment contrast-enhanced CT images were used to construct deep learning and conventional radiomic models. The early-fusion and late-fusion models (LFMs) were compared, and a clinical-radiologic model (CRM) was formed by integrating the better-performing LFM with clinical variables. Using TCGA data and single-cell transcriptomic profiles, the differences between high-score and low-score groups in tumour immune microenvironment, cellular functional states and key signalling pathways were investigated.

RESULTS: The CRM effectively stratified patients' survival across multiple independent cohorts and achieved more granular risk stratification than the existing clinical models. Multi-omic analyses revealed that in the LFM high-score group, myelocytomatosis oncogene was activated, epithelial-mesenchymal transition enhanced, glycolysis upregulated and hypoxia pathway activated. Single-cell transcriptomic data confirmed that virtually all cell types in high-risk patients scored high in hypoxia, and cytotoxic T cells had a reduced cytotoxic activity.

CONCLUSION: The CRM model can non-invasively predict the prognosis of HCC patients treated by TACE therapy.

PMID:41856522 | DOI:10.1136/gutjnl-2025-337938

  •  

Divergent tumor immunity determined by bacteria-cancer cell engagement

In a preclinical breast cancer metastasis model, the same bacteria strain, when present intracellularly versus extracellularly, exerts opposing effects on tumor immunity by inducing divergent neutrophil states, highlighting the intricacy in bacterial-host engagement for shaping tumor immunity.
  •  

daVinci-Env: Open SWE Environment Synthesis at Scale

arXiv:2603.13023v1 Announce Type: cross Abstract: Training capable software engineering (SWE) agents demands large-scale, executable, and verifiable environments that provide dynamic feedback loops for iterative code editing, test execution, and solution refinement. However, existing open-source datasets remain limited in scale and repository diversity, while industrial solutions are opaque with unreleased infrastructure, creating a prohibitive barrier for most academic research groups. We present OpenSWE, the largest fully transparent framework for SWE agent training in Python, comprising 45,320 executable Docker environments spanning over 12.8k repositories, with all Dockerfiles, evaluation scripts, and infrastructure fully open-sourced for reproducibility. OpenSWE is built through a multi-agent synthesis pipeline deployed across a 64-node distributed cluster, automating repository exploration, Dockerfile construction, evaluation script generation, and iterative test analysis. Beyond scale, we propose a quality-centric filtering pipeline that characterizes the inherent difficulty of each environment, filtering out instances that are either unsolvable or insufficiently challenging and retaining only those that maximize learning efficiency. With $891K spent on environment construction and an additional $576K on trajectory sampling and difficulty-aware curation, the entire project represents a total investment of approximately $1.47 million, yielding about 13,000 curated trajectories from roughly 9,000 quality guaranteed environments. Extensive experiments validate OpenSWE's effectiveness: OpenSWE-32B and OpenSWE-72B achieve 62.4% and 66.0% on SWE-bench Verified, establishing SOTA among Qwen2.5 series. Moreover, SWE-focused training yields substantial out-of-domain improvements, including up to 12 points on mathematical reasoning and 5 points on science benchmarks, without degrading factual recall.
  •  

MindfulAgents: Personalizing Mindfulness Meditation via an Expert-Aligned Multi-Agent System

arXiv:2603.06926v1 Announce Type: cross Abstract: Mindfulness meditation is a widely accessible and evidence-based method for supporting mental health. Despite the proliferation of mindfulness meditation apps, sustaining user engagement remains a persistent challenge. Personalizing the meditation experience is a promising strategy to improve engagement, but it often requires costly and unscalable manual effort. We present MindfulAgents, a multi-agent system powered by large language models that (1) generates guided meditation scripts based on an expert-established mindfulness framework, (2) encourages users' reflection on emotional states and mindfulness skills, and (3) enables real-time personalization of the mindfulness meditation experience for each user. In a formative lab study (N=13), MindfulAgents significantly improved in-session engagement (p = 0.011) and self-awareness (p = 0.014), and reduced momentary stress (p = 0.020). Furthermore, a four-week deployment study (N=62) demonstrated a notable increase in long-term engagement (p = 0.002) and level of mindfulness (p = 0.023). Participants reported that MindfulAgents offered more relevant meditation sessions personalized to individual needs in various contexts, supporting sustained practice. Our findings highlight the potential of LLM-driven personalization for enhancing user engagement in digital mindfulness meditation interventions.
  •  

Preference Leakage: A Contamination Problem in LLM-as-a-judge

arXiv:2502.01534v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) as judges and LLM-based data synthesis have emerged as two fundamental LLM-driven data annotation methods in model development. While their combination significantly enhances the efficiency of model training and evaluation, little attention has been given to the potential contamination brought by this new model development paradigm. In this work, we expose preference leakage, a contamination problem in LLM-as-a-judge caused by the relatedness between the synthetic data generators and LLM-based evaluators. To study this issue, we first define three common relatednesses between the data generator LLM and the judge LLM: being the same model, having an inheritance relationship, and belonging to the same model family. Through extensive experiments, we empirically confirm the bias of judges towards their related student models caused by preference leakage across multiple LLM baselines and benchmarks. Further analysis suggests that preference leakage is a pervasive and real-world problem that is harder to detect compared to previously identified biases in LLM-as-a-judge scenarios. All of these findings imply that preference leakage is a widespread and challenging problem in the area of LLM-as-a-judge. We release all codes and data at: https://github.com/David-Li0406/Preference-Leakage.
  •  
❌