❌

Normal view

Explainable Retinal Imaging for Prediction of Multi-Organ Dysfunction in Type 2 Diabetes

arXiv:2605.24912v1 Announce Type: cross Abstract: Background: Type 2 diabetes mellitus (T2DM) is increasingly recognised as a systemic disease characterised by coordinated dysfunction across metabolic, renal, lipid, and inflammatory pathways. Existing clinical assessments often fail to capture this multi-dimensional burden. Methods: We conducted a retrospective study of 1,195 patients using routinely collected laboratory biomarkers. System-level abnormality indices were constructed to quantify organ-specific dysfunction, and multi-system involvement was defined as abnormalities in two or more systems. Supervised machine learning models, including logistic regression, random forest, and gradient boosting, were trained to predict multi-system dysregulation. Model interpretability was achieved using SHapley Additive exPlanations (SHAP). Results: The gradient boosting model demonstrated near-perfect discrimination (AUC = 1.000), significantly outperforming logistic regression (AUC = 0.925). Feature attribution analysis revealed that hyperglycaemia, renal impairment, dyslipidaemia, and inflammation were the dominant drivers of multi-system risk. Dose-response relationships observed in partial dependence analyses further supported the biological plausibility of model predictions. Conclusion: This study presents an interpretable, data-driven framework for quantifying systemic disease burden in T2DM. By linking routine biomarkers to multi-organ dysfunction, our approach provides both predictive accuracy and mechanistic insight, offering potential for improved risk stratification and precision medicine in diabetes care. The data and code used in this study are openly available on GitHub at: https://github.com/MiniHanWang/Type-2-Diabetes-1.git

Explainable Multi-Task Retinal Imaging Reveals Microvascular Signals for Systemic Risk Stratification in Type 2 Diabetes: A Pilot Study

arXiv:2605.24913v1 Announce Type: cross Abstract: Retinal imaging provides a non-invasive window into systemic microvascular health and has emerged as a potential biomarker for systemic diseases. However, whether retinal features encode biologically meaningful systemic signals that can be reliably interpreted using explainable artificial intelligence (XAI) remains unclear. An explainable multi-task deep learning framework was developed to investigate associations between retinal microvascular features and systemic abnormalities in Type 2 Diabetes Mellitus. A total of 11,011 fundus images from 2,719 individuals were analysed using a shared neural network with task-specific heads for glycaemic status, kidney abnormality, and multi-system involvement. Model interpretability was evaluated using Gradient-weighted Class Activation Mapping (Grad-CAM), anatomical masking, and vessel alignment analysis. The framework demonstrated task-dependent predictive performance, with the best discrimination observed for kidney abnormality (AUC up to 0.63), whereas glycaemic status prediction showed limited performance (AUC = 0.49-0.61). Explainability analyses consistently localized model attention to retinal vessels and peripapillary regions. Masking experiments showed that occlusion of vascular regions caused the greatest performance decline, indicating that retinal vessels were the primary predictive source. Different architectures exhibited heterogeneous attention patterns, suggesting multiple representational pathways for systemic signal encoding. This pilot study demonstrates that retinal microvascular features contain measurable signals associated with systemic abnormalities, particularly microvascular damage. By integrating multi-task learning with quantitative XAI validation, this framework advances retinal imaging toward interpretable digital biomarkers for systemic risk stratification in diabetes.

Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation

arXiv:2508.13998v2 Announce Type: replace-cross Abstract: Generalization in embodied AI is hindered by the "seeing-to-doing gap," which stems from data scarcity and embodiment heterogeneity. To address this, we pioneer "pointing" as a unified, embodiment-agnostic intermediate representation, defining four core embodied pointing abilities that bridge high-level vision-language comprehension with low-level action primitives. We introduce Embodied-R1, a 3B Vision-Language Model (VLM) specifically designed for embodied reasoning and pointing. We use a wide range of embodied and general visual reasoning datasets as sources to construct a large-scale dataset, Embodied-Points-200K, which supports key embodied pointing capabilities. We then train Embodied-R1 using a two-stage Reinforced Fine-tuning (RFT) curriculum with a specialized multi-task reward design. Embodied-R1 achieves state-of-the-art performance on 11 embodied spatial and pointing benchmarks. Critically, it demonstrates robust zero-shot generalization by achieving a 56.2% success rate in the SIMPLEREnv and 87.5% across 8 real-world XArm tasks without any task-specific fine-tuning, representing a 62% improvement over strong baselines. Furthermore, the model exhibits high robustness against diverse visual disturbances. Our work shows that a pointing-centric representation, combined with an RFT training paradigm, offers an effective and generalizable pathway to closing the perception-action gap in robotics.

Effect of a Digital-Driven Physician-Pharmacist Collaborative Model for Diabetes in Primary Health Care: Cluster Randomized Trial

Background: Evidence-based physician-pharmacist collaborative clinics have demonstrated significant short-term benefits for patients with type 2 diabetes (T2D), but their long-term effectiveness remains unclear, especially in primary health care settings. Objective: This study aimed to explore the long-term effectiveness and cost-effectiveness of a novel, digital-driven, multifaceted physician-pharmacist collaborative model for managing patients with T2D in underresourced settings. Methods: We conducted a 12-month cluster randomized controlled trial from May 2021 to December 2022 across 6 primary health care settings in China. Guided by the theory of planned behavior, the intervention involved routine therapy from physicians along with pharmaceutical interventions from pharmacists. These were delivered through a combination of face-to-face visits and mobile health care. The intervention group received 4 face-to-face visits and biweekly remote education sessions over the 12 months. We conducted intention-to-treat analyses to estimate differences in clinical and behavior indicators between the intervention and control groups. Primary outcomes included glycosylated hemoglobin and 10-year atherosclerotic cardiovascular risk. Data were analyzed using adjusted generalized estimation equations. Results: This study included 574 patients (291 in the intervention group and 283 in the control group). Over 12 months, patients in the intervention group had significant reductions in hemoglobin A1c (–2.57 vs –1.96, respectively; P<.001; 95% CI –1.027 to –0.238) and 10-year atherosclerotic cardiovascular risk (–1.35 vs 0.01, respectively; P<.001; 95% CI –1.690 to –0.630) compared with the control group. Substantial improvements were also observed in several secondary outcomes, including fasting blood glucose, 2-hour postprandial blood glucose, waist circumference, waist-to-hip ratio, blood pressure, triglyceride, and total cholesterol. Total diabetes-related costs decreased, and patient satisfaction improved significantly in the intervention group. There were no significant differences in BMI, high-density lipoprotein, or low-density lipoprotein. Conclusions: These findings suggest that the physician-pharmacist collaborative model could improve the long-term quality and efficiency of T2D management and reduce medical costs in underresourced areas globally. Patients with T2D, especially those with central obesity or high cardiovascular risk, may benefit more from collaborative clinics. Trial Registration: Chinese Clinical Trial Registry ChiCTR2000031839; https://www.chictr.org.cn/showproj.html?proj=51910

Inhibition of ZBTB7B-mediated ADPGK transcription by NEDD4 impedes glycolysis and progression of lung adenocarcinoma

Oncogenesis, Published online: 11 March 2026; doi:10.1038/s41389-026-00605-5

Inhibition of ZBTB7B-mediated ADPGK transcription by NEDD4 impedes glycolysis and progression of lung adenocarcinoma

WebDevJudge: Evaluating (M)LLMs as Critiques for Web Development Quality

arXiv:2510.18560v3 Announce Type: replace-cross Abstract: The paradigm of LLM-as-a-judge is emerging as a scalable and efficient alternative to human evaluation, demonstrating strong performance on well-defined tasks. However, its reliability in open-ended tasks with dynamic environments and complex interactions remains unexplored. To bridge the gap, we introduce WebDevJudge, a systematic benchmark for assessing LLM-as-a-judge performance in web development, with support for both non-interactive evaluation based on static observations and continuous interactive evaluation with a dynamic web environment. WebDevJudge comprises human preference labels over paired web implementations, annotated with structured and query-grounded rubrics to ensure high-quality ground truth. Using this benchmark, we comprehensively evaluate various evaluators, including LLMs, MLLMs, and agentic workflows. We systematically investigate the impact of different paradigms and guidance mechanisms. Our experiments reveal a significant gap between LLM judges and human experts. In-depth analysis indicates this gap stems from fundamental model limitations, including failures in recognizing functional equivalence, verifying task feasibility, and mitigating bias. Overall, WebDevJudge presents a challenge to LLM-as-a-judge, offering insights to guide future research toward developing more reliable and capable automated evaluators for complicated scenarios. Code and data are available at https://github.com/lcy2723/WebDevJudge.

Emotion-LLaMAv2 and MMEVerse: A New Framework and Benchmark for Multimodal Emotion Understanding

arXiv:2601.16449v2 Announce Type: replace-cross Abstract: Understanding human emotions from multimodal signals poses a significant challenge in affective computing and human-robot interaction. While multimodal large language models (MLLMs) have excelled in general vision-language tasks, their capabilities in emotional reasoning remain limited. The field currently suffers from a scarcity of large-scale datasets with high-quality, descriptive emotion annotations and lacks standardized benchmarks for evaluation. Our preliminary framework, Emotion-LLaMA, pioneered instruction-tuned multimodal learning for emotion reasoning but was restricted by explicit face detectors, implicit fusion strategies, and low-quality training data with limited scale. To address these limitations, we present Emotion-LLaMAv2 and the MMEVerse benchmark, establishing an end-to-end pipeline together with a standardized evaluation setting for emotion recognition and reasoning. Emotion-LLaMAv2 introduces three key advances. First, an end-to-end multiview encoder eliminates external face detection and captures nuanced emotional cues via richer spatial and temporal multiview tokens. Second, a Conv Attention pre-fusion module is designed to enable simultaneous local and global multimodal feature interactions external to the LLM backbone. Third, a perception-to-cognition curriculum instruction tuning scheme within the LLaMA2 backbone unifies emotion recognition and free-form emotion reasoning. To support large-scale training and reproducible evaluation, MMEVerse aggregates twelve publicly available emotion datasets, including IEMOCAP, MELD, DFEW, and MAFW, into a unified multimodal instruction format. The data are re-annotated via a multi-agent pipeline involving Qwen2 Audio, Qwen2.5 VL, and GPT 4o, producing 130k training clips and 36k testing clips across 18 evaluation benchmarks.
❌