❌

Normal view

RAU: Reference-based Anatomical Understanding with Vision Language Models

arXiv:2509.22404v2 Announce Type: replace-cross Abstract: Anatomical understanding, which is the ability to identify, localize, or segment anatomical structures, is critical in medical image analysis; however, its progress is constrained by the scarcity of expert-labeled data. A promising remedy is to leverage an annotated reference image to guide the interpretation of an unlabeled target. Although recent vision-language models (VLMs) exhibit non-trivial visual reasoning, their reference-based understanding and fine-grained localization remain limited. We introduce RAU, a framework for reference-based anatomical understanding with VLMs. We first show that a VLM learns to identify anatomical regions through relative spatial reasoning between reference and target images, trained on a moderately sized dataset. We validate this capability through visual question answering (VQA) and bounding box prediction. Next, we demonstrate that the VLM-derived spatial cues can be seamlessly integrated with the fine-grained segmentation capability of SAM2, enabling localization and pixel-level segmentation of small anatomical regions, such as vessel segments. Across two in-distribution and two out-of-distribution datasets, RAU consistently outperforms a SAM2 fine-tuning baseline using the same memory setup, yielding more accurate segmentations and more reliable localization. More importantly, its generalization ability to unseen modalities makes it scalable to unseen datasets, a property crucial for medical image applications. To the best of our knowledge, RAU is the first to explore the capability of VLMs for reference-based identification, localization, and segmentation of anatomical structures in medical images. Its promising performance highlights the potential of VLM-driven approaches for anatomical understanding in automated clinical workflows.

A World Model of Radiologist Reading for Medical Image Representation Learning

arXiv:2605.23992v1 Announce Type: cross Abstract: Radiologist eye-tracking data provide a rich record of how experts search, compare, and accumulate evidence during image reading; yet, existing methods exploit this signal only partially, either as a static spatial prior or as an auxiliary prediction target decoupled from diagnosis. We propose GazeWorld, a medical imaging world model that treats the image as the world and the radiologist's fixation sequence as a trajectory through it. GazeWorld autoregressively predicts the latent representation of the next fixated patch from all previously visited ones, while a spatial-completion branch covers unvisited regions. At inference, GazeWorld generates a sequence of patch representations from the image alone without requiring real gaze data. Frozen GazeWorld features achieve state-of-the-art diagnostic accuracy across all nine supervised settings on CheXpert, RSNA Pneumonia, and SIIM-ACR Pneumothorax, as well as the highest zero-shot accuracy on all three benchmarks. On the GazeSearch benchmark, a generic decoder trained on the same frozen features outperforms the purpose-built LogitGaze-Med by over 16\% in ScanMatch and 22\% in SED, despite not being explicitly trained to predict gaze. GazeWorld demonstrates that modeling how experts read, not just what they conclude, offers a promising pretraining paradigm for medical imaging AI.

hUCMSC-exosomes attenuate acute lung injury by inhibiting ferroptosis in pulmonary microvascular endothelial cells through ribosomal protein RPS11 upregulation

J Nanobiotechnology. 2026 May 22. doi: 10.1186/s12951-026-04565-1. Online ahead of print.

ABSTRACT

BACKGROUND: Human umbilical cord mesenchymal stem cell-derived exosomes (hUCMSC-Exos) are a promising treatment for acute lung injury (ALI)/acute respiratory distress syndrome (ARDS), but traditional delivery methods have limitations. Therefore, this study presents a noninvasive therapeutic approach for ALI/ARDS, offering new mechanistic insights and identifying potential therapeutic targets.

RESULTS: We established a nebulized LPS-induced ALI model that was characterized by diffuse lung injury and high homogeneity. Following inhalation, hUCMSC-Exos were observed to be internalized by pulmonary microvascular endothelial cells. Analysis revealed that hUCMSC-Exos alleviated ALI by reducing the severity of histological damage, pulmonary oedema, lung inflammation and ferroptosis. Additionally, hUCMSC-Exos improved the mitochondrial function of human pulmonary microvascular endothelial cells (HPMECs) via the transfer of mitochondrial components. Subsequent proteomic sequencing of mitochondria isolated from HPMECs receiving different treatments revealed the significant differential expression of ribosomal proteins among the groups. The most significantly upregulated protein, RPS11, was identified as a key mediator; its knockdown blocked the ability of hUCMSC-Exos to suppress ferroptosis and restore mitochondrial function in HPMECs. Mechanistically, hUCMSC-Exos exert their effects by enhancing mitochondria-encoded protein translation.

CONCLUSIONS: We report a mechanism whereby hUCMSC-Exos upregulate RPS11 to promote mitochondria-encoded protein translation, rescuing mitochondrial function, inhibiting ferroptosis in HPMECs, and ultimately alleviating ALI. Validated across multiple models and supported by multi-omics analyses, our findings collectively establish nebulized hUCMSC-Exos as a promising cell-free therapy targeting mitochondrial homeostasis in HPMECs for the treatment of ALI.

PMID:42174606 | DOI:10.1186/s12951-026-04565-1

Integrative Multi-Omics Analysis Identifies FTO as a Genetic and Epigenetic Link Between Metabolic Susceptibility and Staphylococcus aureus-Induced Airway Remodeling in Chronic Rhinosinusitis

Chem Biol Drug Des. 2026 Apr;107(4):e70297. doi: 10.1111/cbdd.70297.

ABSTRACT

This study identifies fat mass and obesity-associated protein (FTO) as a pivotal link between metabolic predisposition and pathogenesis associated with Staphylococcus aureus in chronic rhinosinusitis (CRS). These findings were established through the application of an integrative multi-omics framework. We demonstrate that S. aureus upregulates FTO, which functions as an m6A demethylase to stabilize the Metastasis Associated Lung Adenocarcinoma Transcript 1 (MALAT1). This molecular axis suppresses GSK-3β and promotes β-catenin nuclear translocation, thereby driving epithelial-mesenchymal transition (EMT) and pathological mucosal remodeling. By mapping the FTO-MALAT1-GSK-3β/β-catenin signaling network, this research elucidates how metabolic susceptibility facilitates infection-triggered epithelial reprogramming. These findings establish FTO as a promising biomarker and potential therapeutic target, providing a systemic foundation for personalized CRS treatment strategies.

PMID:41973807 | DOI:10.1111/cbdd.70297

Proteogenomic Analysis of Coronary Artery Calcification in Human Populations

Arterioscler Thromb Vasc Biol. 2026 Apr 2. doi: 10.1161/ATVBAHA.125.324171. Online ahead of print.

ABSTRACT

BACKGROUND: Joint use of multiple molecular layers can be useful to prioritize targets for mechanistic studies. Application of coronary disease in large populations is an emerging field.

METHODS: We used reported circulating proteomic data (Somascan aptamer-based) from ≈3000 individuals in the CARDIA study (Coronary Artery Risk Development in Young Adults), measuring association with prevalent and 10-year incident coronary artery calcium (CAC) score. We used a multiparametric approach to prioritize circulating protein-CAC associations via genomics of circulating protein levels and coronary artery transcription.

RESULTS: Proteins linked to prevalent/incident CAC in CARDIA implicated pathogenic mechanisms of vascular disease, including fibrosis and inflammation (GDF-15 [growth/differentiation factor 15], CDCP1 [CUB domain-containing protein 1], GSN [gelsolin], THBS2 [thrombospondin-2], chemokines, RNAS6), oxidative lipid metabolism (CILP2), extracellular matrix remodeling and signaling (MMPs [matrix metalloproteinases], TIMP-1, integrins), calcification (Notch 1, ARHGAP36 [Rho GTPase-activating protein 36]), and metabolism (GIP [gastric inhibitory polypeptide]), as well as new proteins not previously reported. Using protein-wide association study genetic approaches, several targets with nominal evidence in CAC proteomics were associated with atherosclerosis or myocardial infarction in over 300K individuals, including PCSK9 (proprotein convertase subtilisin/kexin type 9) and APO C1. Finally, the coronary artery-specific transcriptome-wide association study of CAC yielded genes with previously implicated mechanistic roles in vascular homeostasis, inflammation, and metabolism, as well as genes without previously described function in CAC. Overlap across CAC proteomics and transcriptome-wide association study highlighted genes involved in vascular inflammation (S100A9), cardiac development (HES1), vessel wall structure (SPARCL1), and vascular dysfunction or plaque (NOTCH3, TNFSF12, S100A12).

CONCLUSIONS: These results report population-level multiomics in human coronary calcification, presenting a method to identify disease-relevant targets through integration of human genetic approaches with multiomics.

PMID:41924874 | DOI:10.1161/ATVBAHA.125.324171

Proteogenomic Analysis of Coronary Artery Calcification in Human Populations

Arterioscler Thromb Vasc Biol. 2026 Apr 2. doi: 10.1161/ATVBAHA.125.324171. Online ahead of print.

ABSTRACT

BACKGROUND: Joint use of multiple molecular layers can be useful to prioritize targets for mechanistic studies. Application of coronary disease in large populations is an emerging field.

METHODS: We used reported circulating proteomic data (Somascan aptamer-based) from ≈3000 individuals in the CARDIA study (Coronary Artery Risk Development in Young Adults), measuring association with prevalent and 10-year incident coronary artery calcium (CAC) score. We used a multiparametric approach to prioritize circulating protein-CAC associations via genomics of circulating protein levels and coronary artery transcription.

RESULTS: Proteins linked to prevalent/incident CAC in CARDIA implicated pathogenic mechanisms of vascular disease, including fibrosis and inflammation (GDF-15 [growth/differentiation factor 15], CDCP1 [CUB domain-containing protein 1], GSN [gelsolin], THBS2 [thrombospondin-2], chemokines, RNAS6), oxidative lipid metabolism (CILP2), extracellular matrix remodeling and signaling (MMPs [matrix metalloproteinases], TIMP-1, integrins), calcification (Notch 1, ARHGAP36 [Rho GTPase-activating protein 36]), and metabolism (GIP [gastric inhibitory polypeptide]), as well as new proteins not previously reported. Using protein-wide association study genetic approaches, several targets with nominal evidence in CAC proteomics were associated with atherosclerosis or myocardial infarction in over 300K individuals, including PCSK9 (proprotein convertase subtilisin/kexin type 9) and APO C1. Finally, the coronary artery-specific transcriptome-wide association study of CAC yielded genes with previously implicated mechanistic roles in vascular homeostasis, inflammation, and metabolism, as well as genes without previously described function in CAC. Overlap across CAC proteomics and transcriptome-wide association study highlighted genes involved in vascular inflammation (S100A9), cardiac development (HES1), vessel wall structure (SPARCL1), and vascular dysfunction or plaque (NOTCH3, TNFSF12, S100A12).

CONCLUSIONS: These results report population-level multiomics in human coronary calcification, presenting a method to identify disease-relevant targets through integration of human genetic approaches with multiomics.

PMID:41924874 | DOI:10.1161/ATVBAHA.125.324171

Balancing Efficiency and Empathy: Healthcare Providers' Perspectives on AI-Supported Workflows for Serious Illness Conversations in the Emergency Department

arXiv:2506.00241v2 Announce Type: replace-cross Abstract: Serious Illness Conversations (SICs), discussions about values and care preferences for patients with life-threatening illness, rarely occur in Emergency Departments (EDs), despite evidence that early conversations improve care alignment and reduce unnecessary interventions. We interviewed 11 ED providers to identify challenges in SICs and opportunities for technology support, with a focus on AI. Our analysis revealed a four-stage SIC workflow (identification, preparation, conduction, documentation) and barriers at each stage, including fragmented patient information, limited time and space, lack of conversational guidance, and burdensome documentation. Providers expressed interest in AI systems for synthesizing information, supporting real-time conversations, and automating documentation, but emphasized concerns about preserving human connection and clinical autonomy. This tension highlights the need for technologies that enhance efficiency without undermining the interpersonal nature of SICs. We propose design guidelines for ambient and peripheral AI systems to support providers while preserving the essential humanity of these conversations.

General scales unlock AI evaluation with explanatory and predictive power

Nature, Published online: 01 April 2026; doi:10.1038/s41586-026-10303-2

A fully automated methodology based on rubrics capturing a broad range of cognitive and intellectual demands is illustrated using LLMs and tasks, demonstrating a new way to evaluate the capabilities of AI systems and anticipate their performance.

Cellular Senescence in Gastric Cancer: Molecular Mechanisms, Microenvironment Remodeling and Therapeutic Implications

Aging Dis. 2026 Mar 19. doi: 10.14336/AD.2025.1571. Online ahead of print.

ABSTRACT

Gastric cancer (GC) remains a leading cause of cancer-related morbidity and mortality worldwide, with poor prognosis for advanced-stage patients. Therefore, in-depth exploration of the mechanisms underlying GC initiation and progression, as well as the development of novel therapeutic strategies, is of crucial importance. Cellular senescence is a stable cell cycle arrest program that plays a dual role in GC. It exerts tumor-suppressive effects via growth arrest but also promotes tumor progression and immune evasion by remodeling the tumor microenvironment (TME) through senescence-associated secretory phenotype (SASP). This review comprehensively elucidates the molecular mechanisms of cellular senescence in GC and the core regulatory networks involving gene regulation, epigenetic modifications, metabolic reprogramming, and cell cycle arrest. Additionally, the review highlights how senescent cells foster an immunosuppressive microenvironment via SASP, forming a self-reinforcing feed-forward loop. Regarding therapeutic strategies, we summarize potential approaches targeting cellular senescence, including senescence induction, senescent cell clearance, SASP modulation, and multi-target synergistic therapy by integrating epigenetic regulation, metabolic intervention, and immune microenvironment modulation. Despite progress, numerous challenges remain. Future studies should leverage multi-omics technologies, novel models' development, and large-scale clinical trials to advance the clinical translation of GC cellular senescence research, providing new insights for improving prognosis.

PMID:41910653 | DOI:10.14336/AD.2025.1571

Thinking with Gaze: Sequential Eye-Tracking as Visual Reasoning Supervision for Medical VLMs

arXiv:2603.06697v1 Announce Type: cross Abstract: Vision--language models (VLMs) process images as visual tokens, yet their intermediate reasoning is often carried out in text, which can be suboptimal for visually grounded radiology tasks. Radiologists instead diagnose via sequential visual search; eye-tracking captures this process as time-ordered gaze trajectories that reveal how evidence is acquired over time. We use eye-gaze as supervision to guide VLM reasoning by introducing a small set of dedicated gaze tokens. These tokens are trained to predict gaze-selected image patch indices in temporal order, encouraging the model to follow human-like evidence acquisition and integration. Experiments on MIMIC-EYE and multiple external zero-shot benchmarks show consistent gains over baselines, achieving state-of-the-art in-domain performance and improved out-of-domain robustness. These results highlight temporally ordered gaze as an effective supervision signal for learning visually grounded medical reasoning.

ITO: Images and Texts as One via Synergizing Multiple Alignment and Training-Time Fusion

arXiv:2603.02767v3 Announce Type: replace-cross Abstract: Image-text contrastive pretraining has become a dominant paradigm for visual representation learning, yet existing methods often yield representations that remain partially organized by modality. We propose ITO, a framework addressing this limitation through two synergistic mechanisms. Multimodal multiple alignment enriches supervision by mining diverse image-text correspondences, while a lightweight training-time multimodal fusion module enforces structured cross-modal interaction. Crucially, the fusion module is discarded at inference, preserving the efficiency of standard dual-encoder architectures. Extensive experiments show that ITO consistently outperforms strong baselines across classification, retrieval, and multimodal benchmarks. Our analysis reveals that while multiple alignment drives discriminative power, training-time fusion acts as a critical structural regularizer -- eliminating the modality gap and stabilizing training dynamics to prevent the early saturation often observed in aggressive contrastive learning.

Order Is Not Layout: Order-to-Space Bias in Image Generation

arXiv:2603.03714v1 Announce Type: cross Abstract: We study a systematic bias in modern image generation models: the mention order of entities in text spuriously determines spatial layout and entity--role binding. We term this phenomenon Order-to-Space Bias (OTS) and show that it arises in both text-to-image and image-to-image generation, often overriding grounded cues and causing incorrect layouts or swapped assignments. To quantify OTS, we introduce OTS-Bench, which isolates order effects with paired prompts differing only in entity order and evaluates models along two dimensions: homogenization and correctness. Experiments show that Order-to-Space Bias (OTS) is widespread in modern image generation models, and provide evidence that it is primarily data-driven and manifests during the early stages of layout formation. Motivated by this insight, we show that both targeted fine-tuning and early-stage intervention strategies can substantially reduce OTS, while preserving generation quality.

ITO: Images and Texts as One via Synergizing Multiple Alignment and Training-Time Fusion

arXiv:2603.02767v2 Announce Type: replace-cross Abstract: Image-text contrastive pretraining has become a dominant paradigm for visual representation learning, yet existing methods often yield representations that remain partially organized by modality. We propose ITO, a framework addressing this limitation through two synergistic mechanisms. Multimodal multiple alignment enriches supervision by mining diverse image-text correspondences, while a lightweight training-time multimodal fusion module enforces structured cross-modal interaction. Crucially, the fusion module is discarded at inference, preserving the efficiency of standard dual-encoder architectures. Extensive experiments show that ITO consistently outperforms strong baselines across classification, retrieval, and multimodal benchmarks. Our analysis reveals that while multiple alignment drives discriminative power, training-time fusion acts as a critical structural regularizer -- eliminating the modality gap and stabilizing training dynamics to prevent the early saturation often observed in aggressive contrastive learning.

ITO: Images and Texts as One via Synergizing Multiple Alignment and Training-Time Fusion

arXiv:2603.02767v1 Announce Type: cross Abstract: Image-text contrastive pretraining has become a dominant paradigm for visual representation learning, yet existing methods often yield representations that remain partially organized by modality. We propose ITO, a framework addressing this limitation through two synergistic mechanisms. Multimodal multiple alignment enriches supervision by mining diverse image-text correspondences, while a lightweight training-time multimodal fusion module enforces structured cross-modal interaction. Crucially, the fusion module is discarded at inference, preserving the efficiency of standard dual-encoder architectures. Extensive experiments show that ITO consistently outperforms strong baselines across classification, retrieval, and multimodal benchmarks. Our analysis reveals that while multiple alignment drives discriminative power, training-time fusion acts as a critical structural regularizer -- eliminating the modality gap and stabilizing training dynamics to prevent the early saturation often observed in aggressive contrastive learning.

Proximity-Based Multi-Turn Optimization: Practical Credit Assignment for LLM Agent Training

arXiv:2602.19225v1 Announce Type: new Abstract: Multi-turn LLM agents are becoming pivotal to production systems, spanning customer service automation, e-commerce assistance, and interactive task management, where accurately distinguishing high-value informative signals from stochastic noise is critical for sample-efficient training. In real-world scenarios, a failure in a trivial task may reflect random instability, whereas success in a high-difficulty task signifies a genuine capability breakthrough. Yet, existing group-based policy optimization methods rigidly rely on statistical deviation within discrete batches, frequently misallocating credit when task difficulty fluctuates. To address this issue, we propose Proximity-based Multi-turn Optimization (ProxMO), a practical and robust framework engineered specifically for the constraints of real-world deployment. ProxMO integrates global context via two lightweight mechanisms: success-rate-aware modulation dynamically adapts gradient intensity based on episode-level difficulty, while proximity-based soft aggregation derives baselines through continuous semantic weighting at the step level. Extensive evaluations on ALFWorld and WebShop benchmarks demonstrate that ProxMO yields substantial performance gains over existing baselines with negligible computational cost. Ablation studies further validate the independent and synergistic efficacy of both mechanisms. Crucially, ProxMO offers plug-and-play compatibility with standard GRPO frameworks, facilitating immediate, low-friction adoption in existing industrial training pipelines. Our implementation is available at: \href{https://anonymous.4open.science/r/proxmo-B7E7/README.md}{https://anonymous.4open.science/r/proxmo}.
❌