❌

Normal view

DRScaffold: Boosting Dense-Scene Reasoning in Lightweight Vision Language Models

arXiv:2605.26038v1 Announce Type: cross Abstract: Lightweight vision-language models perform competitively on standard benchmarks yet fail systematically in dense-scene reasoning, where multiple objects, attributes, and relations must be jointly grounded and resolved through multi-step inference. Such capability is critical for real-world applications where models must reliably interpret cluttered environments. Yet existing training signals provide no explicit grounding between reasoning steps and the underlying visual entities and relations, leaving lightweight models free to generate fluent but visually unanchored reasoning chains. To address this gap, we first introduce DRBench, a benchmark of 14,573 questions across 2,943 images, organized into five task categories spanning three progressive reasoning layers. Building on DRBench, we propose DRScaffold, a supervised fine-tuning framework that decomposes the supervision target into four causally ordered stages, enforcing grounded reasoning without architectural modification. Experiments on three lightweight VLMs demonstrate substantial gains on DRBench while preserving or improving performance on general-purpose benchmarks. Notably, Qwen2.5-VL-3B trained with DRScaffold surpasses the frozen Qwen2.5-VL-32B on DRBench, demonstrating that structured supervision can substitute for a significant portion of model scale in dense-scene reasoning. Our code and models are available at https://github.com/irene-shi/DRScaffold .

Characterization of dysbiosis patterns in gut microbiota of digestive system cancers: an umbrella review

14 May 2026 at 18:00

Front Microbiol. 2026 Apr 28;17:1782471. doi: 10.3389/fmicb.2026.1782471. eCollection 2026.

ABSTRACT

Digestive system cancers (DSCs) represent a substantial global health burden. In recent years, the role of gut microbiota in the DSCs has garnered considerable attention, but its change pattern during tumor progression and the specific mechanisms are still not fully understood. We conducted a comprehensive systematic review to characterize patterns of gut microbiota dysbiosis across different DSC types and assess their clinical significance. We systematically searched four English and three Chinese databases up to January 2025 to identify systematic reviews focused on the dynamic characteristics of the gut microbiota during gastrointestinal tumorigenesis. Microbiota biodiversity and taxonomic composition were extracted to identify specific signatures associated with DSCs. The ROBIS tool was used to evaluate the methodological quality of the included studies. Ultimately, 59 studies involving six distinct DSC types were included. Data synthesis and comparison revealed distinct microbiota profiles across DSCs. At the phylum level, Bacillota was decreased in esophageal cancer (EC) and pancreatic ductal adenocarcinoma (PDAC), Pseudomonadota was augmented in EC but exhibited divergent trajectories in colorectal cancer (CRC) and PDAC. Genus-level analyses revealed Veillonella enrichment in EC and PDAC, and Fusobacterium outgrowth in EC, gastric cancer (GC) and CRC. Parvimonas and Streptococcus showed a concordant ascending trend in GC and CRC. Prevotella was overrepresented in EC and GC. This synthesis delineates a qualitative landscape of gut microbiota imbalances associated with various DSCs, highlighting the potential for these microbial shifts to serve as markers for early detection and targeted therapy. Multiomics integration and prospective cohort studies should be prioritized to accelerate clinical translation.

PMID:42131199 | PMC:PMC13161176 | DOI:10.3389/fmicb.2026.1782471

JavisDiT: Joint Audio-Video Diffusion Transformer with Hierarchical Spatio-Temporal Prior Synchronization

24 February 2026 at 13:00
arXiv:2503.23377v2 Announce Type: replace-cross Abstract: This paper introduces JavisDiT, a novel Joint Audio-Video Diffusion Transformer designed for synchronized audio-video generation (JAVG). Based on the powerful Diffusion Transformer (DiT) architecture, JavisDiT simultaneously generates high-quality audio and video content from open-ended user prompts in a unified framework. To ensure audio-video synchronization, we introduce a fine-grained spatio-temporal alignment mechanism through a Hierarchical Spatial-Temporal Synchronized Prior (HiST-Sypo) Estimator. This module extracts both global and fine-grained spatio-temporal priors, guiding the synchronization between the visual and auditory components. Furthermore, we propose a new benchmark, JavisBench, which consists of 10,140 high-quality text-captioned sounding videos and focuses on synchronization evaluation in diverse and complex real-world scenarios. Further, we specifically devise a robust metric for measuring the synchrony between generated audio-video pairs in real-world content. Experimental results demonstrate that JavisDiT significantly outperforms existing methods by ensuring both high-quality generation and precise synchronization, setting a new standard for JAVG tasks. Our code, model, and data are available at https://javisverse.github.io/JavisDiT-page/.
❌