❌

Normal view

Beyond the Aggregation Dilemma: Prior-Retaining Decoupled Learning for Multimodal Graphs

arXiv:2605.24684v1 Announce Type: cross Abstract: Multimodal Attributed Graph Learning (MAGL) integrates intrinsic node attributes with structural topology via graph aggregation. However, as pretrained encoders evolve into Large Foundation Models (LFMs), the landscape of MAGL fundamentally shifts: under high-confidence LFM priors, mandatory aggregation introduces topological noise that overwhelms discriminative signals, triggering a counter-intuitive performance inversion where sophisticated MAGL architectures underperform simple topology-agnostic MLPs. Through systematic empirical and theoretical analysis, we identify that this inversion stems from a fundamental aggregation dilemma characterized by two concurrent pathologies: (1) Representational Pathology (SNR Degradation) - mandatory aggregation dilutes robust intrinsic features with topological noise, causing the noise penalty to outweigh its collaborative benefit; and (2) Optimization Pathology (Gradient Starvation) - topological aggregation attenuates gradient flow, while a shared task loss causes dominant modalities to prematurely suppress weaker ones. To resolve this dilemma, we propose SUPRA (Shared-Unique Prior-Retaining Architecture), a decoupled dual-pathway paradigm. SUPRA processes modality-specific features through topology-agnostic MLPs while capturing structural synergy via a lightweight shared GNN, with auxiliary deep supervision counteracting gradient starvation. Extensive evaluations demonstrate that SUPRA achieves state-of-the-art performance while requiring 3.5x lower peak GPU memory and up to 4.4x faster training time than Multimodal Graph Transformers.

Profiling-Driven Adaptive Distributed Transformer Inference on Embedded Edge Deployment

arXiv:2605.25682v1 Announce Type: cross Abstract: Distributing Transformer inference across embedded edge devices can alleviate individual memory and compute constraints, yet practical benefits on real hardware remain unclear: prior work relies largely on simulations that overlook hardware-specific communication overheads. We present a hardware prototype study on NVIDIA Jetson Orin Nano devices connected over WiFi. Our key finding is that the dominant bottleneck is not just network bandwidth but also the CPU-GPU staging during communication. Because Jetson's integrated GPU architecture lacks the PCIe/NVLink pathway that NCCL requires, all inter-device data communication should be routed through GLOO and staged in CPU memory; an overhead that scales with communication data volume and makes full-tensor exchange slower than single-device inference across the batch sizes for medium sized models such as ViT. We therefore evaluate Prism by combining Segment Means compression with lightweight offline profiling to adaptively select between local and distributed execution at runtime. Experiments show that this strategy reduces latency by 65%-77% and energy consumption by 34%-52% relative to full-tensor exchange in static distributed execution setup, demonstrating that profiling-driven adaptation is essential for practical distributed Transformer inference on embedded hardware.

AgentArk: Distilling Multi-Agent Intelligence into a Single LLM Agent

arXiv:2602.03955v3 Announce Type: replace Abstract: While large language model (LLM) multi-agent systems achieve superior reasoning performance through iterative debate, practical deployment is limited by their high computational cost and error propagation. This paper proposes AgentArk, a novel framework to distill multi-agent dynamics into the weights of a single model, effectively transforming explicit test-time interactions into implicit model capabilities. This equips a single agent with the intelligence of multi-agent systems while remaining computationally efficient. Specifically, we investigate three hierarchical distillation strategies across various models, tasks, scaling, and scenarios: reasoning-enhanced fine-tuning; trajectory-based augmentation; and process-aware distillation. By shifting the burden of computation from inference to training, the distilled models preserve the efficiency of one agent while exhibiting strong reasoning and self-correction performance of multiple agents. They further demonstrate enhanced robustness and generalization across diverse reasoning tasks. We hope this work can shed light on future research on efficient and robust multi-agent development. Our code is at https://github.com/AIFrontierLab/AgentArk.

Safety in Embodied AI: A Survey of Risks, Attacks, and Defenses

arXiv:2605.02900v2 Announce Type: replace-cross Abstract: Embodied Artificial Intelligence (Embodied AI) integrates perception, cognition, planning, and interaction into agents that operate in open-world, safety-critical environments. As these systems gain autonomy and enter domains such as transportation, healthcare, and industrial or assistive robotics, ensuring their safety becomes both technically challenging and socially indispensable. Unlike digital AI systems, embodied agents must act under uncertain sensing, incomplete knowledge, and dynamic human-robot interactions, where failures can directly lead to physical harm. This survey provides a comprehensive and structured review of safety research in embodied AI, examining attacks and defenses across the full embodied pipeline, from perception and cognition to planning, action and interaction, and agentic system. We introduce a multi-level taxonomy that unifies fragmented lines of work and connects embodied-specific safety findings with broader advances in vision, language, and multimodal foundation models. Our review synthesizes insights from over 500 papers spanning adversarial, backdoor, jailbreak, and hardware-level attacks; attack detection, safe training and robust inference; and risk-aware human-agent interaction. This analysis reveals several overlooked challenges, including the fragility of multimodal perception fusion, the instability of planning under jailbreak attacks, and the trustworthiness of human-agent interaction in open-ended scenarios. By organizing the field into a coherent framework and identifying critical research gaps, this survey provides a roadmap for building embodied agents that are not only capable and autonomous but also safe, robust, and reliable in real-world deployment.

PRXL2B facilitates the progression of hepatocellular carcinoma and the therapeutic efficacy of oncolytic adenovirus H101 through the PI3K/AKT/PD-L1 axis

Biosci Trends. 2026 May 21. doi: 10.5582/bst.2026.01000. Online ahead of print.

ABSTRACT

Oncolytic adenovirus H101 has shown antitumor activity in hepatocellular carcinoma (HCC), but the molecular determinants of treatment response remain unclear. In this study, a Hepa1-6 subcutaneous tumor model was established in C57BL/6 mice and treated with intratumoral H101, followed by integrated transcriptomic and proteomic analyses to identify candidate genes associated with H101 response. PRXL2B was selected for further investigation using public multi-omics datasets, tissue microarray-based immunohistochemistry, in vitro functional assays, mechanistic analyses, and in vivo validation experiments. Integrated multi-omics analyses identified PRXL2B as a candidate gene downregulated after H101 treatment. Public datasets and tissue-based validation further showed that PRXL2B was upregulated in HCC tissues. In MHCC97H and HCCLM3 cells, PRXL2B knockdown inhibited proliferation, migration, and invasion, promoted apoptosis and cell-cycle arrest, and enhanced the antitumor effect of H101. Mechanistically, PRXL2B silencing reduced AKT phosphorylation and PD-L1 expression. In vivo, PRXL2B knockdown suppressed tumor growth, and the combination of PRXL2B knockdown and H101 produced the strongest antitumor effect. These findings indicate that PRXL2B promotes malignant phenotypes in HCC and may modulate H101 efficacy through the PI3K/AKT/PD-L1 axis. Targeting PRXL2B may therefore represent a potential strategy to enhance the therapeutic efficacy of oncolytic virus therapy in HCC.

PMID:42161529 | DOI:10.5582/bst.2026.01000

GA-GS: Generation-Assisted Gaussian Splatting for Static Scene Reconstruction

arXiv:2604.04331v1 Announce Type: cross Abstract: Reconstructing static 3D scene from monocular video with dynamic objects is important for numerous applications such as virtual reality and autonomous driving. Current approaches typically rely on background for static scene reconstruction, limiting the ability to recover regions occluded by dynamic objects. In this paper, we propose GA-GS, a Generation-Assisted Gaussian Splatting method for Static Scene Reconstruction. The key innovation of our work lies in leveraging generation to assist in reconstructing occluded regions. We employ a motion-aware module to segment and remove dynamic regions, and thenuse a diffusion model to inpaint the occluded areas, providing pseudo-ground-truth supervision. To balance contributions from real background and generated region, we introduce a learnable authenticity scalar for each Gaussian primitive, which dynamically modulates opacity during splatting for authenticity-aware rendering and supervision. Since no existing dataset provides ground-truth static scene of video with dynamic objects, we construct a dataset named Trajectory-Match, using a fixed-path robot to record each scene with/without dynamic objects, enabling quantitative evaluation in reconstruction of occluded regions. Extensive experiments on both the DAVIS and our dataset show that GA-GS achieves state-of-the-art performance in static scene reconstruction, especially in challenging scenarios with large-scale, persistent occlusions.

Protective Effects of the Ethyl Acetate Fraction from Madeng'ai on Lipopolysaccharide-Induced Acute Lung Injury in Mice: Insights from Integrated Multi-Omics Analysis

J Ethnopharmacol. 2026 Apr 4:121650. doi: 10.1016/j.jep.2026.121650. Online ahead of print.

ABSTRACT

ETHNOPHARMACOLOGICAL RELEVANCE: Madeng'ai (MDA) is a traditional medicinal plant of the Dong ethnic group. Its roots have been widely used in folk medicine for clearing heat and removing toxins, alleviating swelling and relieving pain, dispersing blood stasis and arresting bleeding, as well as promoting wound healing. It is taxonomically classified as a variety of Potentilla freyniana Bornm.

AIM OF THE STUDY: Acute lung injury (ALI) is a life-threatening pulmonary disorder associated with high mortality, underscoring the urgent need to explore novel therapeutic strategies. This study aimed to evaluate the protective effects of the ethyl acetate fraction of MDA (MEA) against LPS-induced ALI in mice and to investigate its underlying mechanisms.

MATERIALS AND METHODS: LC-MS/MS was employed to tentatively identify the bioactive components of MEA. A mouse model of ALI was established by LPS induction. The protective effects of MEA were evaluated through assessments of lung histopathology, inflammatory cytokine levels, and oxidative stress markers. The underlying mechanisms were systematically investigated by integrating transcriptomics, metabolomics, network pharmacology, molecular docking, and Western blotting.

RESULTS: MEA significantly attenuated LPS-induced pulmonary pathological lesions, pulmonary edema, and excessive inflammatory responses in ALI mice. Comprehensive bioinformatics analyses predicted potential mechanisms involving oxidative stress and the regulation of metabolic pathways. Experimental validation via Western blotting confirmed that MEA inhibited TLR4-mediated inflammatory signaling and modulated the PI3K/AKT pathway, thereby exerting multi-pathway protective effects against ALI.

CONCLUSIONS: Collectively, this study confirms that MEA, as a traditional herbal extract, holds potential as an adjuvant therapeutic agent for ALI, providing experimental evidence for the modernization and development of ethnic medicines.

PMID:41941987 | DOI:10.1016/j.jep.2026.121650

Protective Effects of the Ethyl Acetate Fraction from Madeng'ai on Lipopolysaccharide-Induced Acute Lung Injury in Mice: Insights from Integrated Multi-Omics Analysis

J Ethnopharmacol. 2026 Apr 4:121650. doi: 10.1016/j.jep.2026.121650. Online ahead of print.

ABSTRACT

ETHNOPHARMACOLOGICAL RELEVANCE: Madeng'ai (MDA) is a traditional medicinal plant of the Dong ethnic group. Its roots have been widely used in folk medicine for clearing heat and removing toxins, alleviating swelling and relieving pain, dispersing blood stasis and arresting bleeding, as well as promoting wound healing. It is taxonomically classified as a variety of Potentilla freyniana Bornm.

AIM OF THE STUDY: Acute lung injury (ALI) is a life-threatening pulmonary disorder associated with high mortality, underscoring the urgent need to explore novel therapeutic strategies. This study aimed to evaluate the protective effects of the ethyl acetate fraction of MDA (MEA) against LPS-induced ALI in mice and to investigate its underlying mechanisms.

MATERIALS AND METHODS: LC-MS/MS was employed to tentatively identify the bioactive components of MEA. A mouse model of ALI was established by LPS induction. The protective effects of MEA were evaluated through assessments of lung histopathology, inflammatory cytokine levels, and oxidative stress markers. The underlying mechanisms were systematically investigated by integrating transcriptomics, metabolomics, network pharmacology, molecular docking, and Western blotting.

RESULTS: MEA significantly attenuated LPS-induced pulmonary pathological lesions, pulmonary edema, and excessive inflammatory responses in ALI mice. Comprehensive bioinformatics analyses predicted potential mechanisms involving oxidative stress and the regulation of metabolic pathways. Experimental validation via Western blotting confirmed that MEA inhibited TLR4-mediated inflammatory signaling and modulated the PI3K/AKT pathway, thereby exerting multi-pathway protective effects against ALI.

CONCLUSIONS: Collectively, this study confirms that MEA, as a traditional herbal extract, holds potential as an adjuvant therapeutic agent for ALI, providing experimental evidence for the modernization and development of ethnic medicines.

PMID:41941987 | DOI:10.1016/j.jep.2026.121650

Captioning Daily Activity Images in Early Childhood Education: Benchmark and Algorithm

arXiv:2604.01941v1 Announce Type: cross Abstract: Image captioning for Early Childhood Education (ECE) is essential for automated activity understanding and educational assessment. However, existing methods face two key challenges. First, the lack of large-scale, domain-specific datasets limits the model's ability to capture fine-grained semantic concepts unique to ECE scenarios, resulting in generic and imprecise descriptions. Second, conventional training paradigms exhibit limitations in enhancing professional object description capability, as supervised learning tends to favor high-frequency expressions, while reinforcement learning may suffer from unstable optimization on difficult samples. To address these limitations, we introduce ECAC, a large-scale benchmark for ECE daily activity image captioning, comprising 256,121 real-world images annotated with expert-level captions and fine-grained labels. ECAC is further equipped with a domain-oriented evaluation protocol, the Teaching Toy Recognition Score (TTS), to explicitly measure professional object naming accuracy. Furthermore, we propose RSRS (Reward-Conditional Switch of Reinforcement Learning and Supervised Fine-Tuning), a hybrid training framework that dynamically alternates between RL and supervised optimization. By rerouting hard samples with zero rewards to supervised fine-tuning, RSRS effectively mitigates advantage collapse and enables stable optimization for fine-grained recognition. Leveraging ECAC and RSRS, we develop KinderMM-Cap-3B, a domain-adapted multimodal large language model. Extensive experiments demonstrate that our model achieves a TTS of 51.06, substantially outperforming state-of-the-art baselines while maintaining superior caption quality, highlighting its potential for specialized educational applications.

JFTA-Bench: Evaluate LLM's Ability of Tracking and Analyzing Malfunctions Using Fault Trees

arXiv:2603.22978v1 Announce Type: new Abstract: In the maintenance of complex systems, fault trees are used to locate problems and provide targeted solutions. To enable fault trees stored as images to be directly processed by large language models, which can assist in tracking and analyzing malfunctions, we propose a novel textual representation of fault trees. Building on it, we construct a benchmark for multi-turn dialogue systems that emphasizes robust interaction in complex environments, evaluating a model's ability to assist in malfunction localization, which contains $3130$ entries and $40.75$ turns per entry on average. We train an end-to-end model to generate vague information to reflect user behavior and introduce long-range rollback and recovery procedures to simulate user error scenarios, enabling assessment of a model's integrated capabilities in task tracking and error recovery, and Gemini 2.5 pro archives the best performance.

SortedRL: Accelerating RL Training for LLMs through Online Length-Aware Scheduling

arXiv:2603.23414v1 Announce Type: cross Abstract: Scaling reinforcement learning (RL) has shown strong promise for enhancing the reasoning abilities of large language models (LLMs), particularly in tasks requiring long chain-of-thought generation. However, RL training efficiency is often bottlenecked by the rollout phase, which can account for up to 70% of total training time when generating long trajectories (e.g., 16k tokens), due to slow autoregressive generation and synchronization overhead between rollout and policy updates. We propose SortedRL, an online length-aware scheduling strategy designed to address this bottleneck by improving rollout efficiency and maintaining training stability. SortedRL reorders rollout samples based on output lengths, prioritizing short samples forming groups for early updates. This enables large rollout batches, flexible update batches, and near on-policy micro-curriculum construction simultaneously. To further accelerate the pipeline, SortedRL incorporates a mechanism to control the degree of off-policy training through a cache-based mechanism, and is supported by a dedicated RL infrastructure that manages rollout and update via a stateful controller and rollout buffer. Experiments using LLaMA-3.1-8B and Qwen-2.5-32B on diverse tasks, including logical puzzles, and math challenges like AIME 24, Math 500, and Minerval, show that SortedRL reduces RL training bubble ratios by over 50%, while attaining 3.9% to 18.4% superior performance over baseline given same amount of data.

SARE: Sample-wise Adaptive Reasoning for Training-free Fine-grained Visual Recognition

arXiv:2603.17729v2 Announce Type: replace-cross Abstract: Recent advances in Large Vision-Language Models (LVLMs) have enabled training-free Fine-Grained Visual Recognition (FGVR). However, effectively exploiting LVLMs for FGVR remains challenging due to the inherent visual ambiguity of subordinate-level categories. Existing methods predominantly adopt either retrieval-oriented or reasoning-oriented paradigms to tackle this challenge, but both are constrained by two fundamental limitations:(1) They apply the same inference pipeline to all samples without accounting for uneven recognition difficulty, thereby leading to suboptimal accuracy and efficiency; (2) The lack of mechanisms to consolidate and reuse error-specific experience causes repeated failures on similar challenging cases. To address these limitations, we propose SARE, a Sample-wise Adaptive textbfREasoning framework for training-free FGVR. Specifically, SARE adopts a cascaded design that combines fast candidate retrieval with fine-grained reasoning, invoking the latter only when necessary. In the reasoning process, SARE incorporates a self-reflective experience mechanism that leverages past failures to provide transferable discriminative guidance during inference, without any parameter updates. Extensive experiments across 14 datasets substantiate that SARE achieves state-of-the-art performance while substantially reducing computational overhead.

Human-specific features of the cerebellum and ZP2-regulated synapse development

Human-specific transcriptomic and regulatory features are present in the cerebellum, with ZP2 playing a key role in synapse regulation. ZP2 expression is induced by pontine mossy fibers, leading to decreased synaptic proteins and neuronal activity, which provides insights into the evolutionary development of the human cerebellum.

Evaluation Faking: Unveiling Observer Effects in Safety Evaluation of Frontier AI Systems

arXiv:2505.17815v2 Announce Type: replace Abstract: As foundation models grow increasingly more intelligent, reliable and trustworthy safety evaluation becomes more indispensable than ever. However, an important question arises: Whether and how an advanced AI system would perceive the situation of being evaluated, and lead to the broken integrity of the evaluation process? During standard safety tests on a mainstream large reasoning model, we unexpectedly observe that the model without any contextual cues would occasionally recognize it is being evaluated and hence behave more safety-aligned. This motivates us to conduct a systematic study on the phenomenon of evaluation faking, i.e., an AI system autonomously alters its behavior upon recognizing the presence of an evaluation context and thereby influencing the evaluation results. Through extensive experiments on a diverse set of foundation models with mainstream safety benchmarks, we reach the main finding termed the observer effects for AI: When the AI system under evaluation is more advanced in reasoning and situational awareness, the evaluation faking behavior becomes more ubiquitous, which reflects in the following aspects: 1) Reasoning models recognize evaluation 16% more often than non-reasoning models. 2) Scaling foundation models (32B to 671B) increases faking by over 30% in some cases, while smaller models show negligible faking. 3) AI with basic memory is 2.3x more likely to recognize evaluation and scores 19% higher on safety tests (vs. no memory). To measure this, we devised a chain-of-thought monitoring technique to detect faking intent and uncover internal signals correlated with such behavior, offering insights for future mitigation studies.

Effects of Digital Health Interventions on Functional and Psychological Outcomes in Older Patients With Hip Fractures: Systematic Review and Meta-Analysis of Randomized Controlled Trials

Background: Hip fractures in older adults increasingly challenge public health, making traditional rehabilitation very challenging. Digital health interventions (DHIs) have emerged as a promising solution for postoperative rehabilitation. However, evidence on DHIs’ effects on functional and psychological outcomes remains insufficient. Objective: This systematic review aimed to comprehensively examine the effects of DHIs on functional and psychological outcomes in older adults with hip fractures. Methods: Following PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) guidelines, we searched 9 databases (PubMed, Embase, CENTRAL, APA PsycINFO, Web of Science, PEDro, CNKI, WANFANG, and SinoMed) from inception to November 13, 2025. Included studies enrolled adults aged 60 years and older with hip fractures, delivered DHIs, assessed functional and psychological outcomes, set usual care or no intervention as the control, and had a randomized controlled trial design. Studies were excluded if they enrolled nonhospitalized patients in the emergency department, patients discharged to nonhome settings, or had inaccessible full text or insufficient data. Study quality was evaluated using the Cochrane Risk of Bias tool 2.0 (Cochrane Collaboration), and evidence certainty was assessed using GRADE (Grading of Recommendations, Assessment, Development and Evaluation). The literature screening, data extraction, and quality assessment were independently conducted by 2 researchers, and any disputes were resolved by the third researcher. We performed analysis using R version 4.0.3 (R Foundation for Statistical Computing) with a random-effects model. Results: Of 17,723 studies screened, 13 met the inclusion criteria. DHIs, compared to the control, significantly improved hip function (standardized mean difference [SMD] 0.80, 95% CI 0.33-1.26; 95% prediction interval [PI] –0.24 to 1.83; P=.007) and functional independence (SMD 1.23, 95% CI 0.34-2.11; 95% PI –0.98 to 3.34; P=.02). Despite favorable pooled effects, a wide 95% PI spanning positive or negative values signals substantial heterogeneity. No significant difference was observed in balance function, risk of falling, and quality of life. Only a single available study reported a 70% adherence rate in the DHIs group. Subgroup analyses stratified by intervention duration revealed no significant intersubgroup differences for hip function (χ12=0.1; P=.75) or functional independence (χ12=2.93; P=.09). For hip function, the point estimate favored the 3 months subgroup (SMD 0.89, 95% CI 0.36-1.41; I2=7%; P=.41) over the <3 months subgroup. Conversely, for functional independence, the point estimate favored shorter intervention duration (SMD 0.67, 95% CI 0.12-1.23; I²=0%; P=.72). Conclusions: This review incorporates the latest randomized controlled trials and comprehensively assesses functional and psychological outcomes of DHIs in older patients with hip fractures, distinct from prior studies focusing solely on functional outcomes. While the 95% CI supports the potential of DHIs to improve hip function and functional independence, the wide 95% PI indicating substantial real-world response variability, which calls for cautious interpretation, informs the design of targeted DHI-based rehabilitation regimens, warranting further research into optimal techniques and dosages in clinical practice. Trial Registration: PROSPERO CRD42024626186; https://www.crd.york.ac.uk/PROSPERO/view/CRD42024626186

Rel-MOSS: Towards Imbalanced Relational Deep Learning on Relational Databases

arXiv:2603.07916v1 Announce Type: new Abstract: In recent advances, to enable a fully data-driven learning paradigm on relational databases (RDB), relational deep learning (RDL) is proposed to structure the RDB as a heterogeneous entity graph and adopt the graph neural network (GNN) as the predictive model. However, existing RDL methods neglect the imbalance problem of relational data in RDBs and risk under-representing the minority entities, leading to an unusable model in practice. In this work, we investigate, for the first time, class imbalance problem in RDB entity classification and design the relation-centric minority synthetic over-sampling GNN (Rel-MOSS), in order to fill a critical void in the current literature. Specifically, to mitigate the issue of minority-related information being submerged by majority counterparts, we design the relation-wise gating controller to modulate neighborhood messages from each individual relation type. Based on the relational-gated representations, we further propose the relation-guided minority synthesizer for over-sampling, which integrates the entity relational signatures to maintain relational consistency. Extensive experiments on 12 entity classification datasets provide compelling evidence for the superiority of Rel-MOSS, yielding an average improvement of up to 2.46% and 4.00% in terms of Balanced Accuracy and G-Mean, compared with SOTA RDL methods and classic methods for handling class imbalance.

Dexamethasone promotes neutrophil ROS-mediated tumor killing through the glucocorticoid receptor

Oncogene, Published online: 06 March 2026; doi:10.1038/s41388-026-03708-w

Dexamethasone promotes neutrophil ROS-mediated tumor killing through the glucocorticoid receptor

UniG2U-Bench: Do Unified Models Advance Multimodal Understanding?

arXiv:2603.03241v1 Announce Type: cross Abstract: Unified multimodal models have recently demonstrated strong generative capabilities, yet whether and when generation improves understanding remains unclear. Existing benchmarks lack a systematic exploration of the specific tasks where generation facilitates understanding. To this end, we introduce UniG2U-Bench, a comprehensive benchmark categorizing generation-to-understanding (G2U) evaluation into 7 regimes and 30 subtasks, requiring varying degrees of implicit or explicit visual transformations. Extensive evaluation of over 30 models reveals three core findings: 1) Unified models generally underperform their base Vision-Language Models (VLMs), and Generate-then-Answer (GtA) inference typically degrades performance relative to direct inference. 2) Consistent enhancements emerge in spatial intelligence, visual illusions, or multi-round reasoning subtasks, where enhanced spatial and shape perception, as well as multi-step intermediate image states, prove beneficial. 3) Tasks with similar reasoning structures and models sharing architectures exhibit correlated behaviors, suggesting that generation-understanding coupling induces class-consistent inductive biases over tasks, pretraining data, and model architectures. These findings highlight the necessity for more diverse training data and novel paradigms to fully unlock the potential of unified multimodal modeling.
❌