❌

Reading view

ProFashion: Prototype-guided Fashion Video Generation with Multiple Reference Images

arXiv:2505.06537v2 Announce Type: replace-cross Abstract: Fashion video generation aims to synthesize temporally consistent videos from reference images of a designated character. Despite significant progress, existing diffusion-based methods only support a single reference image as input, severely limiting their capability to generate view-consistent fashion videos, especially when there are different patterns on the clothes from different perspectives. Moreover, the widely adopted motion module does not sufficiently model human body movement, leading to sub-optimal spatiotemporal consistency. To address these issues, we propose ProFashion, a fashion video generation framework leveraging multiple reference images to achieve improved view consistency and temporal coherency. To effectively leverage features from multiple reference images while maintaining a reasonable computational cost, we devise a Pose-aware Prototype Aggregator, which selects and aggregates global and fine-grained reference features according to pose information to form frame-wise prototypes, which serve as guidance in the denoising process. To further enhance motion consistency, we introduce a Flow-enhanced Prototype Instantiator, which exploits the human keypoint motion flow to guide an extra spatiotemporal attention process in the denoiser. To demonstrate the effectiveness of ProFashion, we extensively evaluate our method on the MRFashion-7K dataset we collected from the Internet. ProFashion also outperforms previous methods on the UBC Fashion dataset.
  •  

Cerebra: A Multidisciplinary AI Board for Multimodal Dementia Characterization and Risk Assessment

arXiv:2603.21597v2 Announce Type: replace Abstract: Modern clinical practice increasingly depends on reasoning over heterogeneous, evolving, and incomplete patient data. Although recent advances in multimodal foundation models have improved performance on various clinical tasks, most existing models remain static, opaque, and poorly aligned with real-world clinical workflows. We present Cerebra, an interactive multi-agent AI team that coordinates specialized agents for EHR, clinical notes, and medical imaging analysis. These outputs are synthesized into a clinician-facing dashboard that combines visual analytics with a conversational interface, enabling clinicians to interrogate predictions and contextualize risk at the point of care. Cerebra supports privacy-preserving deployment by operating on structured representations and remains robust when modalities are incomplete. We evaluated Cerebra using a massive multi-institutional dataset spanning 3 million patients from four independent healthcare systems. Cerebra consistently outperformed both state-of-the-art single-modality models and large multimodal language model baselines. In dementia risk prediction, it achieved AUROCs up to 0.80, compared with 0.74 for the strongest single-modality model and 0.68 for language model baselines. For dementia diagnosis, it achieved an AUROC of 0.86, and for survival prediction, a C-index of 0.81. In a reader study with experienced physicians, Cerebra significantly improved expert performance, increasing accuracy by 17.5 percentage points in prospective dementia risk estimation. These results demonstrate Cerebra's potential for interpretable, robust decision support in clinical care.
  •  

OmniDiT: Extending Diffusion Transformer to Omni-VTON Framework

arXiv:2603.19643v2 Announce Type: replace-cross Abstract: Despite the rapid advancement of Virtual Try-On (VTON) and Try-Off (VTOFF) technologies, existing VTON methods face challenges with fine-grained detail preservation, generalization to complex scenes, complicated pipeline, and efficient inference. To tackle these problems, we propose OmniDiT, an omni Virtual Try-On framework based on the Diffusion Transformer, which combines try-on and try-off tasks into one unified model. Specifically, we first establish a self-evolving data curation pipeline to continuously produce data, and construct a large VTON dataset Omni-TryOn, which contains over 380k diverse and high-quality garment-model-tryon image pairs and detailed text prompts. Then, we employ the token concatenation and design an adaptive position encoding to effectively incorporate multiple reference conditions. To relieve the bottleneck of long sequence computation, we are the first to introduce Shifted Window Attention into the diffusion model, thus achieving a linear complexity. To remedy the performance degradation caused by local window attention, we utilize multiple timestep prediction and an alignment loss to improve generation fidelity. Experiments reveal that, under various complex scenes, our method achieves the best performance in both the model-free VTON and VTOFF tasks and a performance comparable to current SOTA methods in the model-based VTON task.
  •  

Research on the compatibility mechanism of the Tingli Dazao Xiefei Decoction by multi-organ metabolomics strategy

J Ethnopharmacol. 2026 Mar 21:121548. doi: 10.1016/j.jep.2026.121548. Online ahead of print.

ABSTRACT

ETHNOPHARMACOLOGICAL RELEVANCE: The Tingli Dazao Xiefei Decoction (TD) is a traditional phlegm-eliminating prescription composed of Descurainia sophia (L.) Webb. ex Prantl (TLZ) and Ziziphus jujuba Mill. (DZ), which can relieve lung, heart and kidney injury in asthma. TLZ acts as the monarch drug in the TD. Based on the research mode of "material basis of traditional Chinese medicinal properties can be divided and combined", we have confirmed that the flavonoid glycosides components /the oligosaccharide components/the fatty oil component (FG/Oli/FO) are effective components of TLZ. However, the compatibility mechanism of the TD, and the contribution of the effective components of TLZ to the efficacy were still unclear.

AIM OF THE STUDY: To clarify the compatibility mechanism of TD, and the contribution of the effective components of TLZ to the efficacy from a comprehensive perspective of lung, heart, and kidney.

METHODS: First, we chose the asthma model corresponding to the efficacy of TD in purging the lungs and relieving asthma, and the rats were divided into the normal (NC) group, model (M) group, dexamethasone (DEX) group, and treatment groups of TD/TLZ/DZ/FO+DZ/Oli+DZ/FG+DZ. Second, metabolomics and network pharmacology were applied to elucidate the comprehensive protective effect of TD/FG+DZ/Oli+DZ/FO+DZ. Third, the multi-omics results were validated using Western blotting, RT-qPCR, flow cytometry, and immunofluorescence.

RESULTS: FO+DZ/Oli+DZ/FG+DZ had different degrees of protective effects against lung/heart/kidney injury in asthma. In metabolomics research, the principal component analysis (PCA) and cluster analysis results showed that the TLZ group was closer to TD group than DZ group, the FO+DZ and Oli+DZ group clustered with TD/NC groups in the lung and kidney, and the FO+DZ and FG+DZ group clustered with TD/NC groups in the heart. Pathway enrichment analysis suggested that the comprehensive protective effect of TLZ and its effective components combined with DZ on lung/heart/kidney may be achieved by regulating the arginine and proline metabolism, alanine, aspartate and glutamate metabolism, and unsaturated fatty acid biosynthesis. Multi-organ metabolomics and network pharmacology revealed consistent biological functions in KEGG pathways. Validation experiment showed that TLZ and its effective components combined with DZ could reverse the abnormal expression of proteins and RNA related to inflammation, airway remodeling, excitotoxicity, and energy-supply, apoptosis at different levels. Furthermore, FO+DZ may reduce asthma damage by inhibiting the FABP4/PPAR-γ/NF-κB signaling pathway.

CONCLUSION: TLZ played the key role in TD, and FO had the best therapeutic effect on each organ; the efficacy of Oli was mainly reflected in reducing lung and kidney damage, and FG was mainly involved in enhancing energy metabolism in the heart. These findings proved that traditional Chinese medicine could exert comprehensive efficacy in a 'multi-components trigger multi-channel' way.

PMID:41871629 | DOI:10.1016/j.jep.2026.121548

  •  

Research on the compatibility mechanism of the Tingli Dazao Xiefei Decoction by multi-organ metabolomics strategy

J Ethnopharmacol. 2026 Mar 21:121548. doi: 10.1016/j.jep.2026.121548. Online ahead of print.

ABSTRACT

ETHNOPHARMACOLOGICAL RELEVANCE: The Tingli Dazao Xiefei Decoction (TD) is a traditional phlegm-eliminating prescription composed of Descurainia sophia (L.) Webb. ex Prantl (TLZ) and Ziziphus jujuba Mill. (DZ), which can relieve lung, heart and kidney injury in asthma. TLZ acts as the monarch drug in the TD. Based on the research mode of "material basis of traditional Chinese medicinal properties can be divided and combined", we have confirmed that the flavonoid glycosides components /the oligosaccharide components/the fatty oil component (FG/Oli/FO) are effective components of TLZ. However, the compatibility mechanism of the TD, and the contribution of the effective components of TLZ to the efficacy were still unclear.

AIM OF THE STUDY: To clarify the compatibility mechanism of TD, and the contribution of the effective components of TLZ to the efficacy from a comprehensive perspective of lung, heart, and kidney.

METHODS: First, we chose the asthma model corresponding to the efficacy of TD in purging the lungs and relieving asthma, and the rats were divided into the normal (NC) group, model (M) group, dexamethasone (DEX) group, and treatment groups of TD/TLZ/DZ/FO+DZ/Oli+DZ/FG+DZ. Second, metabolomics and network pharmacology were applied to elucidate the comprehensive protective effect of TD/FG+DZ/Oli+DZ/FO+DZ. Third, the multi-omics results were validated using Western blotting, RT-qPCR, flow cytometry, and immunofluorescence.

RESULTS: FO+DZ/Oli+DZ/FG+DZ had different degrees of protective effects against lung/heart/kidney injury in asthma. In metabolomics research, the principal component analysis (PCA) and cluster analysis results showed that the TLZ group was closer to TD group than DZ group, the FO+DZ and Oli+DZ group clustered with TD/NC groups in the lung and kidney, and the FO+DZ and FG+DZ group clustered with TD/NC groups in the heart. Pathway enrichment analysis suggested that the comprehensive protective effect of TLZ and its effective components combined with DZ on lung/heart/kidney may be achieved by regulating the arginine and proline metabolism, alanine, aspartate and glutamate metabolism, and unsaturated fatty acid biosynthesis. Multi-organ metabolomics and network pharmacology revealed consistent biological functions in KEGG pathways. Validation experiment showed that TLZ and its effective components combined with DZ could reverse the abnormal expression of proteins and RNA related to inflammation, airway remodeling, excitotoxicity, and energy-supply, apoptosis at different levels. Furthermore, FO+DZ may reduce asthma damage by inhibiting the FABP4/PPAR-γ/NF-κB signaling pathway.

CONCLUSION: TLZ played the key role in TD, and FO had the best therapeutic effect on each organ; the efficacy of Oli was mainly reflected in reducing lung and kidney damage, and FG was mainly involved in enhancing energy metabolism in the heart. These findings proved that traditional Chinese medicine could exert comprehensive efficacy in a 'multi-components trigger multi-channel' way.

PMID:41871629 | DOI:10.1016/j.jep.2026.121548

  •  

MoKus: Leveraging Cross-Modal Knowledge Transfer for Knowledge-Aware Concept Customization

arXiv:2603.12743v1 Announce Type: cross Abstract: Concept customization typically binds rare tokens to a target concept. Unfortunately, these approaches often suffer from unstable performance as the pretraining data seldom contains these rare tokens. Meanwhile, these rare tokens fail to convey the inherent knowledge of the target concept. Consequently, we introduce Knowledge-aware Concept Customization, a novel task aiming at binding diverse textual knowledge to target visual concepts. This task requires the model to identify the knowledge within the text prompt to perform high-fidelity customized generation. Meanwhile, the model should efficiently bind all the textual knowledge to the target concept. Therefore, we propose MoKus, a novel framework for knowledge-aware concept customization. Our framework relies on a key observation: cross-modal knowledge transfer, where modifying knowledge within the text modality naturally transfers to the visual modality during generation. Inspired by this observation, MoKus contains two stages: (1) In visual concept learning, we first learn the anchor representation to store the visual information of the target concept. (2) In textual knowledge updating, we update the answer for the knowledge queries to the anchor representation, enabling high-fidelity customized generation. To further comprehensively evaluate our proposed MoKus on the new task, we introduce the first benchmark for knowledge-aware concept customization: KnowCusBench. Extensive evaluations have demonstrated that MoKus outperforms state-of-the-art methods. Moreover, the cross-model knowledge transfer allows MoKus to be easily extended to other knowledge-aware applications like virtual concept creation and concept erasure. We also demonstrate the capability of our method to achieve improvements on world knowledge benchmarks.
  •  

SketchGraphNet: A Memory-Efficient Hybrid Graph Transformer for Large-Scale Sketch Corpora Recognition

arXiv:2603.07521v1 Announce Type: cross Abstract: This work investigates large-scale sketch recognition from a graph-native perspective, where free-hand sketches are directly modeled as structured graphs rather than raster images or stroke sequences. We propose SketchGraphNet, a hybrid graph neural architecture that integrates local message passing with a memory-efficient global attention mechanism, without relying on auxiliary positional or structural encodings. To support systematic evaluation, we construct SketchGraph, a large-scale benchmark comprising 3.44 million graph-structured sketches across 344 categories, with two variants (A and R) to reflect different noise conditions. Each sketch is represented as a spatiotemporal graph with normalized stroke-order attributes. On SketchGraph-A and SketchGraph-R, SketchGraphNet achieves Top-1 accuracies of 83.62% and 87.61%, respectively, under a unified training configuration. MemEffAttn further reduces peak GPU memory by over 40% and training time by more than 30% compared with Performer-based global attention, while maintaining comparable accuracy.
  •  

SFIBA: Spatial-based Full-target Invisible Backdoor Attacks

arXiv:2504.21052v2 Announce Type: replace-cross Abstract: Multi-target backdoor attacks pose significant security threats to deep neural networks, as they can preset multiple target classes through a single backdoor injection. This allows attackers to control the model to misclassify poisoned samples with triggers into any desired target class during inference, exhibiting superior attack performance compared with conventional backdoor attacks. However, existing multi-target backdoor attacks fail to guarantee trigger specificity and stealthiness in black-box settings, resulting in two main issues. First, they are unable to simultaneously target all classes when only training data can be manipulated, limiting their effectiveness in realistic attack scenarios. Second, the triggers often lack visual imperceptibility, making poisoned samples easy to detect. To address these problems, we propose a Spatial-based Full-target Invisible Backdoor Attack, called SFIBA. It restricts triggers for different classes to specific local spatial regions and morphologies in the pixel space to ensure specificity, while employing a frequency-domain-based trigger injection method to guarantee stealthiness. Specifically, for injection of each trigger, we first apply fast fourier transform to obtain the amplitude spectrum of clean samples in local spatial regions. Then, we employ discrete wavelet transform to extract the features from the amplitude spectrum and use singular value decomposition to integrate the trigger. Subsequently, we selectively filter parts of the trigger in pixel space to implement trigger morphology constraints and adjust injection coefficients based on visual effects. We conduct experiments on multiple datasets and models. The results demonstrate that SFIBA can achieve excellent attack performance and stealthiness, while preserving the model's performance on benign samples, and can also bypass existing backdoor defenses.
  •  

Farther the Shift, Sparser the Representation: Analyzing OOD Mechanisms in LLMs

arXiv:2603.03415v1 Announce Type: cross Abstract: In this work, we investigate how Large Language Models (LLMs) adapt their internal representations when encountering inputs of increasing difficulty, quantified as the degree of out-of-distribution (OOD) shift. We reveal a consistent and quantifiable phenomenon: as task difficulty increases, whether through harder reasoning questions, longer contexts, or adding answer choices, the last hidden states of LLMs become substantially sparser. In short, \textbf{\textit{the farther the shift, the sparser the representations}}. This sparsity--difficulty relation is observable across diverse models and domains, suggesting that language models respond to unfamiliar or complex inputs by concentrating computation into specialized subspaces in the last hidden state. Through a series of controlled analyses with a learning dynamic explanation, we demonstrate that this sparsity is not incidental but an adaptive mechanism for stabilizing reasoning under OOD. Leveraging this insight, we design \textit{Sparsity-Guided Curriculum In-Context Learning (SG-ICL)}, a strategy that explicitly uses representation sparsity to schedule few-shot demonstrations, leading to considerable performance enhancements. Our study provides new mechanistic insights into how LLMs internalize OOD challenges. The source code is available at the URL: https://github.com/MingyuJ666/sparsityLLM.
  •  

SWE-CI: Evaluating Agent Capabilities in Maintaining Codebases via Continuous Integration

arXiv:2603.03823v1 Announce Type: cross Abstract: Large language model (LLM)-powered agents have demonstrated strong capabilities in automating software engineering tasks such as static bug fixing, as evidenced by benchmarks like SWE-bench. However, in the real world, the development of mature software is typically predicated on complex requirement changes and long-term feature iterations -- a process that static, one-shot repair paradigms fail to capture. To bridge this gap, we propose \textbf{SWE-CI}, the first repository-level benchmark built upon the Continuous Integration loop, aiming to shift the evaluation paradigm for code generation from static, short-term \textit{functional correctness} toward dynamic, long-term \textit{maintainability}. The benchmark comprises 100 tasks, each corresponding on average to an evolution history spanning 233 days and 71 consecutive commits in a real-world code repository. SWE-CI requires agents to systematically resolve these tasks through dozens of rounds of analysis and coding iterations. SWE-CI provides valuable insights into how well agents can sustain code quality throughout long-term evolution.
  •  

Can Multimodal LLMs See Science Instruction? Benchmarking Pedagogical Reasoning in K-12 Classroom Videos

arXiv:2602.18466v1 Announce Type: cross Abstract: K-12 science classrooms are rich sites of inquiry where students coordinate phenomena, evidence, and explanatory models through discourse; yet, the multimodal complexity of these interactions has made automated analysis elusive. Existing benchmarks for classroom discourse focus primarily on mathematics and rely solely on transcripts, overlooking the visual artifacts and model-based reasoning emphasized by the Next Generation Science Standards (NGSS). We address this gap with SciIBI, the first video benchmark for analyzing science classroom discourse, featuring 113 NGSS-aligned clips annotated with Core Instructional Practices (CIP) and sophistication levels. By evaluating eight state-of-the-art LLMs and Multimodal LLMs, we reveal fundamental limitations: current models struggle to distinguish pedagogically similar practices, suggesting that CIP coding requires instructional reasoning beyond surface pattern matching. Furthermore, adding video input yields inconsistent gains across architectures. Crucially, our evidence-based evaluation reveals that models often succeed through surface shortcuts rather than genuine pedagogical understanding. These findings establish science classroom discourse as a challenging frontier for multimodal AI and point toward human-AI collaboration, where models retrieve evidence to accelerate expert review rather than replace it.
  •  

(PASS) Visual Prompt Locates Good Structure Sparsity through a Recurrent HyperNetwork

arXiv:2407.17412v2 Announce Type: replace-cross Abstract: Large-scale neural networks have demonstrated remarkable performance in different domains like vision and language processing, although at the cost of massive computation resources. As illustrated by compression literature, structural model pruning is a prominent algorithm to encourage model efficiency, thanks to its acceleration-friendly sparsity patterns. One of the key questions of structural pruning is how to estimate the channel significance. In parallel, work on data-centric AI has shown that prompting-based techniques enable impressive generalization of large language models across diverse downstream tasks. In this paper, we investigate a charming possibility - \textit{leveraging visual prompts to capture the channel importance and derive high-quality structural sparsity}. To this end, we propose a novel algorithmic framework, namely \texttt{PASS}. It is a tailored hyper-network to take both visual prompts and network weight statistics as input, and output layer-wise channel sparsity in a recurrent manner. Such designs consider the intrinsic channel dependency between layers. Comprehensive experiments across multiple network architectures and six datasets demonstrate the superiority of \texttt{PASS} in locating good structural sparsity. For example, at the same FLOPs level, \texttt{PASS} subnetworks achieve $1\%\sim 3\%$ better accuracy on Food101 dataset; or with a similar performance of $80\%$ accuracy, \texttt{PASS} subnetworks obtain $0.35\times$ more speedup than the baselines.
  •  

Dialogue is Better Than Monologue: Instructing Medical LLMs via Strategical Conversations

arXiv:2501.17860v2 Announce Type: replace-cross Abstract: Current medical AI systems often fail to replicate real-world clinical reasoning, as they are predominantly trained and evaluated on static text and question-answer tasks. These tuning methods and benchmarks overlook critical aspects like evidence-based reasoning and handling distracting information. To bridge this gap, we introduce a novel benchmark that simulates real-world diagnostic scenarios, integrating noise and difficulty levels aligned with USMLE standards. Moreover, we explore dialogue-based fine-tuning, which transforms static datasets into conversational formats to better capture iterative reasoning processes. Experiments show that dialogue-tuned models outperform traditional methods, with improvements of $9.64\%$ in multi-round reasoning scenarios and $6.18\%$ in accuracy in a noisy environment. Our findings highlight dialogue tuning as a promising approach for advancing clinically aligned and robust medical AI systems.
  •  

SPATIA: Multimodal Generation and Prediction of Spatial Cell Phenotypes

arXiv:2507.04704v2 Announce Type: replace-cross Abstract: Understanding how cellular morphology, gene expression, and spatial context jointly shape tissue function is a central challenge in biology. Image-based spatial transcriptomics technologies now provide high-resolution measurements of cell images and gene expression profiles, but existing methods typically analyze these modalities in isolation or at limited resolution. We address the problem by introducing SPATIA, a multi-level generative and predictive model that learns unified, spatially aware representations by fusing morphology, gene expression, and spatial context from the cell to the tissue level. SPATIA also incorporates a novel spatially conditioned generative framework for predicting cell morphologies under perturbations. Specifically, we propose a confidence-aware flow matching objective that reweights weak optimal-transport pairs based on uncertainty. We further apply morphology-profile alignment to encourage biologically meaningful image generation, enabling the modeling of microenvironment-dependent phenotypic transitions. We assembled a multi-scale dataset consisting of 25.9 million cell-gene pairs across 17 tissues. We benchmark SPATIA against 18 models across 12 tasks, spanning categories such as phenotype generation, annotation, clustering, gene imputation, and cross-modal prediction. SPATIA achieves improved performance over state-of-the-art models, improving generative fidelity by 8% and predictive accuracy by up to 3%.
  •  

ShotFinder: Imagination-Driven Open-Domain Video Shot Retrieval via Web Search

arXiv:2601.23232v3 Announce Type: replace-cross Abstract: In recent years, large language models (LLMs) have made rapid progress in information retrieval, yet existing research has mainly focused on text or static multimodal settings. Open-domain video shot retrieval, which involves richer temporal structure and more complex semantics, still lacks systematic benchmarks and analysis. To fill this gap, we introduce ShotFinder, a benchmark that formalizes editing requirements as keyframe-oriented shot descriptions and introduces five types of controllable single-factor constraints: Temporal order, Color, Visual style, Audio, and Resolution. We curate 1,210 high-quality samples from YouTube across 20 thematic categories, using large models for generation with human verification. Based on the benchmark, we propose ShotFinder, a text-driven three-stage retrieval and localization pipeline: (1) query expansion via video imagination, (2) candidate video retrieval with a search engine, and (3) description-guided temporal localization. Experiments on multiple closed-source and open-source models reveal a significant gap to human performance, with clear imbalance across constraints: temporal localization is relatively tractable, while color and visual style remain major challenges. These results reveal that open-domain video shot retrieval is still a critical capability that multimodal large models have yet to overcome.
  •  
❌