❌

Normal view

M3D-BFS: a Multi-stage Dynamic Fusion Strategy for Sample-Adaptive Multi-Modal Brain Network Analysis

arXiv:2604.01667v1 Announce Type: new Abstract: Multi-modal fusion is of great significance in neuroscience which integrates information from different modalities and can achieve better performance than uni-modal methods in downstream tasks. Current multi-modal fusion methods in brain networks, which mainly focus on structural connectivity (SC) and functional connectivity (FC) modalities, are static in nature. They feed different samples into the same model with identical computation, ignoring inherent difference between input samples. This lack of sample adaptation hinders model's further performance. To this end, we innovatively propose a multi-stage dynamic fusion strategy (M3D-BFS) for sample-adaptive multi-modal brain network analysis. Unlike other static fusion methods, we design different mixture-of-experts (MoEs) for uni- and multi-modal representations where modules can adaptively change as input sample changes during inference. To alleviate issue of MoE where training of experts may be collapsed, we divide our method into 3 stages. We first train uni-modal encoders respectively, then pretrain single experts of MoEs before finally finetuning the whole model. A multi-modal disentanglement loss is designed to enhance the final representations. To the best of our knowledge, this is the first work for dynamic fusion for multi-modal brain network analysis. Extensive experiments on different real-world datasets demonstrates the superiority of M3D-BFS.

Integrative Multi-omics Analysis of Buti Huatan Tang in Chronic Obstructive Pulmonary Disease

J Vis Exp. 2026 Mar 13;(229). doi: 10.3791/70383.

ABSTRACT

This study utilized a multi-omics and computational biology framework to investigate the therapeutic potential of the Traditional Chinese Medicine (TCM) formula Buti Huatan Tang (BTHTT) against chronic obstructive pulmonary disease (COPD). Significant physiological improvements were observed in a rat model following BTHTT intervention. Histological analysis showed a reversal of lung pathological damage, while biochemical assays, and transcriptomics confirmed the normalization of IL-1β and IL-1R2 levels. Additionally, metabolic profiling revealed that BTHTT corrected disruptions in T3 and T4 thyroid hormone levels. A negative correlation was observed between the IL-1β/IL-1R2 axis and these thyroid hormones, indicating that their regulation is associated with the formula's therapeutic effect. Beyond direct measurements, machine learning algorithms identified ten COPD signature genes from clinical databases. Pathway enrichment analysis suggests that BTHTT may act through cytokine-cytokine-receptor interactions and thyroid hormone synthesis pathways. Furthermore, while 283 components were identified in vivo, compounds such as tanshinone IIA and cryptotanshinone are currently considered candidate active substances. Their role as primary drivers is supported by a model in which they stably bind to IL-1R2; this inference is based on molecular docking and molecular dynamics (MD) simulations rather than direct experimental isolation. Overall, the data support a model in which BTHTT exerts a multi-target effect on COPD by modulating inflammation and metabolic homeostasis. This integrated approach provides a refined scientific basis for the clinical application of BTHTT and highlights specific pathways for future experimental validation.

PMID:41911070 | DOI:10.3791/70383

Integrative Multi-omics Analysis of Buti Huatan Tang in Chronic Obstructive Pulmonary Disease

J Vis Exp. 2026 Mar 13;(229). doi: 10.3791/70383.

ABSTRACT

This study utilized a multi-omics and computational biology framework to investigate the therapeutic potential of the Traditional Chinese Medicine (TCM) formula Buti Huatan Tang (BTHTT) against chronic obstructive pulmonary disease (COPD). Significant physiological improvements were observed in a rat model following BTHTT intervention. Histological analysis showed a reversal of lung pathological damage, while biochemical assays, and transcriptomics confirmed the normalization of IL-1β and IL-1R2 levels. Additionally, metabolic profiling revealed that BTHTT corrected disruptions in T3 and T4 thyroid hormone levels. A negative correlation was observed between the IL-1β/IL-1R2 axis and these thyroid hormones, indicating that their regulation is associated with the formula's therapeutic effect. Beyond direct measurements, machine learning algorithms identified ten COPD signature genes from clinical databases. Pathway enrichment analysis suggests that BTHTT may act through cytokine-cytokine-receptor interactions and thyroid hormone synthesis pathways. Furthermore, while 283 components were identified in vivo, compounds such as tanshinone IIA and cryptotanshinone are currently considered candidate active substances. Their role as primary drivers is supported by a model in which they stably bind to IL-1R2; this inference is based on molecular docking and molecular dynamics (MD) simulations rather than direct experimental isolation. Overall, the data support a model in which BTHTT exerts a multi-target effect on COPD by modulating inflammation and metabolic homeostasis. This integrated approach provides a refined scientific basis for the clinical application of BTHTT and highlights specific pathways for future experimental validation.

PMID:41911070 | DOI:10.3791/70383

From Editor to Dense Geometry Estimator

arXiv:2509.04338v2 Announce Type: replace-cross Abstract: Leveraging visual priors from pre-trained text-to-image (T2I) generative models has shown success in dense prediction. However, dense prediction is inherently an image-to-image task, suggesting that image editing models, rather than T2I generative models, may be a more suitable foundation for fine-tuning. Motivated by this, we conduct a systematic analysis of the fine-tuning behaviors of both editors and generators for dense geometry estimation. Our findings show that editing models possess inherent structural priors, which enable them to converge more stably by ``refining" their innate features, and ultimately achieve higher performance than their generative counterparts. Based on these findings, we introduce \textbf{FE2E}, a framework that pioneeringly adapts an advanced editing model based on Diffusion Transformer (DiT) architecture for dense geometry prediction. Specifically, to tailor the editor for this deterministic task, we reformulate the editor's original flow matching loss into the ``consistent velocity" training objective. And we use logarithmic quantization to resolve the precision conflict between the editor's native BFloat16 format and the high precision demand of our tasks. Additionally, we leverage the DiT's global attention for a cost-free joint estimation of depth and normals in a single forward pass, enabling their supervisory signals to mutually enhance each other. Without scaling up the training data, FE2E achieves impressive performance improvements in zero-shot monocular depth and normal estimation across multiple datasets. Notably, it achieves over 35\% performance gains on the ETH3D dataset and outperforms the DepthAnything series, which is trained on 100$\times$ data. The project page can be accessed \href{https://amap-ml.github.io/FE2E/}{here}.

<i>LRRK2</i>-targeting antisense oligonucleotide in Parkinson’s disease: a phase 1 randomized controlled trial

Nature Medicine, Published online: 24 March 2026; doi:10.1038/s41591-026-04262-4

The first-in-human clinical trial of the LRRK2-targeting antisense oligonucleotide BIIB094 in Parkinson’s disease demonstrates that the treatment is well tolerated and produces dose-dependent reductions in cerebrospinal fluid levels of LRRK2 and phosphorylated Rab10, indicating successful target engagement.

Spatial Omics in Gastrointestinal Oncology: Recent Advances, Therapeutic Insights, and Clinical Translation

J Cancer. 2026 Jan 30;17(3):515-523. doi: 10.7150/jca.127381. eCollection 2026.

ABSTRACT

Gastrointestinal (GI) cancers remain a leading cause of cancer-related morbidity and mortality worldwide, largely due to their molecular heterogeneity, complex tumor microenvironment (TME), and variable treatment responses. In recent years, the emergence of spatially resolved omics technologies-encompassing spatial transcriptomics, proteomics, metabolomics, and epigenomics-has revolutionized the ability to interrogate tumor architecture with unprecedented resolution. These methods enable precise mapping of cellular and molecular interactions within intact tissue contexts, thereby uncovering spatially defined niches that influence tumor progression, immune evasion, and therapeutic resistance. In GI malignancies such as colorectal, gastric, and esophageal cancers, spatial omics have provided critical insights into cancer-stromal-immune crosstalk, identified predictive biomarkers for immunotherapy and targeted agents, and guided the development of novel therapeutic strategies. This review synthesizes the latest advances in spatial omics applied to GI oncology over the past five years, with an emphasis on their integration into early diagnosis, treatment stratification, and real-time monitoring of therapeutic efficacy. We also discuss current challenges, including standardization, data integration, and clinical validation, as well as future directions for incorporating spatial profiling into routine oncology practice. By bridging the gap between bench discoveries and bedside applications, spatial omics hold transformative potential for achieving truly personalized treatment in gastrointestinal cancers.

PMID:41869445 | PMC:PMC13003551 | DOI:10.7150/jca.127381

Towards unified brain-to-text decoding across speech production and perception

arXiv:2603.12628v1 Announce Type: new Abstract: Speech production and perception are the main ways humans communicate daily. Prior brain-to-text decoding studies have largely focused on a single modality and alphabetic languages. Here, we present a unified brain-to-sentence decoding framework for both speech production and perception in Mandarin Chinese. The framework exhibits strong generalization ability, enabling sentence-level decoding when trained only on single-character data and supporting characters and syllables unseen during training. In addition, it allows direct and controlled comparison of neural dynamics across modalities. Mandarin speech is decoded by first classifying syllable components in Hanyu Pinyin, namely initials and finals, from neural signals, followed by a post-trained large language model (LLM) that maps sequences of toneless Pinyin syllables to Chinese sentences. To enhance LLM decoding, we designed a three-stage post-training and two-stage inference framework based on a 7-billion-parameter LLM, achieving overall performance that exceeds larger commercial LLMs with hundreds of billions of parameters or more. In addition, several characteristics were observed in Mandarin speech production and perception: speech production involved neural responses across broader cortical regions than auditory perception; channels responsive to both modalities exhibited similar activity patterns, with speech perception showing a temporal delay relative to production; and decoding performance was broadly comparable across hemispheres. Our work not only establishes the feasibility of a unified decoding framework but also provides insights into the neural characteristics of Mandarin speech production and perception. These advances contribute to brain-to-text decoding in logosyllabic languages and pave the way toward neural language decoding systems supporting multiple modalities.

FedBPrompt: Federated Domain Generalization Person Re-Identification via Body Distribution Aware Visual Prompts

arXiv:2603.12912v1 Announce Type: cross Abstract: Federated Domain Generalization for Person Re-Identification (FedDG-ReID) learns domain-invariant representations from decentralized data. While Vision Transformer (ViT) is widely adopted, its global attention often fails to distinguish pedestrians from high similarity backgrounds or diverse viewpoints -- a challenge amplified by cross-client distribution shifts in FedDG-ReID. To address this, we propose Federated Body Distribution Aware Visual Prompt (FedBPrompt), introducing learnable visual prompts to guide Transformer attention toward pedestrian-centric regions. FedBPrompt employs a Body Distribution Aware Visual Prompts Mechanism (BAPM) comprising: Holistic Full Body Prompts to suppress cross-client background noise, and Body Part Alignment Prompts to capture fine-grained details robust to pose and viewpoint variations. To mitigate high communication costs, we design a Prompt-based Fine-Tuning Strategy (PFTS) that freezes the ViT backbone and updates only lightweight prompts, significantly reducing communication overhead while maintaining adaptability. Extensive experiments demonstrate that BAPM effectively enhances feature discrimination and cross-domain generalization, while PFTS achieves notable performance gains within only a few aggregation rounds. Moreover, both BAPM and PFTS can be easily integrated into existing ViT-based FedDG-ReID frameworks, making FedBPrompt a flexible and effective solution for federated person re-identification. The code is available at https://github.com/leavlong/FedBPrompt.

Breast Cancer Screening Knowledge and Sentiments in Singaporean Women: Mixed Methods Study Using Topic Modeling, Sentiment Analysis, and Structured Questionnaire Data

Background: Mammography screening uptake in Singapore remains below 40% despite campaigns and subsidies. Natural language processing (NLP) can extract nuanced attitudes from free text that fixed response options miss, revealing latent factors influencing breast cancer (BC) screening behavior. Objective: This study characterized women’s attitudes toward mammography using mixed methods data, examined associations between BC awareness and screening willingness, and identified barriers and facilitators through NLP of free-text responses. Methods: We conducted a cross-sectional study within the multicenter cohort in Singapore (October 2021-December 2023). In total, 4169 women aged 35‐59 years (median 48, IQR 43‐54) were recruited via convenience sampling (3 hospitals and 2 polyclinics). Participants completed online structured questionnaires on demographics and screening history, then a BC education quiz with feedback. Participants answering >80% correctly were classified as “BC-aware.” Posteducation, participants reported screening willingness (motivated or neutral) with optional free-text explanations. Logistic regression models (adjusted for study site, age, ethnicity, marital status, housing, and education) examined the associations with willingness. For 3819 English-language respondents, biterm topic modeling identified themes and sentiment analysis quantified emotional tone. Statistical significance: =.05. Results: Overall, 79% (3287/4169) were BC-aware, and 94% (3908/4169) reported increased motivation posteducation. BC-aware women had higher screening motivation than BC-unaware women (adjusted odds ratio [aOR] 2.88, 95% CI 2.19‐3.80;

RLJP: Legal Judgment Prediction via First-Order Logic Rule-enhanced with Large Language Models

arXiv:2505.21281v2 Announce Type: replace Abstract: Legal Judgment Prediction (LJP) is a pivotal task in legal AI. Existing semantic-enhanced LJP models integrate judicial precedents and legal knowledge for high performance. But they neglect legal reasoning logic, a critical component of legal judgments requiring rigorous logical analysis. Although some approaches utilize legal reasoning logic for high-quality predictions, their logic rigidity hinders adaptation to case-specific logical frameworks, particularly in complex cases that are lengthy and detailed. This paper proposes a rule-enhanced legal judgment prediction framework based on first-order logic (FOL) formalism and comparative learning (CL) to develop an adaptive adjustment mechanism for legal judgment logic and further enhance performance in LJP. Inspired by the process of human exam preparation, our method follows a three-stage approach: first, we initialize judgment rules using the FOL formalism to capture complex reasoning logic accurately; next, we propose a Confusion-aware Contrastive Learning (CACL) to dynamically optimize the judgment rules through a quiz consisting of confusable cases; finally, we utilize the optimized judgment rules to predict legal judgments. Experimental results on two public datasets show superior performance across all metrics. The code is publicly available{https://anonymous.4open.science/r/RLJP-FDF1}.

R1-Code-Interpreter: LLMs Reason with Code via Supervised and Multi-stage Reinforcement Learning

arXiv:2505.21668v3 Announce Type: replace Abstract: Practical guidance on training Large Language Models (LLMs) to leverage Code Interpreter across diverse tasks remains lacking. We present R1-Code-Interpreter, an extension of a text-only LLM trained via multi-turn supervised fine-tuning (SFT) and reinforcement learning (RL) to autonomously generate multiple code queries during step-by-step reasoning. Unlike prior RL + tool-use efforts focused on narrow domains such as math or retrieval, we curate 144 diverse reasoning and planning tasks and show that training a general-purpose Code Interpreter across them presents significant challenges due to task heterogeneity and scarcity of effective samples. To address this, we introduce a multi-stage curriculum learning approach that partitions training samples by measured improvement potential. The RL training prioritizes samples with higher potential and gradually shifts to lower-potential ones, increasing the average RL gains from merely +3.4% to +9.3% across Qwen-2.5 models (3/7/14B). Our final model, R1-CI-14B, improves average accuracy on the 37 test tasks from 44.1% to 72.4%, outperforming text-only GPT-4o (58.6%) and GPT-4o with Code Interpreter (70.9%). Notably, R1-CI-14B also exhibits emergent self-checking behavior through code generation. Datasets, Codes, and Models are available at https://github.com/yongchao98/R1-Code-Interpreter and https://huggingface.co/yongchao98.

From Static Spectra to Operando Infrared Dynamics: Physics Informed Flow Modeling and a Benchmark

arXiv:2602.18551v1 Announce Type: cross Abstract: The Solid Electrolyte Interphase (SEI) is critical to the performance of lithium-ion batteries, yet its analysis via Operando Infrared (IR) spectroscopy remains experimentally complex and expensive, which limits its accessibility for standard research facilities. To overcome this bottleneck, we formulate a novel task, Operando IR Prediction, which aims to forecast the time-resolved evolution of spectral ``fingerprints'' from a single static spectrum. To facilitate this, we introduce OpIRSpec-7K, the first large-scale operando dataset comprising 7,118 high-quality samples across 10 distinct battery systems, alongside OpIRBench, a comprehensive evaluation benchmark with carefully designed protocols. Addressing the limitations of standard spectrum, video, and sequence models in capturing voltage-driven chemical dynamics and complex composition, we propose Aligned Bi-stream Chemical Constraint (ABCC), an end-to-end physics-aware framework. It reformulates MeanFlow and introduces a novel Chemical Flow to explicitly model reaction trajectories, employs a two-stream disentanglement mechanism for solvent-SEI separation, and enforces physics and spectrum constraints such as mass conservation and peak shifts. ABCC significantly outperforms state-of-the-art static, sequential, and generative baselines. ABCC even generalizes to unseen systems and enables interpretable downstream recovery of SEI formation pathways, supporting AI-driven electrochemical discovery.

Taming Preconditioner Drift: Unlocking the Potential of Second-Order Optimizers for Federated Learning on Non-IID Data

arXiv:2602.19271v1 Announce Type: cross Abstract: Second-order optimizers can significantly accelerate large-scale training, yet their naive federated variants are often unstable or even diverge on non-IID data. We show that a key culprit is \emph{preconditioner drift}: client-side second-order training induces heterogeneous \emph{curvature-defined geometries} (i.e., preconditioner coordinate systems), and server-side model averaging updates computed under incompatible metrics, corrupting the global descent direction. To address this geometric mismatch, we propose \texttt{FedPAC}, a \emph{preconditioner alignment and correction} framework for reliable federated second-order optimization. \texttt{FedPAC} explicitly decouples parameter aggregation from geometry synchronization by: (i) \textbf{Alignment} (i.e.,aggregating local preconditioners into a global reference and warm-starting clients via global preconditioner); and (ii) \textbf{Correction} (i.e., steering local preconditioned updates using a global preconditioned direction to suppress long-term drift). We provide drift-coupled non-convex convergence guarantees with linear speedup under partial participation. Empirically, \texttt{FedPAC} consistently improves stability and accuracy across vision and language tasks, achieving up to $5.8\%$ absolute accuracy gain on CIFAR-100 with ViTs. Code is available at https://anonymous.4open.science/r/FedPAC-8B24.

GeoEyes: On-Demand Visual Focusing for Evidence-Grounded Understanding of Ultra-High-Resolution Remote Sensing Imagery

arXiv:2602.14201v1 Announce Type: cross Abstract: The "thinking-with-images" paradigm enables multimodal large language models (MLLMs) to actively explore visual scenes via zoom-in tools. This is essential for ultra-high-resolution (UHR) remote sensing VQA, where task-relevant cues are sparse and tiny. However, we observe a consistent failure mode in existing zoom-enabled MLLMs: Tool Usage Homogenization, where tool calls collapse into task-agnostic patterns, limiting effective evidence acquisition. To address this, we propose GeoEyes, a staged training framework consisting of (1) a cold-start SFT dataset, UHR Chain-of-Zoom (UHR-CoZ), which covers diverse zooming regimes, and (2) an agentic reinforcement learning method, AdaZoom-GRPO, that explicitly rewards evidence gain and answer improvement during zoom interactions. The resulting model learns on-demand zooming with proper stopping behavior and achieves substantial improvements on UHR remote sensing benchmarks, with 54.23% accuracy on XLRS-Bench.
❌