❌

Normal view

Med-CMR: A Fine-Grained Benchmark Integrating Visual Evidence and Clinical Logic for Medical Complex Multimodal Reasoning

arXiv:2512.00818v2 Announce Type: replace Abstract: MLLMs MLLMs are beginning to appear in clinical workflows, but their ability to perform complex medical reasoning remains unclear. We present Med-CMR, a fine-grained Medical Complex Multimodal Reasoning benchmark. Med-CMR distinguishes from existing counterparts by three core features: 1) Systematic capability decomposition, splitting medical multimodal reasoning into fine-grained visual understanding and multi-step reasoning to enable targeted evaluation; 2) Challenging task design, with visual understanding across three key dimensions (small-object detection, fine-detail discrimination, spatial understanding) and reasoning covering four clinically relevant scenarios (temporal prediction, causal reasoning, long-tail generalization, multi-source integration); 3) Broad, high-quality data coverage, comprising 20,653 Visual Question Answering (VQA) pairs spanning 11 organ systems and 12 imaging modalities, validated via a rigorous two-stage (human expert + model-assisted) review to ensure clinical authenticity. We evaluate 18 state-of-the-art MLLMs with Med-CMR, revealing GPT-5 as the top-performing commercial model: 57.81 accuracy on multiple-choice questions (MCQs) and a 48.70 open-ended score, outperforming Gemini 2.5 Pro (49.87 MCQ accuracy, 45.98 open-ended score) and leading open-source model Qwen3-VL-235B-A22B (49.34 MCQ accuracy, 42.62 open-ended score). However, specialized medical MLLMs do not reliably outperform strong general models, and long-tail generalization emerges as the dominant failure mode. Med-CMR thus provides a stress test for visual-reasoning integration and rare-case robustness in medical MLLMs, and a rigorous yardstick for future clinical systems.

WiseMind: a knowledge-guided multi-agent framework for accurate and empathetic psychiatric diagnosis

npj Digital Medicine, Published online: 25 March 2026; doi:10.1038/s41746-026-02559-9

WiseMind: a knowledge-guided multi-agent framework for accurate and empathetic psychiatric diagnosis

Comprehensive multi omics profiling and Mendelian randomization assessment of lipid metabolites in lung cancer prognosis

Discov Oncol. 2026 Mar 23. doi: 10.1007/s12672-026-04893-6. Online ahead of print.

ABSTRACT

BACKGROUND: Lung cancer remains the leading cause of cancer-related mortality worldwide. This study aimed to develop prognostic prediction models for lung squamous cell carcinoma (LUSC) through multi-omics integration using Mendelian randomization analysis.This study addresses a critical gap in lung cancer research through two complementary approaches in major lung cancer subtypes: (1) hypothesis-generating multi-omics analysis in LUSC to identify prognostic biomarkers and characterize the metabolic-immune landscape. This integrated framework provides both predictive tools for personalized medicine and mechanistic insights into metabolic causality.

METHODS: Multi-omics analysis was performed using TCGA data, including RNA-seq, DNA methylation, and whole-exome sequencing. Machine learning models incorporating 15 algorithms were developed and externally validated in two independent GEO cohorts. Mendelian randomization analysis assessed causal relationships between 32 lipid metabolites and SCLC risk. RT-qPCR experiments validated key prognostic genes in lung squamous cell carcinoma (LUSC) cell lines.

RESULTS: The optimal machine learning model (StepCox [forward] + Random Survival Forest) demonstrated superior performance with C-index of 0.73 in internal testing and 0.71 and 0.68 in external validation cohorts. High CD8 + T cell and M1 macrophage infiltration was associated with favorable prognosis. Most lipid metabolites showed no significant causal associations with SCLC risk after multiple testing correction, though two phosphatidylcholine metabolites demonstrated potential protective effects. RT-qPCR validation confirmed significant upregulation of all four key genes in LUSC cell lines.

CONCLUSIONS: This study successfully developed robust machine learning-based prognostic models for LUSC with clinical utility for risk stratification and provided evidence that lipid alterations in lung cancer are likely downstream consequences rather than causal drivers of tumorigenesis.

PMID:41870745 | DOI:10.1007/s12672-026-04893-6

Comprehensive multi omics profiling and Mendelian randomization assessment of lipid metabolites in lung cancer prognosis

23 March 2026 at 18:00

Discov Oncol. 2026 Mar 23. doi: 10.1007/s12672-026-04893-6. Online ahead of print.

ABSTRACT

BACKGROUND: Lung cancer remains the leading cause of cancer-related mortality worldwide. This study aimed to develop prognostic prediction models for lung squamous cell carcinoma (LUSC) through multi-omics integration using Mendelian randomization analysis.This study addresses a critical gap in lung cancer research through two complementary approaches in major lung cancer subtypes: (1) hypothesis-generating multi-omics analysis in LUSC to identify prognostic biomarkers and characterize the metabolic-immune landscape. This integrated framework provides both predictive tools for personalized medicine and mechanistic insights into metabolic causality.

METHODS: Multi-omics analysis was performed using TCGA data, including RNA-seq, DNA methylation, and whole-exome sequencing. Machine learning models incorporating 15 algorithms were developed and externally validated in two independent GEO cohorts. Mendelian randomization analysis assessed causal relationships between 32 lipid metabolites and SCLC risk. RT-qPCR experiments validated key prognostic genes in lung squamous cell carcinoma (LUSC) cell lines.

RESULTS: The optimal machine learning model (StepCox [forward] + Random Survival Forest) demonstrated superior performance with C-index of 0.73 in internal testing and 0.71 and 0.68 in external validation cohorts. High CD8 + T cell and M1 macrophage infiltration was associated with favorable prognosis. Most lipid metabolites showed no significant causal associations with SCLC risk after multiple testing correction, though two phosphatidylcholine metabolites demonstrated potential protective effects. RT-qPCR validation confirmed significant upregulation of all four key genes in LUSC cell lines.

CONCLUSIONS: This study successfully developed robust machine learning-based prognostic models for LUSC with clinical utility for risk stratification and provided evidence that lipid alterations in lung cancer are likely downstream consequences rather than causal drivers of tumorigenesis.

PMID:41870745 | DOI:10.1007/s12672-026-04893-6

Thinking in Streaming Video

arXiv:2603.12938v1 Announce Type: cross Abstract: Real-time understanding of continuous video streams is essential for interactive assistants and multimodal agents operating in dynamic environments. However, most existing video reasoning approaches follow a batch paradigm that defers reasoning until the full video context is observed, resulting in high latency and growing computational cost that are incompatible with streaming scenarios. In this paper, we introduce ThinkStream, a framework for streaming video reasoning based on a Watch--Think--Speak paradigm that enables models to incrementally update their understanding as new video observations arrive. At each step, the model performs a short reasoning update and decides whether sufficient evidence has accumulated to produce a response. To support long-horizon streaming, we propose Reasoning-Compressed Streaming Memory (RCSM), which treats intermediate reasoning traces as compact semantic memory that replaces outdated visual tokens while preserving essential context. We further train the model using a Streaming Reinforcement Learning with Verifiable Rewards scheme that aligns incremental reasoning and response timing with the requirements of streaming interaction. Experiments on multiple streaming video benchmarks show that ThinkStream significantly outperforms existing online video models while maintaining low latency and memory usage. Code, models and data will be released at https://github.com/johncaged/ThinkStream

Latilactobacillus curvatus IM01 Alleviates Allergic Airway Inflammation Through Microbial and Metabolic Crosstalk Along the Gut-Lung Axis

Nutrients. 2026 Mar 4;18(5):834. doi: 10.3390/nu18050834.

ABSTRACT

Background: Gut microbiota dysbiosis is critically implicated in the pathogenesis of allergic airway inflammation (AAI) via the gut-lung axis. While Latilactobacillus curvatus is a promising probiotic candidate, its specific immunomodulatory mechanisms in respiratory diseases remain poorly understood. Objective: In this study, we investigated the protective effects and underlying mechanisms of L. curvatus IM01 in an ovalbumin (OVA)-induced murine AAI model using an integrated multi-omics approach. Results: Our results demonstrated that oral administration of L. curvatus IM01 significantly attenuated airway inflammation, suppressed Th2-type immune responses, and reduced serum IgE levels. Crucially, our multi-omics integration revealed a coherent gut-lung axis narrative driven by microbial and metabolic crosstalk. Specifically, 16S rRNA sequencing indicated that L. curvatus IM01 was closely linked to structural shifts in the gut microbial community, notably characterized by an enrichment trend for beneficial genera such as Odoribacter and Lactobacillus. This microbial restructuring was closely associated with a modulated cecal metabolic profile, as untargeted metabolomics exhibited a clear trend toward the restoration of key systemically active immunoregulatory metabolites, including indolelactic acid (ILA) and choline, which have been previously linked to the alleviation of AAI symptoms. Further linking this metabolic shift to respiratory immune tolerance, lung transcriptomic analysis showed that the treatment is strongly associated with the promotion of the differentiation of CD4+ T cells into Foxp3+ regulatory T cells (Tregs). Conclusions: Collectively, these findings suggest a novel potential pathway by which L. curvatus IM01 modulates the gut-lung axis through coordinated microbial and metabolic interventions, highlighting its potential as a therapeutic functional food ingredient for AAI.

PMID:41830004 | PMC:PMC12987261 | DOI:10.3390/nu18050834

Latilactobacillus curvatus IM01 Alleviates Allergic Airway Inflammation Through Microbial and Metabolic Crosstalk Along the Gut-Lung Axis

Nutrients. 2026 Mar 4;18(5):834. doi: 10.3390/nu18050834.

ABSTRACT

Background: Gut microbiota dysbiosis is critically implicated in the pathogenesis of allergic airway inflammation (AAI) via the gut-lung axis. While Latilactobacillus curvatus is a promising probiotic candidate, its specific immunomodulatory mechanisms in respiratory diseases remain poorly understood. Objective: In this study, we investigated the protective effects and underlying mechanisms of L. curvatus IM01 in an ovalbumin (OVA)-induced murine AAI model using an integrated multi-omics approach. Results: Our results demonstrated that oral administration of L. curvatus IM01 significantly attenuated airway inflammation, suppressed Th2-type immune responses, and reduced serum IgE levels. Crucially, our multi-omics integration revealed a coherent gut-lung axis narrative driven by microbial and metabolic crosstalk. Specifically, 16S rRNA sequencing indicated that L. curvatus IM01 was closely linked to structural shifts in the gut microbial community, notably characterized by an enrichment trend for beneficial genera such as Odoribacter and Lactobacillus. This microbial restructuring was closely associated with a modulated cecal metabolic profile, as untargeted metabolomics exhibited a clear trend toward the restoration of key systemically active immunoregulatory metabolites, including indolelactic acid (ILA) and choline, which have been previously linked to the alleviation of AAI symptoms. Further linking this metabolic shift to respiratory immune tolerance, lung transcriptomic analysis showed that the treatment is strongly associated with the promotion of the differentiation of CD4+ T cells into Foxp3+ regulatory T cells (Tregs). Conclusions: Collectively, these findings suggest a novel potential pathway by which L. curvatus IM01 modulates the gut-lung axis through coordinated microbial and metabolic interventions, highlighting its potential as a therapeutic functional food ingredient for AAI.

PMID:41830004 | PMC:PMC12987261 | DOI:10.3390/nu18050834

Multi-Omics and Single-Cell Mendelian Randomization Reveal a Potential Role of VNN2 in Lung Adenocarcinoma in Resting Natural Killer Cells

World J Oncol. 2026 Mar 5;17(2):247-255. doi: 10.14740/wjon2689. eCollection 2026 Apr.

ABSTRACT

BACKGROUND: We aimed to evaluate the potential association between genetically predicted vanin-2 (VNN2) expression and lung adenocarcinoma (LUAD) risk, and to explore the immune cell subtype that may underlie this relationship.

METHODS: We integrated whole-blood expression quantitative trait loci (eQTL) data from eQTLGen, plasma protein quantitative trait loci (pQTL) data from deCODE, and LUAD genome-wide association study (GWAS) data from European-ancestry cohorts, together with differential expression analysis using GEPIA2, to identify candidate genes for subsequent single-cell eQTL (sc-eQTL) Mendelian randomization (MR) analysis. For the sc-eQTL analysis, VNN2-associated eQTLs from 14 immune cell types profiled in the OneK1K single-cell eQTL resource were tested for associations with LUAD risk.

RESULTS: Bulk-level MR analysis showed that genetically predicted increases in VNN2 expression and protein levels were significantly associated with a reduced risk of LUAD (eQTL-MR: odds ratio (OR) = 0.964, 95% confidence interval (95% CI), 0.934-0.995; P = 0.024; pQTL-MR: OR = 0.946, 95% CI, 0.921-0.970; P = 2.87 × 10-5). Transcriptomic analyses confirmed significant downregulation of VNN2 in LUAD tumors compared with normal lung tissues. sc-eQTL MR identified the strongest association in resting natural killer (rNK) cells (OR = 0.896, 95% CI, 0.829-0.967; P = 0.005).

CONCLUSIONS: Multi-omics and sc-eQTL MR analyses indicated that genetically predicted increases in VNN2 expression were associated with a reduced risk of LUAD, with the most pronounced effect observed in rNK cells. These findings suggest a potential cell type-specific role of VNN2 in LUAD susceptibility and warrant further studies to validate its biological relevance and clinical implications.

PMID:41822323 | PMC:PMC12978397 | DOI:10.14740/wjon2689

ViLAM: Distilling Vision-Language Reasoning into Attention Maps for Social Robot Navigation

arXiv:2503.09820v2 Announce Type: replace-cross Abstract: We introduce ViLAM, a novel method for distilling vision-language reasoning from large Vision-Language Models (VLMs) into spatial attention maps for socially compliant robot navigation. Unlike traditional methods that rely on expert demonstrations or human-annotated datasets, ViLAM performs knowledge distillation and fine-tuning at the intermediate layer representation (attention) level by aligning attention maps from a pretrained vision-action model with socially guided attention maps derived from a large VLM. These distilled attention maps highlight key navigational regions in a scene and serve as socially informed spatial cost maps for motion planning. To achieve this, we introduce a novel attention-level distillation loss that fuses knowledge from both sources, generating augmented attention maps with enhanced social awareness. These refined attention maps are then used as a traversability costmap within a socially aware local planner for navigation. We validate our approach through real-world experiments on a Husky wheeled robot, and demonstrate 14.2% - 50% improvements in success rate over existing methods.

FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference

arXiv:2505.13109v5 Announce Type: replace-cross Abstract: Large language models (LLMs) are widely deployed with rapidly expanding context windows to support increasingly demanding applications. However, long contexts pose significant deployment challenges, primarily due to the KV cache whose size grows proportionally with context length. While KV cache compression methods have been proposed to address this issue, KV dropping methods incur considerable accuracy loss, and KV retrieval methods suffer from significant efficiency bottlenecks. We propose FreeKV, a training-free algorithm-system co-optimization framework to enhance KV retrieval efficiency while preserving accuracy. On the algorithm side, FreeKV introduces speculative retrieval to shift the KV selection and recall processes out of the critical path, combined with fine-grained correction to ensure accuracy. On the system side, FreeKV employs hybrid KV layouts across CPU and GPU memory to eliminate fragmented data transfers, and leverages double-buffered streamed recall to further improve efficiency, enabling effective overlap with computation, full latency hiding, and practical speedups from speculative recall. Experiments demonstrate that FreeKV achieves near-lossless accuracy across various scenarios and models, delivering up to a 13$\times$ speedup compared to SOTA KV retrieval methods. Code is available at https://github.com/sjtu-zhao-lab/FreeKV.

Mozi: Governed Autonomy for Drug Discovery LLM Agents

arXiv:2603.03655v1 Announce Type: new Abstract: Tool-augmented large language model (LLM) agents promise to unify scientific reasoning with computation, yet their deployment in high-stakes domains like drug discovery is bottlenecked by two critical barriers: unconstrained tool-use governance and poor long-horizon reliability. In dependency-heavy pharmaceutical pipelines, autonomous agents often drift into irreproducible trajectories, where early-stage hallucinations multiplicatively compound into downstream failures. To overcome this, we present Mozi, a dual-layer architecture that bridges the flexibility of generative AI with the deterministic rigor of computational biology. Layer A (Control Plane) establishes a governed supervisor--worker hierarchy that enforces role-based tool isolation, limits execution to constrained action spaces, and drives reflection-based replanning. Layer B (Workflow Plane) operationalizes canonical drug discovery stages -- from Target Identification to Lead Optimization -- as stateful, composable skill graphs. This layer integrates strict data contracts and strategic human-in-the-loop (HITL) checkpoints to safeguard scientific validity at high-uncertainty decision boundaries. Operating on the design principle of ``free-form reasoning for safe tasks, structured execution for long-horizon pipelines,'' Mozi provides built-in robustness mechanisms and trace-level audibility to completely mitigate error accumulation. We evaluate Mozi on PharmaBench, a curated benchmark for biomedical agents, demonstrating superior orchestration accuracy over existing baselines. Furthermore, through end-to-end therapeutic case studies, we demonstrate Mozi's ability to navigate massive chemical spaces, enforce stringent toxicity filters, and generate highly competitive in silico candidates, effectively transforming the LLM from a fragile conversationalist into a reliable, governed co-scientist.

Unbiased Dynamic Pruning for Efficient Group-Based Policy Optimization

arXiv:2603.04135v1 Announce Type: cross Abstract: Group Relative Policy Optimization (GRPO) effectively scales LLM reasoning but incurs prohibitive computational costs due to its extensive group-based sampling requirement. While recent selective data utilization methods can mitigate this overhead, they could induce estimation bias by altering the underlying sampling distribution, compromising theoretical rigor and convergence behavior. To address this limitation, we propose Dynamic Pruning Policy Optimization (DPPO), a framework that enables dynamic pruning while preserving unbiased gradient estimation through importance sampling-based correction. By incorporating mathematically derived rescaling factors, DPPO significantly accelerates GRPO training without altering the optimization objective of the full-batch baseline. Furthermore, to mitigate the data sparsity induced by pruning, we introduce Dense Prompt Packing, a window-based greedy strategy that maximizes valid token density and hardware utilization. Extensive experiments demonstrate that DPPO consistently accelerates training across diverse models and benchmarks. For instance, on Qwen3-4B trained on MATH, DPPO achieves 2.37$\times$ training speedup and outperforms GRPO by 3.36% in average accuracy across six mathematical reasoning benchmarks.

High-order Knowledge Based Network Controllability Robustness Prediction: A Hypergraph Neural Network Approach

arXiv:2603.02265v1 Announce Type: cross Abstract: In order to evaluate the invulnerability of networks against various types of attacks and provide guidance for potential performance enhancement as well as controllability maintenance, network controllability robustness (NCR) has attracted increasing attention in recent years. Traditionally, controllability robustness is determined by attack simulations, which are computationally time-consuming and only applicable to small-scale networks. Although some machine learning-based methods for predicting network controllability robustness have been proposed, they mainly focus on pairwise interactions in complex networks, and the underlying relationships between high-order structural information and controllability robustness have not been explored. In this paper, a dual hypergraph attention neural network model based on high-order knowledge (NCR-HoK) is proposed to accomplish robustness learning and controllability robustness curve prediction. Through a node feature encoder, hypergraph construction with high-order relations, and a dedicated dual hypergraph attention module, the proposed method can effectively learn three types of network information simultaneously: explicit structural information in the original graph, high-order connection information in local neighborhoods, and hidden features in the embedding space. Notably, we explore for the first time the impact of high-order knowledge on network controllability robustness. Compared with state-of-the-art methods for network robustness learning, the proposed method achieves superior performance on both synthetic and real-world networks with low computational overhead.

Rigidity-Aware Geometric Pretraining for Protein Design and Conformational Ensembles

arXiv:2603.02406v1 Announce Type: cross Abstract: Generative models have recently advanced $\textit{de novo}$ protein design by learning the statistical regularities of natural structures. However, current approaches face three key limitations: (1) Existing methods cannot jointly learn protein geometry and design tasks, where pretraining can be a solution; (2) Current pretraining methods mostly rely on local, non-rigid atomic representations for property prediction downstream tasks, limiting global geometric understanding for protein generation tasks; and (3) Existing approaches have yet to effectively model the rich dynamic and conformational information of protein structures. To overcome these issues, we introduce $\textbf{RigidSSL}$ ($\textit{Rigidity-Aware Self-Supervised Learning}$), a geometric pretraining framework that front-loads geometry learning prior to generative finetuning. Phase I (RigidSSL-Perturb) learns geometric priors from 432K structures from the AlphaFold Protein Structure Database with simulated perturbations. Phase II (RigidSSL-MD) refines these representations on 1.3K molecular dynamics trajectories to capture physically realistic transitions. Underpinning both phases is a bi-directional, rigidity-aware flow matching objective that jointly optimizes translational and rotational dynamics to maximize mutual information between conformations. Empirically, RigidSSL variants improve designability by up to 43\% while enhancing novelty and diversity in unconditional generation. Furthermore, RigidSSL-Perturb improves the success rate by 5.8\% in zero-shot motif scaffolding and RigidSSL-MD captures more biophysically realistic conformational ensembles in G protein-coupled receptor modeling. The code is available at: https://github.com/ZhanghanNi/RigidSSL.git.

xLLM Technical Report

arXiv:2510.14686v2 Announce Type: replace-cross Abstract: We introduce xLLM, an intelligent and efficient Large Language Model (LLM) inference framework designed for high-performance, large-scale enterprise-grade serving, with deep optimizations for diverse AI accelerators. To address these challenges, xLLM builds a novel decoupled service-engine architecture. At the service layer, xLLM-Service features an intelligent scheduling module that efficiently processes multimodal requests and co-locates online and offline tasks through unified elastic scheduling to maximize cluster utilization. This module also relies on a workload-adaptive dynamic Prefill-Decode (PD) disaggregation policy and a novel Encode-Prefill-Decode (EPD) disaggregation policy designed for multimodal inputs. Furthermore, it incorporates a distributed architecture to provide global KV Cache management and robust fault-tolerant capabilities for high availability. At the engine layer, xLLM-Engine co-optimizes system and algorithm designs to fully saturate computing resources. This is achieved through comprehensive multi-layer execution pipeline optimizations, an adaptive graph mode and an xTensor memory management. xLLM-Engine also further integrates algorithmic enhancements such as optimized speculative decoding and dynamic EPLB, collectively serving to substantially boost throughput and inference efficiency. Extensive evaluations demonstrate that xLLM delivers significantly superior performance and resource efficiency. Under identical TPOT constraints, xLLM achieves throughput up to 1.7x that of MindIE and 2.2x that of vLLM-Ascend with Qwen-series models, while maintaining an average throughput of 1.7x that of MindIE with Deepseek-series models. xLLM framework is publicly available at https://github.com/jd-opensource/xllm and https://github.com/jd-opensource/xllm-service.

Generative Reasoning Re-ranker

arXiv:2602.07774v4 Announce Type: replace-cross Abstract: Recent studies increasingly explore Large Language Models (LLMs) as a new paradigm for recommendation systems due to their scalability and world knowledge. However, existing work has three key limitations: (1) most efforts focus on retrieval and ranking, while the reranking phase, critical for refining final recommendations, is largely overlooked; (2) LLMs are typically used in zero-shot or supervised fine-tuning settings, leaving their reasoning abilities, especially those enhanced through reinforcement learning (RL) and high-quality reasoning data, underexploited; (3) items are commonly represented by non-semantic IDs, creating major scalability challenges in industrial systems with billions of identifiers. To address these gaps, we propose the Generative Reasoning Reranker (GR2), an end-to-end framework with a three-stage training pipeline tailored for reranking. First, a pretrained LLM is mid-trained on semantic IDs encoded from non-semantic IDs via a tokenizer achieving $\ge$99% uniqueness. Next, a stronger larger-scale LLM generates high-quality reasoning traces through carefully designed prompting and rejection sampling, which are used for supervised fine-tuning to impart foundational reasoning skills. Finally, we apply Decoupled Clip and Dynamic sAmpling Policy Optimization (DAPO), enabling scalable RL supervision with verifiable rewards designed specifically for reranking. Experiments on two real-world datasets demonstrate GR2's effectiveness: it surpasses the state-of-the-art OneRec-Think by 2.4% in Recall@5 and 1.3% in NDCG@5. Ablations confirm that advanced reasoning traces yield substantial gains across metrics. We further find that RL reward design is crucial in reranking: LLMs tend to exploit reward hacking by preserving item order, motivating conditional verifiable rewards to mitigate this behavior and optimize reranking performance.

MOTIF: Learning Action Motifs for Few-shot Cross-Embodiment Transfer

arXiv:2602.13764v1 Announce Type: cross Abstract: While vision-language-action (VLA) models have advanced generalist robotic learning, cross-embodiment transfer remains challenging due to kinematic heterogeneity and the high cost of collecting sufficient real-world demonstrations to support fine-tuning. Existing cross-embodiment policies typically rely on shared-private architectures, which suffer from limited capacity of private parameters and lack explicit adaptation mechanisms. To address these limitations, we introduce MOTIF for efficient few-shot cross-embodiment transfer that decouples embodiment-agnostic spatiotemporal patterns, termed action motifs, from heterogeneous action data. Specifically, MOTIF first learns unified motifs via vector quantization with progress-aware alignment and embodiment adversarial constraints to ensure temporal and cross-embodiment consistency. We then design a lightweight predictor that predicts these motifs from real-time inputs to guide a flow-matching policy, fusing them with robot-specific states to enable action generation on new embodiments. Evaluations across both simulation and real-world environments validate the superiority of MOTIF, which significantly outperforms strong baselines in few-shot transfer scenarios by 6.5% in simulation and 43.7% in real-world settings. Code is available at https://github.com/buduz/MOTIF.
❌