❌

Normal view

Pelican-Sim 1.0: A General World Model Simulator for Embodied Intelligence

arXiv:2609.12036v1 Announce Type: cross Abstract: In this technical report, we propose Pelican-Sim 1.0, a general world model simulator for embodied intelligence that predicts future observations from visual context and robot actions to support downstream learning and decision making. The model incorporates four key design features: (1) Unified action representation: a 28-dimensional action value space covering most mainstream embodiments, keeping one model valid across heterogeneous devices. (2) Action-visual injection: URDF- and camera-rendered action videos bridge actions and pixels, giving markedly better controllability across embodiments, scenes, and tasks (PSNR +0.904 over alternative fusion baselines). (3) Sparse mixture-of-experts (MoE): sparse MoE layers add capacity for heterogeneous dynamics and absorb the action modality while reducing inter-modality conflict (FVD -6.530 vs. the dense backbone). (4) Efficient rollout generation: causal adaptation and few-step distillation yield a four-step autoregressive simulator, achieving a 5.67-fold speedup over the 35-step model. Benefiting from these designs, we train on approximately one million real-world and simulated trajectories and obtain large gains in action controllability and video quality: PSNR improves over the strongest evaluated baselines by 4.636 on AgiBotWorld Beta, 2.080 on RoboMIND, and 10.343 on RoboTwin, with the adapted EWMBench DYN score up 0.426 on RoboTwin. Relying on this, four downstream applications on RoboTwin succeed: 500 generated trajectories added to 50 demonstrations per task raise policy success from 70% to 93%; policy evaluation reaches a Pearson correlation of 0.994 across five checkpoints; and relative success gains reach 47.7% for action selection and 20.3% for policy improvement. Qualitative generalization across trajectory, scene, object, embodiment, and viewpoint shifts highlights its potential as a general-purpose world model simulator.

From Black Box to Biological Insight: AttentioFuse Unlocks Multi-Omics Dynamics in Lung Cancer

Cancers (Basel). 2026 Mar 9;18(5):878. doi: 10.3390/cancers18050878.

ABSTRACT

BACKGROUND: Lung adenocarcinoma (LUAD) and squamous cell carcinoma (LUSC), the major subtypes of non-small cell lung cancer (NSCLC), exhibit distinct molecular landscapes that demand precision in prognosis and therapy. While deep learning models can achieve high predictive accuracy, their black-box nature limits clinical translation.

METHODS: We introduce AttentioFuse, an interpretable deep learning framework employing a Reactome-guided mid-fusion strategy for multi-omics integration. AttentioFuse builds on three pillars: (i) dual-phase learning with omics-specific encoders to preserve modality-unique patterns, (ii) hierarchical attention mechanisms (cross-omics, feature-level, and fusion-layer) to quantify layer contributions dynamically, and (iii) integrated explainability combining DeepSHAP and global attention weights for gene-to-pathway interpretation. Two depth variants are instantiated under identical priors: a three-layer configuration (3F) for main discrimination and a five-layer configuration (AttentioFuse-5X) for deeper hierarchical interpretation; the 5X variant is trained end-to-end and yields comparable accuracy while enhancing pathway-level resolution.

RESULTS: Evaluated on The Cancer Genome Atlas (TCGA) LUAD/LUSC cohorts, AttentioFuse matches state-of-the-art performance in TNM staging while uncovering actionable biological insights, including pan-NSCLC AKT/mTOR metabolic control, histology-divergent Notch signaling roles, and additional pathways related to developmental reactivation, microbiota-associated metastasis, and extracellular matrix remodeling.

CONCLUSIONS: By design, AttentioFuse-5X bridges predictive performance with hierarchical, pathway-resolved explanations, advancing oncology by transforming black-box predictions into biologically grounded decision support.

PMID:41827812 | PMC:PMC12985206 | DOI:10.3390/cancers18050878

From Black Box to Biological Insight: AttentioFuse Unlocks Multi-Omics Dynamics in Lung Cancer

14 March 2026 at 18:00

Cancers (Basel). 2026 Mar 9;18(5):878. doi: 10.3390/cancers18050878.

ABSTRACT

BACKGROUND: Lung adenocarcinoma (LUAD) and squamous cell carcinoma (LUSC), the major subtypes of non-small cell lung cancer (NSCLC), exhibit distinct molecular landscapes that demand precision in prognosis and therapy. While deep learning models can achieve high predictive accuracy, their black-box nature limits clinical translation.

METHODS: We introduce AttentioFuse, an interpretable deep learning framework employing a Reactome-guided mid-fusion strategy for multi-omics integration. AttentioFuse builds on three pillars: (i) dual-phase learning with omics-specific encoders to preserve modality-unique patterns, (ii) hierarchical attention mechanisms (cross-omics, feature-level, and fusion-layer) to quantify layer contributions dynamically, and (iii) integrated explainability combining DeepSHAP and global attention weights for gene-to-pathway interpretation. Two depth variants are instantiated under identical priors: a three-layer configuration (3F) for main discrimination and a five-layer configuration (AttentioFuse-5X) for deeper hierarchical interpretation; the 5X variant is trained end-to-end and yields comparable accuracy while enhancing pathway-level resolution.

RESULTS: Evaluated on The Cancer Genome Atlas (TCGA) LUAD/LUSC cohorts, AttentioFuse matches state-of-the-art performance in TNM staging while uncovering actionable biological insights, including pan-NSCLC AKT/mTOR metabolic control, histology-divergent Notch signaling roles, and additional pathways related to developmental reactivation, microbiota-associated metastasis, and extracellular matrix remodeling.

CONCLUSIONS: By design, AttentioFuse-5X bridges predictive performance with hierarchical, pathway-resolved explanations, advancing oncology by transforming black-box predictions into biologically grounded decision support.

PMID:41827812 | PMC:PMC12985206 | DOI:10.3390/cancers18050878

Give Them an Inch and They Will Take a Mile:Understanding and Measuring Caller Identity Confusion in MCP-Based AI Systems

arXiv:2603.07473v1 Announce Type: cross Abstract: The Model Context Protocol (MCP) is an open and standardized interface that enables large language models (LLMs) to interact with external tools and services, and is increasingly adopted by AI agents. However, the security of MCP-based systems remains largely unexplored.In this work, we conduct a large-scale security analysis of MCP servers integrated within MCP clients. We show that treating MCP servers as trusted entities without authenticating the caller identity is fundamentally insecure. Since MCP servers often cannot distinguish who is invoking a request, a single authorization decision may implicitly grant access to multiple, potentially untrusted callers.Our empirical study reveals that most MCP servers rely on persistent authorization states, allowing tool invocations after an initial authorization without re-authentication, regardless of the caller. In addition, many MCP servers fail to enforce authentication at the per-tool level, enabling unauthorized access to sensitive operations.These findings demonstrate that one-time authorization and server-level trust significantly expand the attack surface of MCP-based systems, highlighting the need for explicit caller authentication and fine-grained authorization mechanisms.
❌