❌

Normal view

Emotional intelligence in large language models is fragmented across perception, cognition, and interaction

arXiv:2605.24686v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly integrated into emotionally sensitive domains, the structural integrity of their emotional intelligence (EI) becomes a critical frontier for safety and alignment. Current benchmarks often conflate superficial politeness with deep affective reasoning, failing to distinguish between perceptual accuracy and interactive efficacy. Here, we introduce FACET (Functional Affective Competence and Empathy Test), a psychometrically grounded framework comprising 480 expert-crafted items. Unlike previous metrics, FACET is theoretically anchored in the Mayer-Salovey-Caruso four-branch ability model, operationalizing EI through perception, facilitation, understanding, and management of emotions. Through an evaluation of nine frontier models (including GPT-5, Claude-Sonnet-4), we demonstrate that emotional intelligence is not a monolithic capability but is fragmented across cognitive and interactive dimensions. While frontier models demonstrate robust proficiency in objective emotion recognition and social reasoning, this does not consistently translate to interactive success. We categorize these discrepancies into three distinct performance profiles: cognitive-dominant, interactive-dominant, and context-dependent. These typologies indicate that emotional skills do not scale uniformly with general intelligence or model size; rather, they are shaped by specific alignment paradigms. Notably, we identify hidden emotion recognition as a universal performance bottleneck across all architectures. Our results suggest that current RLHF processes may optimize for "stochastic empathy", a statistical mimicry of emotional syntax, at the expense of integrated affective reasoning. These findings challenge the assumption of linear emotional scaling and provide a rigorous roadmap for developing socially aware agents capable of genuine clinical resonance.

Multi-omics and experimental validation identify USP54 as a prognostic deubiquitinase promoting pancreatic ductal adenocarcinoma progression within the immune microenvironment

Front Immunol. 2026 Mar 18;17:1791707. doi: 10.3389/fimmu.2026.1791707. eCollection 2026.

ABSTRACT

BACKGROUND: Pancreatic ductal adenocarcinoma (PDAC) is a highly lethal malignancy with a complex tumor ecosystem that contributes to its progression. Deubiquitinases (DUBs) are vital regulators in cancer. However, the overall activity of DUBs and their role in driving PDAC progression within immune microenvironment remain largely unknown.

METHODS: We employed an integrative multi-omics strategy combining machine learning (ML) on bulk transcriptomic data, single-cell RNA sequencing and spatial transcriptomic profiling. We applied Coxnet and Fuzzy SVM for prognostic modeling, inferCNV for malignant cell identification, SCENIC for transcription factor regulon analysis, LIANA+ for inferring inter-cellular communication networks and cell2location for spatial deconvolution. USP54 expression was detected by real-time quantitative PCR, western blotting and immunohistochemistry. USP54 function was validated through in vitro and in vivo assays.

RESULTS: ML-based pathway analysis revealed post-translational modification as a major prognostic category, within which elevated DUBs activity emerged as an independent adverse prognostic factor. At the single-cell level, USP54 was upregulated along the trajectory of malignant ductal cells and correlated with an inflamed tumor microenvironment. Cell-cell communication analysis predicted signaling from monocytes/macrophages to tumor cells via the THBS1-integrin ligand-receptor pair. This immune-derived signaling potentially converged on KLF5-positive tumor cells, with KLF5 identified as a putative transcriptional activator of USP54. Spatial transcriptomics validated the co-localization of USP54 expression, elevated DUB activity, and KRAS signaling within specific tumor niches adjacent to THBS1-enriched immune regions. High USP54 expression was frequently observed in PDAC tissues and associated with poor patient survival. More importantly, in both BxPC-3 and PANC-1 cell lines, USP54 knockdown suppressed cell proliferation and metastasis, whereas its overexpression enhanced these malignant phenotypes. Subcutaneous xenograft growth and tail vein injection experiments validated these findings in vivo.

CONCLUSIONS: Our comprehensive multi-omics analysis and experimental validation identify the deubiquitinase USP54 as a novel promoter of PDAC progression within a spatially organized tumor-immune microenvironment. These findings suggest USP54 as both a candidate prognostic biomarker and a potential therapeutic target for this lethal malignancy.

PMID:41929495 | PMC:PMC13038871 | DOI:10.3389/fimmu.2026.1791707

IGASA: Integrated Geometry-Aware and Skip-Attention Modules for Enhanced Point Cloud Registration

arXiv:2603.12719v1 Announce Type: cross Abstract: Point cloud registration (PCR) is a fundamental task in 3D vision and provides essential support for applications such as autonomous driving, robotics, and environmental modeling. Despite its widespread use, existing methods often fail when facing real-world challenges like heavy noise, significant occlusions, and large-scale transformations. These limitations frequently result in compromised registration accuracy and insufficient robustness in complex environments. In this paper, we propose IGASA as a novel registration framework constructed upon a Hierarchical Pyramid Architecture (HPA) designed for robust multi-scale feature extraction and fusion. The framework integrates two pivotal components consisting of the Hierarchical Cross-Layer Attention (HCLA) module and the Iterative Geometry-Aware Refinement (IGAR) module. The HCLA module utilizes skip attention mechanisms to align multi-resolution features and enhance local geometric consistency. Simultaneously, the IGAR module is designed for the fine matching phase by leveraging reliable correspondences established during coarse matching. This synergistic integration within the architecture allows IGASA to adapt effectively to diverse point cloud structures and intricate transformations. We evaluate the performance of IGASA on four widely recognized benchmark datasets including 3D(Lo)Match, KITTI, and nuScenes. Our extensive experiments consistently demonstrate that IGASA significantly surpasses state-of-the-art methods and achieves notable improvements in registration accuracy. This work provides a robust foundation for advancing point cloud registration techniques while offering valuable insights for practical 3D vision applications. The code for IGASA is available in \href{https://github.com/DongXu-Zhang/IGASA}{https://github.com/DongXu-Zhang/IGASA}.

CMHANet: A Cross-Modal Hybrid Attention Network for Point Cloud Registration

arXiv:2603.12721v1 Announce Type: cross Abstract: Robust point cloud registration is a fundamental task in 3D computer vision and geometric deep learning, essential for applications such as large-scale 3D reconstruction, augmented reality, and scene understanding. However, the performance of established learning-based methods often degrades in complex, real world scenarios characterized by incomplete data, sensor noise, and low overlap regions. To address these limitations, we propose CMHANet, a novel Cross-Modal Hybrid Attention Network. Our method integrates the fusion of rich contextual information from 2D images with the geometric detail of 3D point clouds, yielding a comprehensive and resilient feature representation. Furthermore, we introduce an innovative optimization function based on contrastive learning, which enforces geometric consistency and significantly improves the model's robustness to noise and partial observations. We evaluated CMHANet on the 3DMatch and the challenging 3DLoMatch datasets. \rev{Additionally, zero-shot evaluations on the TUM RGB-D SLAM dataset verify the model's generalization capability to unseen domains.} The experimental results demonstrate that our method achieves substantial improvements in both registration accuracy and overall robustness, outperforming current techniques. We also release our code in \href{https://github.com/DongXu-Zhang/CMHANet}{https://github.com/DongXu-Zhang/CMHANet}.

SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition

arXiv:2511.21471v2 Announce Type: replace Abstract: Spatial cognition is fundamental to real-world multimodal intelligence, allowing models to effectively interact with the physical environment. While multimodal large language models (MLLMs) have made significant strides, existing benchmarks often oversimplify spatial cognition, reducing it to a single-dimensional metric, which fails to capture the hierarchical structure and interdependence of spatial abilities. To address this gap, we propose a hierarchical spatial cognition framework that decomposes spatial intelligence into five progressively complex levels from basic observation to high-level planning. Building upon this taxonomy, we construct SpatialBench, a large-scale, fine-grained benchmark covering 15 tasks aligned with these cognitive levels. To provide a unified evaluation across heterogeneous tasks, we further introduce a high-level capability-oriented metric that reliably assesses a model's overall spatial reasoning ability. Extensive experiments over massive MLLMs reveal distinct performance stratification across cognitive levels: models exhibit strong perceptual grounding yet remain limited in symbolic reasoning, causal inference, and planning. Additional human tests demonstrate that humans perform selective, goal-directed abstraction, while MLLMs tend to over-attend to surface details without coherent spatial intent. Our work establishes the first systematic framework for measuring hierarchical spatial cognition in MLLMs, laying the foundation for future spatially intelligent systems.

Incentivizing Agentic Reasoning in LLM Judges via Tool-Integrated Reinforcement Learning

By: Ran Xu Β· Jingjing Chen Β· Jiayu Ye Β· Yu Wu Β· Jun Yan Β· Carl Yang Β· Hongkun Yu
24 February 2026 at 13:00
arXiv:2510.23038v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are widely used as judges to evaluate response quality, providing a scalable alternative to human evaluation. However, most LLM judges operate solely on intrinsic text-based reasoning, limiting their ability to verify complex constraints or perform accurate computation. Motivated by the success of tool-integrated reasoning (TIR) in numerous tasks, we propose TIR-Judge, an end-to-end RL framework for training LLM judges that integrates a code executor for precise evaluation. TIR-Judge is built on three principles: (i) diverse training across verifiable and non-verifiable domains, (ii) flexible judgment formats (pointwise, pairwise, listwise), and (iii) iterative RL that bootstraps directly from the initial model without distillation. On seven public benchmarks, TIR-Judge surpasses strong reasoning-based judges by up to 6.4% (pointwise) and 7.7% (pairwise), and achieves listwise performance comparable to Claude-Opus-4 despite having only 8B parameters. Remarkably, TIR-Judge-Zero - trained entirely without distilled judge trajectories, matches the performance of distilled variants, demonstrating that tool-augmented judges can self-evolve through iterative reinforcement learning.
❌