❌

Normal view

JFTA-Bench: Evaluate LLM's Ability of Tracking and Analyzing Malfunctions Using Fault Trees

arXiv:2603.22978v1 Announce Type: new Abstract: In the maintenance of complex systems, fault trees are used to locate problems and provide targeted solutions. To enable fault trees stored as images to be directly processed by large language models, which can assist in tracking and analyzing malfunctions, we propose a novel textual representation of fault trees. Building on it, we construct a benchmark for multi-turn dialogue systems that emphasizes robust interaction in complex environments, evaluating a model's ability to assist in malfunction localization, which contains $3130$ entries and $40.75$ turns per entry on average. We train an end-to-end model to generate vague information to reflect user behavior and introduce long-range rollback and recovery procedures to simulate user error scenarios, enabling assessment of a model's integrated capabilities in task tracking and error recovery, and Gemini 2.5 pro archives the best performance.

SortedRL: Accelerating RL Training for LLMs through Online Length-Aware Scheduling

arXiv:2603.23414v1 Announce Type: cross Abstract: Scaling reinforcement learning (RL) has shown strong promise for enhancing the reasoning abilities of large language models (LLMs), particularly in tasks requiring long chain-of-thought generation. However, RL training efficiency is often bottlenecked by the rollout phase, which can account for up to 70% of total training time when generating long trajectories (e.g., 16k tokens), due to slow autoregressive generation and synchronization overhead between rollout and policy updates. We propose SortedRL, an online length-aware scheduling strategy designed to address this bottleneck by improving rollout efficiency and maintaining training stability. SortedRL reorders rollout samples based on output lengths, prioritizing short samples forming groups for early updates. This enables large rollout batches, flexible update batches, and near on-policy micro-curriculum construction simultaneously. To further accelerate the pipeline, SortedRL incorporates a mechanism to control the degree of off-policy training through a cache-based mechanism, and is supported by a dedicated RL infrastructure that manages rollout and update via a stateful controller and rollout buffer. Experiments using LLaMA-3.1-8B and Qwen-2.5-32B on diverse tasks, including logical puzzles, and math challenges like AIME 24, Math 500, and Minerval, show that SortedRL reduces RL training bubble ratios by over 50%, while attaining 3.9% to 18.4% superior performance over baseline given same amount of data.

SARE: Sample-wise Adaptive Reasoning for Training-free Fine-grained Visual Recognition

arXiv:2603.17729v2 Announce Type: replace-cross Abstract: Recent advances in Large Vision-Language Models (LVLMs) have enabled training-free Fine-Grained Visual Recognition (FGVR). However, effectively exploiting LVLMs for FGVR remains challenging due to the inherent visual ambiguity of subordinate-level categories. Existing methods predominantly adopt either retrieval-oriented or reasoning-oriented paradigms to tackle this challenge, but both are constrained by two fundamental limitations:(1) They apply the same inference pipeline to all samples without accounting for uneven recognition difficulty, thereby leading to suboptimal accuracy and efficiency; (2) The lack of mechanisms to consolidate and reuse error-specific experience causes repeated failures on similar challenging cases. To address these limitations, we propose SARE, a Sample-wise Adaptive textbfREasoning framework for training-free FGVR. Specifically, SARE adopts a cascaded design that combines fast candidate retrieval with fine-grained reasoning, invoking the latter only when necessary. In the reasoning process, SARE incorporates a self-reflective experience mechanism that leverages past failures to provide transferable discriminative guidance during inference, without any parameter updates. Extensive experiments across 14 datasets substantiate that SARE achieves state-of-the-art performance while substantially reducing computational overhead.

Human-specific features of the cerebellum and ZP2-regulated synapse development

Human-specific transcriptomic and regulatory features are present in the cerebellum, with ZP2 playing a key role in synapse regulation. ZP2 expression is induced by pontine mossy fibers, leading to decreased synaptic proteins and neuronal activity, which provides insights into the evolutionary development of the human cerebellum.

Evaluation Faking: Unveiling Observer Effects in Safety Evaluation of Frontier AI Systems

arXiv:2505.17815v2 Announce Type: replace Abstract: As foundation models grow increasingly more intelligent, reliable and trustworthy safety evaluation becomes more indispensable than ever. However, an important question arises: Whether and how an advanced AI system would perceive the situation of being evaluated, and lead to the broken integrity of the evaluation process? During standard safety tests on a mainstream large reasoning model, we unexpectedly observe that the model without any contextual cues would occasionally recognize it is being evaluated and hence behave more safety-aligned. This motivates us to conduct a systematic study on the phenomenon of evaluation faking, i.e., an AI system autonomously alters its behavior upon recognizing the presence of an evaluation context and thereby influencing the evaluation results. Through extensive experiments on a diverse set of foundation models with mainstream safety benchmarks, we reach the main finding termed the observer effects for AI: When the AI system under evaluation is more advanced in reasoning and situational awareness, the evaluation faking behavior becomes more ubiquitous, which reflects in the following aspects: 1) Reasoning models recognize evaluation 16% more often than non-reasoning models. 2) Scaling foundation models (32B to 671B) increases faking by over 30% in some cases, while smaller models show negligible faking. 3) AI with basic memory is 2.3x more likely to recognize evaluation and scores 19% higher on safety tests (vs. no memory). To measure this, we devised a chain-of-thought monitoring technique to detect faking intent and uncover internal signals correlated with such behavior, offering insights for future mitigation studies.

Effects of Digital Health Interventions on Functional and Psychological Outcomes in Older Patients With Hip Fractures: Systematic Review and Meta-Analysis of Randomized Controlled Trials

Background: Hip fractures in older adults increasingly challenge public health, making traditional rehabilitation very challenging. Digital health interventions (DHIs) have emerged as a promising solution for postoperative rehabilitation. However, evidence on DHIs’ effects on functional and psychological outcomes remains insufficient. Objective: This systematic review aimed to comprehensively examine the effects of DHIs on functional and psychological outcomes in older adults with hip fractures. Methods: Following PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) guidelines, we searched 9 databases (PubMed, Embase, CENTRAL, APA PsycINFO, Web of Science, PEDro, CNKI, WANFANG, and SinoMed) from inception to November 13, 2025. Included studies enrolled adults aged 60 years and older with hip fractures, delivered DHIs, assessed functional and psychological outcomes, set usual care or no intervention as the control, and had a randomized controlled trial design. Studies were excluded if they enrolled nonhospitalized patients in the emergency department, patients discharged to nonhome settings, or had inaccessible full text or insufficient data. Study quality was evaluated using the Cochrane Risk of Bias tool 2.0 (Cochrane Collaboration), and evidence certainty was assessed using GRADE (Grading of Recommendations, Assessment, Development and Evaluation). The literature screening, data extraction, and quality assessment were independently conducted by 2 researchers, and any disputes were resolved by the third researcher. We performed analysis using R version 4.0.3 (R Foundation for Statistical Computing) with a random-effects model. Results: Of 17,723 studies screened, 13 met the inclusion criteria. DHIs, compared to the control, significantly improved hip function (standardized mean difference [SMD] 0.80, 95% CI 0.33-1.26; 95% prediction interval [PI] –0.24 to 1.83; P=.007) and functional independence (SMD 1.23, 95% CI 0.34-2.11; 95% PI –0.98 to 3.34; P=.02). Despite favorable pooled effects, a wide 95% PI spanning positive or negative values signals substantial heterogeneity. No significant difference was observed in balance function, risk of falling, and quality of life. Only a single available study reported a 70% adherence rate in the DHIs group. Subgroup analyses stratified by intervention duration revealed no significant intersubgroup differences for hip function (χ12=0.1; P=.75) or functional independence (χ12=2.93; P=.09). For hip function, the point estimate favored the 3 months subgroup (SMD 0.89, 95% CI 0.36-1.41; I2=7%; P=.41) over the <3 months subgroup. Conversely, for functional independence, the point estimate favored shorter intervention duration (SMD 0.67, 95% CI 0.12-1.23; I²=0%; P=.72). Conclusions: This review incorporates the latest randomized controlled trials and comprehensively assesses functional and psychological outcomes of DHIs in older patients with hip fractures, distinct from prior studies focusing solely on functional outcomes. While the 95% CI supports the potential of DHIs to improve hip function and functional independence, the wide 95% PI indicating substantial real-world response variability, which calls for cautious interpretation, informs the design of targeted DHI-based rehabilitation regimens, warranting further research into optimal techniques and dosages in clinical practice. Trial Registration: PROSPERO CRD42024626186; https://www.crd.york.ac.uk/PROSPERO/view/CRD42024626186

Rel-MOSS: Towards Imbalanced Relational Deep Learning on Relational Databases

arXiv:2603.07916v1 Announce Type: new Abstract: In recent advances, to enable a fully data-driven learning paradigm on relational databases (RDB), relational deep learning (RDL) is proposed to structure the RDB as a heterogeneous entity graph and adopt the graph neural network (GNN) as the predictive model. However, existing RDL methods neglect the imbalance problem of relational data in RDBs and risk under-representing the minority entities, leading to an unusable model in practice. In this work, we investigate, for the first time, class imbalance problem in RDB entity classification and design the relation-centric minority synthetic over-sampling GNN (Rel-MOSS), in order to fill a critical void in the current literature. Specifically, to mitigate the issue of minority-related information being submerged by majority counterparts, we design the relation-wise gating controller to modulate neighborhood messages from each individual relation type. Based on the relational-gated representations, we further propose the relation-guided minority synthesizer for over-sampling, which integrates the entity relational signatures to maintain relational consistency. Extensive experiments on 12 entity classification datasets provide compelling evidence for the superiority of Rel-MOSS, yielding an average improvement of up to 2.46% and 4.00% in terms of Balanced Accuracy and G-Mean, compared with SOTA RDL methods and classic methods for handling class imbalance.

Dexamethasone promotes neutrophil ROS-mediated tumor killing through the glucocorticoid receptor

Oncogene, Published online: 06 March 2026; doi:10.1038/s41388-026-03708-w

Dexamethasone promotes neutrophil ROS-mediated tumor killing through the glucocorticoid receptor

UniG2U-Bench: Do Unified Models Advance Multimodal Understanding?

arXiv:2603.03241v1 Announce Type: cross Abstract: Unified multimodal models have recently demonstrated strong generative capabilities, yet whether and when generation improves understanding remains unclear. Existing benchmarks lack a systematic exploration of the specific tasks where generation facilitates understanding. To this end, we introduce UniG2U-Bench, a comprehensive benchmark categorizing generation-to-understanding (G2U) evaluation into 7 regimes and 30 subtasks, requiring varying degrees of implicit or explicit visual transformations. Extensive evaluation of over 30 models reveals three core findings: 1) Unified models generally underperform their base Vision-Language Models (VLMs), and Generate-then-Answer (GtA) inference typically degrades performance relative to direct inference. 2) Consistent enhancements emerge in spatial intelligence, visual illusions, or multi-round reasoning subtasks, where enhanced spatial and shape perception, as well as multi-step intermediate image states, prove beneficial. 3) Tasks with similar reasoning structures and models sharing architectures exhibit correlated behaviors, suggesting that generation-understanding coupling induces class-consistent inductive biases over tasks, pretraining data, and model architectures. These findings highlight the necessity for more diverse training data and novel paradigms to fully unlock the potential of unified multimodal modeling.

Revisiting Graph Neural Networks for Graph-level Tasks: Taxonomy, Empirical Study, and Future Directions

arXiv:2501.00773v2 Announce Type: replace-cross Abstract: Graphs are fundamental data structures for modeling complex interactions in domains such as social networks, molecular structures, and biological systems. Graph-level tasks, which involve predicting properties or labels for entire graphs, are crucial for applications like molecular property prediction and subgraph counting. While Graph Neural Networks (GNNs) have shown significant promise for these tasks, their evaluations are often limited by narrow datasets, task coverage, and inconsistent experimental setups, hindering their generalizability. In this paper, we present a comprehensive experimental study of GNNs on graph-level tasks, systematically categorizing them into five types: node-based, hierarchical pooling-based, subgraph-based, graph learning-based, and self-supervised learning-based GNNs. To address these challenges, we propose a unified evaluation framework OpenGLT for graph-level GNNs. OpenGLT standardizes the evaluation process across diverse datasets, multiple graph tasks (e.g., classification and regression), and real-world scenarios, including noisy, imbalanced, and few-shot graphs. Extensive experiments are conducted on 16 baseline models across five categories, evaluated on 13 graph classification and 13 graph regression datasets. These experiments provide comprehensive insights into the strengths and weaknesses of existing GNN architectures.

ParaCook: On Time-Efficient Planning for Multi-Agent Systems

arXiv:2510.11608v2 Announce Type: replace Abstract: Large Language Models (LLMs) exhibit strong reasoning abilities for planning long-horizon, real-world tasks, yet existing agent benchmarks focus on task completion while neglecting time efficiency in parallel and asynchronous operations. To address this, we present ParaCook, a benchmark for time-efficient collaborative planning. Inspired by the Overcooked game, ParaCook provides an environment for various challenging interaction planning of multi-agent systems that are instantiated as cooking tasks, with a simplified action space to isolate the core challenge of strategic parallel planning. Through a comprehensive evaluation of state-of-the-art LLMs, we find that current approaches achieve suboptimal plans, which struggle with parallel actions or coordination. Our analysis also reveals LLMs' potential on abstract tasks where they can focus on high-level parallel optimization. ParaCook provides a scalable evaluation framework with adjustable complexity, establishing a foundation for developing and assessing time efficiency-aware multi-agent planning. The code and data are available at https://github.com/zsq259/ParaCook.
❌