❌

Normal view

Evaluation Faking: Unveiling Observer Effects in Safety Evaluation of Frontier AI Systems

arXiv:2505.17815v2 Announce Type: replace Abstract: As foundation models grow increasingly more intelligent, reliable and trustworthy safety evaluation becomes more indispensable than ever. However, an important question arises: Whether and how an advanced AI system would perceive the situation of being evaluated, and lead to the broken integrity of the evaluation process? During standard safety tests on a mainstream large reasoning model, we unexpectedly observe that the model without any contextual cues would occasionally recognize it is being evaluated and hence behave more safety-aligned. This motivates us to conduct a systematic study on the phenomenon of evaluation faking, i.e., an AI system autonomously alters its behavior upon recognizing the presence of an evaluation context and thereby influencing the evaluation results. Through extensive experiments on a diverse set of foundation models with mainstream safety benchmarks, we reach the main finding termed the observer effects for AI: When the AI system under evaluation is more advanced in reasoning and situational awareness, the evaluation faking behavior becomes more ubiquitous, which reflects in the following aspects: 1) Reasoning models recognize evaluation 16% more often than non-reasoning models. 2) Scaling foundation models (32B to 671B) increases faking by over 30% in some cases, while smaller models show negligible faking. 3) AI with basic memory is 2.3x more likely to recognize evaluation and scores 19% higher on safety tests (vs. no memory). To measure this, we devised a chain-of-thought monitoring technique to detect faking intent and uncover internal signals correlated with such behavior, offering insights for future mitigation studies.

Effects of Digital Health Interventions on Functional and Psychological Outcomes in Older Patients With Hip Fractures: Systematic Review and Meta-Analysis of Randomized Controlled Trials

Background: Hip fractures in older adults increasingly challenge public health, making traditional rehabilitation very challenging. Digital health interventions (DHIs) have emerged as a promising solution for postoperative rehabilitation. However, evidence on DHIs’ effects on functional and psychological outcomes remains insufficient. Objective: This systematic review aimed to comprehensively examine the effects of DHIs on functional and psychological outcomes in older adults with hip fractures. Methods: Following PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) guidelines, we searched 9 databases (PubMed, Embase, CENTRAL, APA PsycINFO, Web of Science, PEDro, CNKI, WANFANG, and SinoMed) from inception to November 13, 2025. Included studies enrolled adults aged 60 years and older with hip fractures, delivered DHIs, assessed functional and psychological outcomes, set usual care or no intervention as the control, and had a randomized controlled trial design. Studies were excluded if they enrolled nonhospitalized patients in the emergency department, patients discharged to nonhome settings, or had inaccessible full text or insufficient data. Study quality was evaluated using the Cochrane Risk of Bias tool 2.0 (Cochrane Collaboration), and evidence certainty was assessed using GRADE (Grading of Recommendations, Assessment, Development and Evaluation). The literature screening, data extraction, and quality assessment were independently conducted by 2 researchers, and any disputes were resolved by the third researcher. We performed analysis using R version 4.0.3 (R Foundation for Statistical Computing) with a random-effects model. Results: Of 17,723 studies screened, 13 met the inclusion criteria. DHIs, compared to the control, significantly improved hip function (standardized mean difference [SMD] 0.80, 95% CI 0.33-1.26; 95% prediction interval [PI] –0.24 to 1.83; P=.007) and functional independence (SMD 1.23, 95% CI 0.34-2.11; 95% PI –0.98 to 3.34; P=.02). Despite favorable pooled effects, a wide 95% PI spanning positive or negative values signals substantial heterogeneity. No significant difference was observed in balance function, risk of falling, and quality of life. Only a single available study reported a 70% adherence rate in the DHIs group. Subgroup analyses stratified by intervention duration revealed no significant intersubgroup differences for hip function (χ12=0.1; P=.75) or functional independence (χ12=2.93; P=.09). For hip function, the point estimate favored the 3 months subgroup (SMD 0.89, 95% CI 0.36-1.41; I2=7%; P=.41) over the <3 months subgroup. Conversely, for functional independence, the point estimate favored shorter intervention duration (SMD 0.67, 95% CI 0.12-1.23; I²=0%; P=.72). Conclusions: This review incorporates the latest randomized controlled trials and comprehensively assesses functional and psychological outcomes of DHIs in older patients with hip fractures, distinct from prior studies focusing solely on functional outcomes. While the 95% CI supports the potential of DHIs to improve hip function and functional independence, the wide 95% PI indicating substantial real-world response variability, which calls for cautious interpretation, informs the design of targeted DHI-based rehabilitation regimens, warranting further research into optimal techniques and dosages in clinical practice. Trial Registration: PROSPERO CRD42024626186; https://www.crd.york.ac.uk/PROSPERO/view/CRD42024626186

Rel-MOSS: Towards Imbalanced Relational Deep Learning on Relational Databases

arXiv:2603.07916v1 Announce Type: new Abstract: In recent advances, to enable a fully data-driven learning paradigm on relational databases (RDB), relational deep learning (RDL) is proposed to structure the RDB as a heterogeneous entity graph and adopt the graph neural network (GNN) as the predictive model. However, existing RDL methods neglect the imbalance problem of relational data in RDBs and risk under-representing the minority entities, leading to an unusable model in practice. In this work, we investigate, for the first time, class imbalance problem in RDB entity classification and design the relation-centric minority synthetic over-sampling GNN (Rel-MOSS), in order to fill a critical void in the current literature. Specifically, to mitigate the issue of minority-related information being submerged by majority counterparts, we design the relation-wise gating controller to modulate neighborhood messages from each individual relation type. Based on the relational-gated representations, we further propose the relation-guided minority synthesizer for over-sampling, which integrates the entity relational signatures to maintain relational consistency. Extensive experiments on 12 entity classification datasets provide compelling evidence for the superiority of Rel-MOSS, yielding an average improvement of up to 2.46% and 4.00% in terms of Balanced Accuracy and G-Mean, compared with SOTA RDL methods and classic methods for handling class imbalance.

Dexamethasone promotes neutrophil ROS-mediated tumor killing through the glucocorticoid receptor

Oncogene, Published online: 06 March 2026; doi:10.1038/s41388-026-03708-w

Dexamethasone promotes neutrophil ROS-mediated tumor killing through the glucocorticoid receptor

UniG2U-Bench: Do Unified Models Advance Multimodal Understanding?

arXiv:2603.03241v1 Announce Type: cross Abstract: Unified multimodal models have recently demonstrated strong generative capabilities, yet whether and when generation improves understanding remains unclear. Existing benchmarks lack a systematic exploration of the specific tasks where generation facilitates understanding. To this end, we introduce UniG2U-Bench, a comprehensive benchmark categorizing generation-to-understanding (G2U) evaluation into 7 regimes and 30 subtasks, requiring varying degrees of implicit or explicit visual transformations. Extensive evaluation of over 30 models reveals three core findings: 1) Unified models generally underperform their base Vision-Language Models (VLMs), and Generate-then-Answer (GtA) inference typically degrades performance relative to direct inference. 2) Consistent enhancements emerge in spatial intelligence, visual illusions, or multi-round reasoning subtasks, where enhanced spatial and shape perception, as well as multi-step intermediate image states, prove beneficial. 3) Tasks with similar reasoning structures and models sharing architectures exhibit correlated behaviors, suggesting that generation-understanding coupling induces class-consistent inductive biases over tasks, pretraining data, and model architectures. These findings highlight the necessity for more diverse training data and novel paradigms to fully unlock the potential of unified multimodal modeling.

Revisiting Graph Neural Networks for Graph-level Tasks: Taxonomy, Empirical Study, and Future Directions

arXiv:2501.00773v2 Announce Type: replace-cross Abstract: Graphs are fundamental data structures for modeling complex interactions in domains such as social networks, molecular structures, and biological systems. Graph-level tasks, which involve predicting properties or labels for entire graphs, are crucial for applications like molecular property prediction and subgraph counting. While Graph Neural Networks (GNNs) have shown significant promise for these tasks, their evaluations are often limited by narrow datasets, task coverage, and inconsistent experimental setups, hindering their generalizability. In this paper, we present a comprehensive experimental study of GNNs on graph-level tasks, systematically categorizing them into five types: node-based, hierarchical pooling-based, subgraph-based, graph learning-based, and self-supervised learning-based GNNs. To address these challenges, we propose a unified evaluation framework OpenGLT for graph-level GNNs. OpenGLT standardizes the evaluation process across diverse datasets, multiple graph tasks (e.g., classification and regression), and real-world scenarios, including noisy, imbalanced, and few-shot graphs. Extensive experiments are conducted on 16 baseline models across five categories, evaluated on 13 graph classification and 13 graph regression datasets. These experiments provide comprehensive insights into the strengths and weaknesses of existing GNN architectures.

ParaCook: On Time-Efficient Planning for Multi-Agent Systems

arXiv:2510.11608v2 Announce Type: replace Abstract: Large Language Models (LLMs) exhibit strong reasoning abilities for planning long-horizon, real-world tasks, yet existing agent benchmarks focus on task completion while neglecting time efficiency in parallel and asynchronous operations. To address this, we present ParaCook, a benchmark for time-efficient collaborative planning. Inspired by the Overcooked game, ParaCook provides an environment for various challenging interaction planning of multi-agent systems that are instantiated as cooking tasks, with a simplified action space to isolate the core challenge of strategic parallel planning. Through a comprehensive evaluation of state-of-the-art LLMs, we find that current approaches achieve suboptimal plans, which struggle with parallel actions or coordination. Our analysis also reveals LLMs' potential on abstract tasks where they can focus on high-level parallel optimization. ParaCook provides a scalable evaluation framework with adjustable complexity, establishing a foundation for developing and assessing time efficiency-aware multi-agent planning. The code and data are available at https://github.com/zsq259/ParaCook.
❌