❌

Normal view

Performance of Conventional EEG Biomarkers Across Different Clinical Phases of Major Depressive Disorder: A Comprehensive Evaluation

arXiv:2603.03864v1 Announce Type: new Abstract: While EEG features differentiate Major Depressive Disorder (MDD) from healthy controls (HC), their clinical utility as biomarkers depends on a monotonic trajectory across the disease spectrum, from the acute (AC) phase to the maintenance (MA) phase and finally to the healthy baseline. However, the progression of the MA phase remains poorly understood in traditional marker analysis. Analyzing EEG data from 74 individuals (24 AC, 23 MA, and 27 HC), this study provides a comprehensive evaluation of classic ERP and resting-state indices across AC, MA, and HC groups. Our results demonstrate that almost no conventional metrics strictly satisfy the criterion of monotonic progression, likely due to profound inter-individual heterogeneity. These findings highlight the inherent limitations of group-level feature extraction and provide critical insights for developing future paradigms and algorithms to identify neurobiological markers with genuine clinical utility.

TrustMH-Bench: A Comprehensive Benchmark for Evaluating the Trustworthiness of Large Language Models in Mental Health

arXiv:2603.03047v1 Announce Type: cross Abstract: While Large Language Models (LLMs) demonstrate significant potential in providing accessible mental health support, their practical deployment raises critical trustworthiness concerns due to the domains high-stakes and safety-sensitive nature. Existing evaluation paradigms for general-purpose LLMs fail to capture mental health-specific requirements, highlighting an urgent need to prioritize and enhance their trustworthiness. To address this, we propose TrustMH-Bench, a holistic framework designed to systematically quantify the trustworthiness of mental health LLMs. By establishing a deep mapping from domain-specific norms to quantitative evaluation metrics, TrustMH-Bench evaluates models across eight core pillars: Reliability, Crisis Identification and Escalation, Safety, Fairness, Privacy, Robustness, Anti-sycophancy, and Ethics. We conduct extensive experiments across six general-purpose LLMs and six specialized mental health models. Experimental results indicate that the evaluated models underperform across various trustworthiness dimensions in mental health scenarios, revealing significant deficiencies. Notably, even generally powerful models (e.g., GPT-5.1) fail to maintain consistently high performance across all dimensions. Consequently, systematically improving the trustworthiness of LLMs has become a critical task. Our data and code are released.
❌