❌

Normal view

Received — 10 September 2026 ⏭ IEEE Journal of Biomedical and Health Informatics - new TOC

DualBAN: Unifying Intra- and Inter-Molecular Features for Compound-Protein Interaction Prediction

Accurate identification of compound–protein interactions (CPIs) is critical for drug discovery. In recent years, neural network–based CPI prediction methods have demonstrated remarkable performance. However, most existing approaches primarily focus on interaction patterns derived from known data, without fully leveraging both intra- and inter-molecular interaction information within compound–protein pairs. This limitation constrains the representation learning capability of models for compounds and proteins, hindering further improvements in predictive accuracy. In this paper, we propose DualBAN, a novel CPI prediction model that integrates intra- and inter-molecular interaction information from both compounds and proteins. Specifically, DualBAN employs pretrained biological large language models to obtain sequence features and extracts atomic features of compounds and residue representations of proteins. To comprehensively capture intra- and inter-molecular interactions, DualBAN fuses atomic and residue representations using a bilinear attention network and combines sequence representations through cross-attention, jointly utilizing both components for CPI prediction. Extensive experiments demonstrate that the proposed DualBAN significantly outperforms state-of-the-art methods on CPI prediction tasks and maintains robust performance under cross-domain and cold-start settings.

$\text{P}^{{2}}$RS: A Quantitative Rating Scale for Pain Assessment Based on Pulse Wave Characterization

For pain intensity assessment, currently there are mainly 11 rating scales, from primitive Visual Analog Scale (VAS) to elaborate Measure of Intermittent and Constant Osteoarthritis Pain (ICOAP). However, they all depend on a self-report mechanism, making their results so subjective that the consistency, comparability and reference value are barely satisfactory. Inspired by the phenomenon that discomfort may give rise to the throbbing of radial artery, we develop an objective rating scale innovatively, quantifying the severity of pain by the degree of “lateral instability” of an arterial pulse wave. In attempting to monitor this lateral instability, a sort of ultra-small piezoresistive pressure sensor is fabricated in an area of 0.4 × 0.4 $\text{mm}^{{2}}$ . With 18 of such sensors, we build a flexible tactile sensing dense-array with a pitch of only 0.65 mm. Overlying the radial artery perpendicularly to the blood flow direction, the dense-array succeeds in observing the cross-section of a pulse wave. The barycenter of the cross-section of each wave cycle is taken as the feature point to represent its lateral shape and drift. The standard deviation of the barycenters’ horizontal coordinates is thereby calculated as the pulsatile perceptual rating scale ($\text{P}^{{2}}$RS) to reflect the degree of lateral instability, that is, our scale of pain intensity. Among 86 clinical samples, the pain threshold is 0.11, which is concluded by a binary classification model based on a support vector machine. In terms of its consistency with previous rating scales, the average correlation coefficient reaches 0.804 among 43 pain samples.

Graph-Informed and FiLM-Enhanced Multimodal Fusion for Myocardial Infarction Prediction

Accurate and timely diagnosis of cardiovascular diseases, particularly myocardial infarction (MI), remains a critical clinical challenge. Existing electrocardiogram (ECG) analysis methods often rely solely on asingle data modality, such as raw signals or waveform images, which limits their ability to capture thebroader physiological context. To address this limitation, we propose GFM-MIP, a Graph-informed and FiLM-enhanced Multimodal Fusion framework for myocardial infarction prediction. GFM-MIP integrates 12-lead ECG time-series signals, ECG images, and laboratory test results through a unified architecture. Specifically, it employs a Graphormer encoder to model inter-lead dependencies in ECG signals and a Vision Transformer to extract morphological patterns from ECG images, both modulated by patient-specific laboratory features using Feature-wise Linear Modulation (FiLM). A Transformer-based fusion module captures cross-modal interactions, while a contrastive learning objective encourages alignment between signal and image modalities. Experimental results on a real-world clinical dataset and three public benchmarks demonstrate that GFM-MIP consistently outperforms state-of-the-art baselines across multiple evaluation metrics. Ablation studies further validate the contribution of each modality and architectural component. The proposed framework offers a clinically meaningful and scalable solution for robust, multimodal cardiovascular diagnosis.

Bond-Aware Molecular Graph Learning With Multi-Graph Interleaved Message Passing

Graph neural networks (GNNs) have demonstrated remarkable capabilities in molecular property prediction. Existing approaches adopt GNNs by modeling molecules as homogeneous graphs. However, the bonds between atoms can be heterogeneous, whose characterization and role in molecular graph representation learning remain unexplored. To address the heterogeneity issue inherent in molecular graphs, in this work, we build the bond-centric graphs and propose a novel multi-graph learning model, which captures the bond heterogeneity via augmented bond graph view and bond coding for atom features. Different from conventional multi-view learning that focus on late-stage view fusion, our method integrates cross-graph information during the node representation learning phase. Towards this end, we introduce the interleaved message passing graph neural network (IMPGNN), allowing the messages passing across three views of the molecular graph. Moreover, we introduce a novel structure-aware pooling mechanims for graph representation, which yields up to 45.7% gains over simple sum pooling. Comparative experiments on two standard molecular property prediction tasks reveal that our method surpasses all competing approaches (including multimodal models) on 75% of the evaluated benchmark datasets.

Nonparametric Dynamic Granger Causality Based on Multi-Space Spectrum Fusion for Time-Varying Directed Brain Network Construction

Nonparametric estimation of time-varying directed networks can unveil the intricate transient organization of directed brain communication while circumventing constraints imposed by prescribed model-driven methods. A robust time-frequency representation – the foundation of its causality inference – is critical for enhancing its reliability. This study proposed a novel method, i.e., nonparametric dynamic Granger causality based on Multi-space Spectrum Fusion (ndGCMSF), which integrates complementary spectrum information from different spaces to generate enhanced spectral representations to estimate dynamic causalities across brain regions. Systematic simulations and validations demonstrate that ndGCMSF exhibits superior noise resistance and a powerful ability to capture subtle dynamic changes in directed brain networks. Particularly, ndGCMSF revealed that during motor imagery, the laterality in the hemisphere ipsilateral to the hemiplegic limb emerges upon task beginning and diminishes upon task accomplishment. These intrinsic variations further provide features for assessing motor functions. The ndGCMSF offers powerful functional patterns to derive effective brain networks in dynamically changing operational settings and contributes to extensive areas involving dynamical and directed communications.

WGB-GLFI: A Novel Graph-Based Global-Local Feature Interaction Framework for Automated Seizure Detection

Epilepsy detection faces significant challenges due to unpredictable seizures, ranging from brief awareness lapses to severe convulsions, posing risks to patients’ safety and quality of life. In recent years, deep learning has become a mainstream approach in this field, leveraging advanced computational resources and EEG datasets. However, a key challenge remains: existing methods often lack unified spatial modeling and struggle to effectively handle local detailed features, thereby limiting their accuracy and robustness. To address these issues, we propose the Weighted Graph Building Global-Local Feature Interaction (WGB-GLFI) framework, which integrates spatial connectivity and dynamic patterns through a Weighted Graph Building (WGB) module and a Global-Local Feature Interaction (GLFI) module. This approach excels by comprehensively capturing the dynamic spatial relationships during epileptic seizures and achieving seamless global-local feature integration, significantly enhancing seizure detection performance. Its effectiveness has been validated across multiple datasets, including CHB-MIT, Siena Scalp, and private datasets, demonstrating robust and reliable results. Evaluated on these datasets, our model achieves accuracy rates of 99.28%, 99.21%, and 99.30%, respectively. The reliability and robustness of our framework provide epilepsy patients with faster and more reliable seizure detection, which helps to intervene in a timely manner and improve the quality of life of patients.

RT-SAM: Visual-Prompt Fusion and Uncertainty Enhancement for Nasopharyngeal Carcinoma Radiotherapy Target Delineation

Precise delineation of the clinical target volume (CTV) and nodal CTV (CTV$_{{\mathit{nd}}}$) is crucial for effective radiotherapy planning in nasopharyngeal carcinoma (NPC). Manual contouring is labor-intensive and subject to substantial inter-observer variability, particularly in regions with complex anatomy and indistinct boundaries. This study presents RT-SAM, a novel framework that adapts the Medical Segment Anything Model 2 (MedSAM-2) for automated CTV (i.e., primary CTV and CTV$_{nd}$) contouring in NPC computed tomography (CT) images. The framework synergistically integrates a generalist foundation model (MedSAM-2) with a domain-specific specialist network (2D U-Net) through three principal contributions: (1) automated generation of multi-modal prompts—comprising mask, bounding box, and point representations—derived from specialist network predictions to guide the generalist model; (2) a Visual-Prompt Fusion Attention (ViPFA) mechanism that optimizes feature-prompt interactions through bidirectional cross-modal attention; and (3) an Uncertainty-Enhanced Prediction Adjustment (UEPA) mechanism that enhances model robustness via confidence-based refinement and selective domain adaptation. Comprehensive evaluation on a multi-center cohort of 256 clinical NPC cases from Sun Yat-sen University Cancer Center and 212 public NPC cases from the SegRap2025 lymph node CTV dataset using 5-fold cross-validation demonstrates that RT-SAM achieves a mean DICE coefficient of 0.796 $\pm$ 0.033 (mean $\pm$ standard deviation), significantly outperforming current state-of-the-art methods. Clinical validation by eight radiation oncologists demonstrates that RT-SAM contours are clinically indistinguishable from expert delineations in blinded Turing assessments, achieve superior quality ratings in 75% of comparisons with mean scores of 2.73 for RT-SAM versus 2.66 for manual expert contours, and attain clinically acceptable ratings in over 97% of cases. These results demonstrate that RT-SAM is a clinically feasible solution for automated CTV contouring, with strong potential to standardize treatment planning and mitigate inter-observer variability in NPC radiotherapy.

R2GenCSR: Mining Contextual and Residual Information for LLMs-Based Radiology Report Generation

Inspired by the tremendous success of Large Language Models (LLMs), existing Radiology report generation methods attempt to leverage large models to achieve better performance. They usually adopt a Transformer to extract the visual features of a given X-ray image, and then, feed them into the LLM for text generation. How to extract more effective information for the LLMs to help them improve final results is an urgent problem that needs to be solved. Additionally, the use of visual Transformer models also brings high computational complexity. To address these issues, this paper proposes a novel context-guided efficient radiology report generation framework. Specifically, we introduce the Mamba as the vision backbone with linear complexity, and the performance obtained is comparable to that of the strong Transformer model. More importantly, we perform context retrieval from the training set for samples within each mini-batch during the training phase, utilizing both positively and negatively related samples to enhance feature representation and discriminative learning. Subsequently, we feed the vision tokens, context information, and prompt statements to invoke the LLM for generating high-quality medical reports. Extensive experiments on three X-ray report generation datasets (i.e., IU X-Ray, MIMIC-CXR, CheXpert Plus) fully validated the effectiveness of our proposed model.

Multi-Level Asymmetric Contrastive Learning for Medical Image Segmentation Pre-Training

Medical image segmentation. is a fundamental yet challenging task due to the arduous process of acquiring large volumes of high-quality labeled data from experts. Contrastive learning offers a promising but still problematic solution to this dilemma. Firstly existing medical contrastive learning strategies focus on extracting image-level representation, which ignores abundant multi-level representations. Furthermore they underutilize the decoder either by random initialization or separate pre-training from the encoder, thereby neglecting the potential collaboration between the encoder and decoder. To address these issues, we propose a novel multi-level asymmetric contrastive learning framework named MACL for enhancing medical image segmentation. Specifically, we design an asymmetric contrastive learning structure to pre-train encoder and decoder simultaneously to provide better initialization for segmentation models. Moreover, we develop a multi-level contrastive learning strategy that integrates correspondences across feature-level, image-level, and pixel-level representations to ensure the encoder and decoder capture comprehensive details from representations of varying scales and granularities during the pre-training phase. Finally, experiments on 8 medical image datasets indicate our MACL framework outperforms existing 11 contrastive learning strategies. i.e. Our MACL achieves a superior performance with more precise predictions from visualization figures and 1.72%, 7.87%, 2.49% and 1.48% Dice higher than previous best results on ACDC, MMWHS, HVSMR and CHAOS with 10% labeled data, respectively. And our MACL also has a strong generalization ability among 5 variant U-Net backbones.

Revealing Sleep Dynamics With PCT-CRV: A Novel Approach for Automatic Sleep Staging and Tracking Transitions Using PSG Signals

Polysomnography (PSG)-based accurate sleep staging is essential to monitor sleep quality and sleep-related disorders. Despite previous attempts for improving the performance of automatic sleep staging, there are certain limitations: 1) neglecting synchronization patterns in their time-frequency (TF) domain, 2) not utilizing both local and global features within sleep epochs, and 3) neglecting correlation patterns for tracking transitions between sleep stages. To address them, we propose a novel framework based on the polynomial chirplet transform-derived characteristic response vector (PCT-CRV) for the assessment of sleep stages. In this work, we perform the time-domain PCT (TPCT) and frequency-domain PCT (FPCT) to enhance the TF representation of nonstationary PSG signals. From these PCT representations, we construct correlation matrices across their frequency bins within short-time windows to obtain characteristic response vectors (CRVs), which are the sums of eigenvectors, weighted by their corresponding eigenvalues. Subsequently, a comprehensive set of local and global features is derived from PCT-CRVs, which is subjected to various machine learning-based classifiers. Our PCT-CRV excels on three datasets, surpassing existing methods, and outperforming wavelet-based and synchrosqueezed-based CRV methods. Furthermore, to track transitions of sleep stages, we form sub-band PCT-CRVs using eigenvectors with maximum information, depending upon the physics of our problem. We hypothesize that sleep stages are characterized by specific correlation profiles, within different frequency bins. Hence, sub-band PCT-CRVs corresponding to the dominant eigenvectors, would detect transition of sleep stages across all epochs. All these results highlight the efficacy of our method in tracking sleep stage transitions and improving their classification performance.

Fall Warning Method Based on Multimodal Sensor Fusion and Gait Phase Detection

Falls are a common and serious cause of injury among the elderly and individuals with mobility impairments. In particular, under complex gait conditions, the early detection of imbalance is crucial for fall prevention. To address the limitations of existing methods in fall phase identification and the scarcity of real fall data, this study proposes a fall warning method based on multimodal sensor fusion and gait phase detection. By combining data from plantar pressure sensors and inertial measurement units, a gait phase detection module is introduced to achieve fine division of the gait cycle, enhancing the system’s ability to detect early imbalance features. Additionally, a hybrid dataset integrating simulation data with real data is constructed, and multiple linear regression is used to accurately map simulation and real data, mitigating the issue of limited samples. Experimental results demonstrate that the proposed method achieves an accuracy of 94.8%, a recall of 92.8%, and a precision of 94.2%. It further maintains stable performance in cross-subject tests and multi-scenario evaluations, demonstrating strong reliability and generalization capability.

Mamba-Based Prototypical Contrastive Learning With Augmented Feature Separation for Common and Rare Arrhythmia Classification

Early diagnosis of arrhythmia, a common cardiovascular condition, is crucial for improving prognosis. Electrocardiogram (ECG) is widely used as a non-invasive diagnostic tool. However, Computer-Aided Diagnosis of rare arrhythmias faces significant challenges due to the severe scarcity of samples for these rare disease classes. To tackle this, we propose a Mamba-based Prototypical Contrastive Learning framework, which can simultaneously identify both common and rare classes under the setting of generalized Few-Shot Learning (FSL). It primarily consists of: (1) the Mamba-based Spatio-Temporal Feature Fusion Network (MST), which integrates spatial features from multi-scale convolutions and temporal dynamics from bidirectional Mamba for ECG modeling; (2) the Prototypical Contrastive Learning framework with Augmented Feature Separation (PCAS), which employs a prototype augmentation strategy with an Augmented Prototype Consistency Loss to optimize prototype representations, and an Separation-Tuned Contrastive Loss to enhance intra-class compactness and inter-class distinctnessy, mitigating the risk of class collapse. Extensive experiments on publicly available datasets PTBXL and Chapman demonstrate the effectiveness of MST-PCAS, achieving superior rare-class recognition accuracies of 79.13% and 50.72%, respectively, for ECG arrhythmia classification.

ReMol: A Chemical Reaction Knowledge-Guided Self-Supervised Molecular Image Representation Learning Framework

Molecular representation learning (MRL) is critical in computational chemistry and drug discovery, paving the way for efficient molecular properties and biological activity prediction. However, existing sequence-based or graph-based MRL methods emphasize static and intrinsic molecular topological features while ignoring dynamic and interactive chemical knowledge, resulting in insufficient generalization ability. MolR meets this challenge by leveraging the equivalence of molecules participating in chemical reactions in embedding space to assist in learning molecular representations. However, it can suffer from unsatisfactory performance because it lacks reaction center information and the complex relationship between reactions, which provides a deeper understanding of the chemical processes. We propose ReMol, an elaborate chemical reaction knowledge-guided self-supervised molecular image representation learning framework to address this issue. The ReMol framework integrates comprehensive reaction inductive biases, including reaction templates and consistency, and diversity in chemical reactions. Experimental results demonstrate that our framework achieves state-of-the-art results compared with cutting-edge methods on various challenging downstream tasks, such as chemical reaction and molecular property prediction tasks. Overall, our work offers a robust tool for advancing chemistry research, with the potential to make significant contributions to both molecular representation learning and drug discovery.
  • ✇IEEE Journal of Biomedical and Health Informatics - new TOC
  • SkeDiff: Skeleton 3D CT Diffusion Reconstruction Using 2D X-Ray
    For orthopedic diagnostics, both 2D X-ray and 3D CT imaging play essential roles. X-ray imaging is widely accessible, clinically effective, easy to operate, and has lower radiation exposure than CT. However, its inherent 2D nature limits comprehensive visualization of skeletal structures, which 3D CT provides. To bridge this gap, we propose SkeDiff, an algorithm for reconstructing 3D CT images of the skeleton from orthogonal 2D X-ray projections. To fully leverage the information in X-ray images
     

SkeDiff: Skeleton 3D CT Diffusion Reconstruction Using 2D X-Ray

For orthopedic diagnostics, both 2D X-ray and 3D CT imaging play essential roles. X-ray imaging is widely accessible, clinically effective, easy to operate, and has lower radiation exposure than CT. However, its inherent 2D nature limits comprehensive visualization of skeletal structures, which 3D CT provides. To bridge this gap, we propose SkeDiff, an algorithm for reconstructing 3D CT images of the skeleton from orthogonal 2D X-ray projections. To fully leverage the information in X-ray images for guiding the diffusion process, we design a cross-dimensional conditional encoder, $E_{Cond}$, to extract 2D priors for the 3D diffusion model, $DM_{3DL}$. This encoder integrates a CNN-Mamba hybrid architecture to enhance feature extraction and nonlinear mapping. Additionally, we introduce a 3D UKAN diffusion backbone, which employs Kolmogorov-Arnold network (KAN) to improve feature representation through learnable nonlinear activations. Furthermore, we propose a diffusion-based scoliosis classifier, $D_{SC}$, enabling scoliosis classification during the 3D CT reconstruction process. Experiments show that SkeDiff outperforms recent algorithms on spine, hip, and knee datasets.

Topology-Aware Diffusion Schrödinger Bridge for Unpaired H&E-to-IHC Stain Translation

Unpaired H&E-to-IHC Stain Translation aims to generate immunohistochemistry (IHC) staining from Hematoxylin and Eosin (H&E) staining. It offers clearer diagnostic insights and potentially expands access to advanced pathology services in resource-limited areas. This task faces two primary challenges: capturing target domain style characteristics and preserving topological features in histological images. Recently, Schrödinger Bridge (SB)-based methods have offered a solution for unpaired image-to-image translation, addressing the mode collapse and artifact issues in CycleGAN-based approaches, as well as the Gaussian prior assumption limitation in diffusion-based methods. While SB-based methods suffer from the curse of dimensionality with high-resolution images, the Unpaired Neural Schrödinger Bridge (UNSB) overcomes this challenge and achieves state-of-the-art (SOTA) performance on natural images. However, UNSB has two key issues in histological images: (1) loss of topological features and (2) IHC staining representation. UNSB focuses only on the optimal path from source to target domains, ignoring local structure paths. Convolutional neural networks (CNNs) do not perfectly preserve critical anatomical structures due to limitations like receptive field size or model capacity. To address these challenges, we introduce the Topology-aware Diffusion Schrödinger Bridge (TDSB), integrating a Topology Guidance (TG) module and Dual-Domain Adaptive Patch-based noise contrastive estimation (DDAP). Experiments on seven translation tasks across three datasets show that our method achieves SOTA performance in unpaired H&E-to-IHC stain translation. Clinical evaluation through pathologists' assessments further validates the effectiveness of our method.

A Novel Eigen-Volume-based Co-Activation Pattern Framework for Dynamic Functional Biomarkers of Multiple Sclerosis

Imaging biomarkers are essential for monitoring multiple sclerosis (MS), wand resting-state functional MRI (rs-fMRI) offers functional insights that complement structural imaging. This study investigates whether a novel co-activation pattern (CAP) approach for dynamic rs-fMRI can function as a dual-purpose biomarker in MS, aiding diagnosis and tracking disease severity. RS-fMRI scans from 25 relapsing-remitting MS patients and 41 healthy controls (HCs) were analyzed using a novel CAP-based approach. CAPs derived from individual time frames to capture dynamic brain activity patterns incorporated a bivariate similarity assessment, eigen-volume-based dimensionality reduction, and consensus clustering. We evaluated the framework in two analyses: (1) a diagnostic evaluation, using dynamic CAP features—dwell time, persistence, and transition probabilities—for group comparisons and classification; and (2) a severity-prediction analysis, relating these CAP-derived measures to clinical disability (EDSS) in MS using LASSO regression. Method performance was benchmarked against standard CAP and sliding-window (SW) approaches. It revealed significant differences in brain activity between MS and HCs, within the default mode, sensorimotor, and language networks (p 0.75) and yielded better classification performance than standard CAP and SW approaches in classifying MS from HCs. These results suggest that dynamic brain activity patterns are altered in MS and linked to clinical disability. The proposed CAP provided improved performance in distinguishing MS patients, offering enhanced clinical monitoring. Transition probabilities emerged as a potential biomarker for tracking MS progression, with network shifts reflecting disease severity. As MS advances, increased transitions toward sensory, motor, and executive networks suggest compensatory recruitment. Conversely, reduced transitions from default mode and salience networks to sensorimotor and frontoparietal systems were associated with greater disability and diminished adaptive reorganization.
  • ✇IEEE Journal of Biomedical and Health Informatics - new TOC
  • EIT to CT Cross-Modality Translation Using Diffusion Transformer
    Computed Tomography (CT) plays a crucial role in medical imaging due to its superior spatial resolution and diagnostic capability, particularly in pulmonary applications. However, the associated radiation risk, high equipment cost, and limitations in accessibility present significant challenges. In contrast, Electrical Impedance Tomography (EIT) offers a radiation-free, cost-effective, and portable imaging alternative, making it attractive in bedside and resource-limited settings. In this work,
     

EIT to CT Cross-Modality Translation Using Diffusion Transformer

Computed Tomography (CT) plays a crucial role in medical imaging due to its superior spatial resolution and diagnostic capability, particularly in pulmonary applications. However, the associated radiation risk, high equipment cost, and limitations in accessibility present significant challenges. In contrast, Electrical Impedance Tomography (EIT) offers a radiation-free, cost-effective, and portable imaging alternative, making it attractive in bedside and resource-limited settings. In this work, we demonstrate the feasibility of generating high-resolution synthetic CT (sCT) images from raw EIT voltages, enhancing EIT’s clinical application by effectively unifying the benefits of both modalities. We propose an EIT-to-CT translation approach that leverages a diffusion transformer conditioned on the voltage measurements to guide sCT image generation through the reverse diffusion process. The model was trained on a synthetic dataset and evaluated on both the simulated and experimental data, resulting in anatomically coherent and visually realistic sCT images. Quantitative assessment demonstrates the method’s effectiveness on both pixel-level and distribution-based metrics. It achieved a Correlation Coefficient of 0.9282, Peak Signal-to-Noise ratio of 24.64 dB, Feature Similarity Index of 0.8321, and deep Image Structure and Texture Similarity of 0.1289, indicating a high level of structural consistency. Similarly, it attained strong alignment with the targeted CT distribution with a Fréchet Inception Distance of 33.36 and Kernel Inception Distance of 0.0110. The clinical relevance of the sCT images was further validated through a downstream lung cancer classification task, where models trained on real CT showed comparable performance on sCT, and a pilot reader study, where clinical experts rated over 90% of the generated images as anatomically correct and realistic.

BioFPT: Biosignal Feature Pyramid Transformer for Self-Supervised Representation Learning From ECG Signals

Electrocardiogram (ECG) analysis represents a promising field for deep learning applications in clinical diagnostics. However, practical use of current methods is still constrained by their heavy reliance on large amounts of labeled data, as well as limitations in processing efficiency and signal quality. To address these challenges, we present BioFPT (Biosignal Feature Pyramid Transformer), a novel self-supervised learning framework designed for ECG signal. The proposed framework incorporates a Split-Mask-Join (SMJ) transformation for a Pre-training strategy, complemented by an overlapping embedding mechanism that eliminates positional encoding requirements. The efficiency of the architectural design is enhanced through a Spatial Reduction Attention (SRA) transformer, which achieves a reduction in computational complexity without performance degradation. Comprehensive evaluation of seven public ECG datasets comprising over 94,000 subjects demonstrates BioFPT’s effectiveness, an accuracy improvement of 4.2% and a parameter reduction of 14.8% compared to state-of-the-art models. Furthermore, it maintains robust performance across diverse pathological conditions and signal qualities. The proposed architecture represents a significant advancement in self-supervised ECG analysis, particularly suitable for scenarios with limited labeled data availability. Moreover, its versatile architecture shows promise for broader applications across various biosignals.

CLDAE: A Two Stage EEG-Based Emotion Recognition Framework Combining Contrastive Learning and Dual-Attention Encoder

Electroencephalogram (EEG)-based emotion recognition systems face a persistent challenge in maintaining robust performance across subjects (generalization) and within subjects (personalization). Existing models for cross-subject recognition generally struggle to adapt to individual-specific neural signatures, while models with optimized within-subject performance typically require a large amount of personalized data. To address these limitations, this study proposes an EEG-based emotion recognition framework, CLDAE, that integrates a contrastive learning strategy and a dual-attention feature extraction mechanism. The CLDAE framework includes two stages: contrastive learning pre-training and emotion recognition fine-tuning. During the pre-training stage, a data augmentation method that combines EEG signals from different subjects is used to generate new training samples. Moreover, to extract discriminative features from the augmented data, the dual-attention encoder combines temporal and channel attention mechanisms. After pre-training, the CLDAE is fine-tuned for final recognition tasks. The proposed CLDAE is verified by experiments on two public datasets (DEAP and SEED-IV) and a private dataset (MAN). The experimental results demonstrate that the CLDAE achieves competitive performance in both within-subject and cross-subject emotion recognition, with 95.12% and 75.29% accuracy on the MAN dataset, respectively; thus, outperforming the baseline methods. These results validate the effectiveness of the proposed framework in both within-subject and cross-subject emotion recognition.

A Multi-Scale Attention-Based Reconstruction Fusion Network for Motor Imagery Classification

Motor imagery (MI) is a widely used cognitive paradigm in brain–computer interface (BCI) systems, where accurate and efficient MI decoding is essential for real-time human–machine interaction. However, the non-stationary nature and pronounced inter-subject variability of electroencephalography (EEG) signals pose significant challenges to reliable decoding. To address these issues, we propose a multi-scale attention-based reconstruction fusion network (MSARFNet) for MI-EEG decoding. The proposed framework employs parallel multi-scale convolutional branches to extract discriminative spatio-temporal features at different temporal resolutions. An attention-based reconstruction fusion module is then introduced to selectively diminish non-dominant information while promoting effective interaction among multi-scale features. Furthermore, local–global temporal encoding strategy is designed to enhance transient MI-related responses through local temporal context aggregation and subsequently capture long-range temporal dependencies via global temporal modeling. Subject-dependent experiments conducted on the BCI Competition IV 2a and 2b datasets demonstrate that MSARFNet achieves average classification accuracies of 84.64% and 87.96%, respectively, outperforming several state-of-the-art methods. These results indicate that MSARFNet provides an effective and robust solution for EEG-based MI decoding.
❌